Diffusion Models for Image Generation: DDPM, Latent Diffusion, and Classifier Guidance
The Intuition Behind Diffusion
Diffusion models generate images by learning to reverse a gradual noising process. Imagine dropping ink into water: the ink spreads and eventually becomes uniformly distributed. If you could reverse this process perfectly, you could start from uniform noise and recover the original ink pattern. Diffusion models learn exactly this reversal, but in the space of images.
The forward process adds Gaussian noise to an image over T timesteps until it becomes indistinguishable from pure noise. The reverse process, parameterized by a neural network, learns to predict and remove the noise added at each step. At generation time, you start with pure noise and apply the learned denoising network T times, gradually refining random noise into a coherent image.
This approach has fundamental advantages over GANs and VAEs. There's no adversarial training instability, no mode collapse, and the training objective is a simple regression loss. The tradeoff is computational cost: generating one image requires hundreds or thousands of forward passes through the neural network, compared to a single pass for GANs.
DDPM: The Mathematical Foundation
Denoising Diffusion Probabilistic Models formalize this intuition. The forward process defines a Markov chain that gradually adds noise to data according to a variance schedule beta_1, ..., beta_T:
q(x_t | x_{t-1}) = N(x_t; sqrt(1 - beta_t) * x_{t-1}, beta_t * I)
A key mathematical property simplifies training enormously: you can sample x_t at any timestep t directly from x_0 without computing intermediate steps. Defining alpha_t = 1 - beta_t and alpha_bar_t as the cumulative product of alphas:
q(x_t | x_0) = N(x_t; sqrt(alpha_bar_t) * x_0, (1 - alpha_bar_t) * I)
This means x_t = sqrt(alpha_bar_t) * x_0 + sqrt(1 - alpha_bar_t) * epsilon, where epsilon is standard Gaussian noise. Training samples a random timestep, applies the corresponding noise level, and trains the network to predict the noise that was added.
import torch
import torch.nn as nn
import numpy as np
class DDPMSchedule:
"""DDPM noise schedule and sampling utilities."""
def __init__(self, T: int = 1000, schedule: str = "cosine"):
self.T = T
if schedule == "linear":
self.betas = torch.linspace(1e-4, 0.02, T)
elif schedule == "cosine":
# Cosine schedule from Nichol & Dhariwal
steps = torch.arange(T + 1, dtype=torch.float64) / T
alpha_bar = torch.cos((steps + 0.008) / 1.008 * np.pi / 2) ** 2
alpha_bar = alpha_bar / alpha_bar[0]
betas = 1 - alpha_bar[1:] / alpha_bar[:-1]
self.betas = torch.clamp(betas, max=0.999).float()
self.alphas = 1.0 - self.betas
self.alpha_bar = torch.cumprod(self.alphas, dim=0)
self.sqrt_alpha_bar = torch.sqrt(self.alpha_bar)
self.sqrt_one_minus_alpha_bar = torch.sqrt(1.0 - self.alpha_bar)
def q_sample(self, x_0: torch.Tensor, t: torch.Tensor,
noise: torch.Tensor = None) -> torch.Tensor:
"""Sample x_t from q(x_t | x_0) directly."""
if noise is None:
noise = torch.randn_like(x_0)
sqrt_ab = self.sqrt_alpha_bar[t][:, None, None, None]
sqrt_1_ab = self.sqrt_one_minus_alpha_bar[t][:, None, None, None]
return sqrt_ab * x_0 + sqrt_1_ab * noise
def diffusion_training_step(model, x_0, schedule, optimizer):
"""One training step for DDPM."""
batch_size = x_0.shape[0]
# Sample random timesteps
t = torch.randint(0, schedule.T, (batch_size,), device=x_0.device)
# Sample noise and create noisy input
noise = torch.randn_like(x_0)
x_t = schedule.q_sample(x_0, t, noise)
# Predict noise
predicted_noise = model(x_t, t)
# Simple MSE loss on noise prediction
loss = nn.functional.mse_loss(predicted_noise, noise)
optimizer.zero_grad()
loss.backward()
optimizer.step()
return loss.item()
The training objective is remarkably simple: predict the noise that was added. Despite this simplicity, the model implicitly learns the score function (gradient of the log probability) of the data distribution at every noise level, which provides the information needed to denoise step by step.
For the normalization layers used inside the U-Net, see Batch Normalization vs Layer Normalization in Transformers.
The U-Net: Architecture of the Denoiser
The neural network that predicts noise in diffusion models is almost always a U-Net: an encoder-decoder architecture with skip connections between corresponding encoder and decoder layers. The skip connections are critical because they allow the network to preserve spatial information that would otherwise be lost in the downsampling bottleneck.
Modern diffusion U-Nets differ from the original medical imaging U-Net in several important ways. They incorporate attention mechanisms (self-attention at lower resolutions and cross-attention for conditioning), they use group normalization instead of batch normalization, and they condition on the timestep through an embedding that modulates the network's behavior at each layer.
Timestep conditioning is typically done through sinusoidal embeddings (similar to positional encodings in transformers), which are projected and added to the hidden representations. Some architectures use adaptive layer normalization (AdaLN) where the timestep embedding modulates the scale and shift parameters of the normalization layers, giving the network more expressive control over how its behavior changes across timesteps.
The attention mechanisms in the U-Net leverage the same principles as Attention Mechanism Variants Beyond Self-Attention, applied to 2D spatial representations rather than 1D sequences.
Latent Diffusion: Trading Pixels for Efficiency
The computational bottleneck of pixel-space diffusion motivated Latent Diffusion Models (LDMs), the architecture behind Stable Diffusion. Instead of running the diffusion process on high-resolution images, LDMs first encode images into a compressed latent space using a pretrained VAE, run diffusion in that latent space, then decode the result back to pixel space.
A 512x512x3 image encodes to a 64x64x4 latent representation — a 48x compression in spatial dimensions. Since the U-Net's compute scales quadratically with spatial resolution, this compression reduces the per-step compute by roughly 2300x. The VAE is trained separately with a perceptual loss and a KL regularization term that keeps the latent space smooth and well-structured for the diffusion process to operate on.
The separation of perceptual compression (VAE) from semantic composition (diffusion) is architecturally elegant. The VAE handles the low-level details — texture, exact pixel values, high-frequency information — while the diffusion model focuses on the high-level composition: object placement, scene structure, semantic content. Each component can be optimized independently for its specific job.
Efficient training of these large models benefits from the techniques discussed in Distributed Training with DeepSpeed ZeRO Configuration.
Classifier-Free Guidance
Classifier-free guidance was the breakthrough that made text-to-image diffusion practical. Earlier work used classifier guidance, which required training a separate classifier on noisy images at every noise level and using its gradients to steer the diffusion process toward a target class. This was cumbersome and limited to predefined class labels.
Classifier-free guidance eliminates the external classifier. During training, the text conditioning is randomly dropped (replaced with a null embedding) for some fraction of samples, typically 10%. This trains the model to generate both conditionally (with text) and unconditionally (without text). At inference, both predictions are computed, and the final prediction extrapolates beyond the conditional prediction:
epsilon_guided = epsilon_uncond + w * (epsilon_cond - epsilon_uncond)
The guidance scale w controls the strength of the conditioning. At w=1, you get the standard conditional prediction. At w=7-15 (common for text-to-image), the model strongly follows the text prompt, producing sharper, more prompt-adherent images at the cost of reduced diversity. Values above 15-20 tend to produce oversaturated, artifact-heavy images.
@torch.no_grad()
def guided_sampling_step(model, x_t, t, text_embedding,
null_embedding, guidance_scale=7.5,
schedule=None):
"""One step of classifier-free guided DDIM sampling."""
# Conditional prediction
noise_cond = model(x_t, t, text_embedding)
# Unconditional prediction
noise_uncond = model(x_t, t, null_embedding)
# Classifier-free guidance
noise_pred = noise_uncond + guidance_scale * (noise_cond - noise_uncond)
# DDIM update step
alpha_bar_t = schedule.alpha_bar[t]
alpha_bar_prev = schedule.alpha_bar[t - 1] if t > 0 else torch.tensor(1.0)
# Predict x_0
x_0_pred = (x_t - torch.sqrt(1 - alpha_bar_t) * noise_pred) / \
torch.sqrt(alpha_bar_t)
x_0_pred = torch.clamp(x_0_pred, -1, 1)
# Compute x_{t-1} (deterministic DDIM)
x_prev = torch.sqrt(alpha_bar_prev) * x_0_pred + \
torch.sqrt(1 - alpha_bar_prev) * noise_pred
return x_prev
DDIM: Faster Sampling Without Retraining
DDPM sampling is slow because it requires iterating through all T timesteps. DDIM (Denoising Diffusion Implicit Models) showed that the same trained model can be sampled with a non-Markovian process that skips steps. Instead of going through all 1000 timesteps, you can define a subsequence (say, every 20th timestep) and sample along that subset.
The key insight is that the DDPM training objective doesn't actually require Markovian sampling. The noise prediction model learns to estimate the original clean image from any noise level, and DDIM exploits this by taking larger steps in the denoising direction. With 50 DDIM steps, image quality is close to 1000-step DDPM quality. With 20 steps, there's noticeable quality degradation but the results are still useful for rapid iteration during development.
DDIM also has the advantage of deterministic sampling: given the same starting noise, it always produces the same output. This makes interpolation in the noise space meaningful. You can smoothly interpolate between two noise vectors and the generated images will smoothly morph between the two corresponding outputs. DDPM's stochastic sampling doesn't have this property.
Text Conditioning Architecture
Text-to-image diffusion models need a way to inject text information into the U-Net. The standard approach uses a pretrained text encoder (CLIP or T5) to convert the text prompt into a sequence of token embeddings, which are then injected into the U-Net through cross-attention layers.
Stable Diffusion 1.x uses CLIP's text encoder, which produces 77-token embeddings of 768 dimensions. Stable Diffusion XL uses both CLIP-L and OpenCLIP-G text encoders, concatenating their outputs for richer text representations. The switch to T5 text encoders (as in Imagen and PixArt) provides even stronger text understanding because T5 is trained on more diverse text data than CLIP, which was trained primarily on image-text pairs.
For a comparison of the text encoders used in these pipelines, see Embedding Models Comparison.
Practical Training Considerations
Training diffusion models involves several practical choices that significantly affect quality:
- Loss weighting: The simple MSE loss weights all timesteps equally, but different timesteps contribute differently to image quality. Weighting the loss by 1/SNR (signal-to-noise ratio) at each timestep focuses training on the most informative noise levels. The "min-SNR" weighting strategy clips the maximum weight to prevent extreme gradients at low noise levels.
- EMA (Exponential Moving Average): Maintaining an exponential moving average of the model weights during training produces smoother, higher-quality outputs. The EMA model is used for inference while the online model continues training. Decay rates of 0.9999 are standard.
- Mixed precision: FP16 or BF16 training with FP32 accumulation is essential for large models. BF16 is preferred over FP16 because it avoids the need for loss scaling and handles the wide range of values in diffusion models more gracefully.
- Noise offset: Adding a small offset (0.1) to the noise at each timestep improves the model's ability to generate very dark or very bright images, addressing a known bias in standard diffusion training toward medium-brightness outputs.
Data augmentation strategies specific to image generation are explored in Data Augmentation Strategies for Model Robustness.
Comparison of Diffusion Model Variants
| Model | Space | Steps (default) | Text Encoder | Resolution |
|---|---|---|---|---|
| DDPM (original) | Pixel | 1000 | None (unconditional) | 256x256 |
| Guided Diffusion | Pixel | 250 | Classifier | 256-512 |
| DALL-E 2 | CLIP latent | 100 | CLIP | 1024x1024 |
| Stable Diffusion 1.5 | VAE latent | 50 | CLIP-L | 512x512 |
| Stable Diffusion XL | VAE latent | 40 | CLIP-L + OpenCLIP-G | 1024x1024 |
| Imagen | Pixel (cascaded) | ~1000 | T5-XXL | 1024x1024 |
| PixArt-alpha | VAE latent | 20 | T5-XXL | 1024x1024 |
Beyond Static Images: Current Frontiers
The diffusion framework extends naturally to video generation, 3D shape synthesis, and even audio. Video diffusion models add temporal attention layers between the spatial layers, operating on 3D (spatial + temporal) instead of 2D representations. The core challenge is maintaining temporal consistency: individual frames should be high quality AND flow smoothly together.
Consistency models and progressive distillation are pushing the sampling step count toward one. Consistency models learn to map any point on a diffusion trajectory directly to the clean data point, enabling single-step generation. The quality gap compared to multi-step sampling is closing rapidly, and 2-4 step generation is already competitive with 20-step DDIM for many applications.
For deploying these models efficiently, quantization plays a crucial role. See Quantization-Aware vs Post-Training Benchmarks for strategies that maintain image quality while reducing model size and inference cost.