🎨 GANs, Diffusion & Flow Matching
Follow adversarial learning, denoising and flow matching through small calculations, then compare their sampling and deployment costs.
On this page
Before you start
Review Probability & Statistics for distributions and expectations, Calculus & Optimization for gradients, and Deep Learning Basics for neural-network training. Numerical Computing helps with integration error and precision.
You will learn to distinguish a training objective from a sampler, compute one noising/denoising example and one flow trajectory, and assess quality, coverage and generation cost without assuming a universal winner.
The problem: generate more than one plausible answer
A generator must produce varied samples from a distribution, not reconstruct one memorized image. As a toy dataset, imagine equal probability near −2 and +2 on a line. Producing only +2 looks realistic for each sample but misses half the distribution. Producing the mean, zero, represents neither mode.
GANs, diffusion and flow matching offer different ways to learn a sampling process. These are overlapping research families, not clean eras in which one made every earlier technique obsolete. Architecture is another axis: a transformer or convolutional network can parameterize several of these objectives.
GANs: learn through a critic
A generator G(z) maps random input z to a sample. A discriminator D learns to distinguish real samples from generated ones. The original objective is:
min_G max_D E_real[log D(x)] + E_z[log(1 − D(G(z)))]
D's output is a probability in this formulation. Under ideal capacity and optimization assumptions, the distribution-matching solution has p_G=p_data and an optimal discriminator equal to 0.5 on their support. Finite neural networks and alternating updates do not guarantee reaching that solution.
If the generator covers only the +2 mode, the discriminator's current weakness may reward it while the missing −2 mode remains absent. This mode collapse is a failure of learned distribution coverage, not proof that the theoretical objective prefers collapse.
The original saturating generator loss can provide weak gradients when D confidently rejects generated samples. A commonly used non-saturating generator loss, −E[log D(G(z))], supplies a different training gradient while retaining the desired distribution-matching target. WGAN uses a real-valued critic with Lipschitz constraints to estimate a Wasserstein objective; gradient-penalty variants regularize that constraint on sampled points rather than guarantee it globally.
Diffusion: make supervised denoising tasks
Start from clean data x_0. Let β_t be a variance schedule, α_t=1−β_t, and α_bar_t the product of α values up to noise level t. For a common Gaussian forward process:
x_t = sqrt(α_bar_t) × x_0 + sqrt(1−α_bar_t) × ε, with ε ∼ N(0,I).
The scalars α and β are dimensionless; x and ε are expressed in the model's normalized data coordinates. A network can learn to predict ε from x_t, t and optional conditioning c, minimizing an expected squared noise error. Other parameterizations predict clean data, a score or another velocity-like target with suitable transformations and weighting.
A one-coordinate calculation
Take x_0=2, α_bar_t=0.64 and sampled ε=−1. Then x_t=0.8×2+0.6×(−1)=1.
Given predicted noise ε_hat=−0.5, the implied clean estimate is:
x_0_hat = (x_t − sqrt(1−α_bar_t)×ε_hat)/sqrt(α_bar_t)
x_0_hat = (1−0.6×(−0.5))/0.8 = 1.625.
With the actual sampled noise −1, the same algebra recovers 2. The network does not get to observe that noise label at inference; it estimates from a noisy input that may admit many possible clean sources. This calculation explains the parameterization, not a complete reverse sampler.
The sampler is a separate choice
A DDPM-style reverse process uses learned predictions in a stochastic transition chain. DDIM constructs a related family of sampling updates; its zero-noise setting is deterministic conditional on the initial latent and numerical execution. DDIM can also include stochasticity. Both quality and runtime depend on selected noise levels, model accuracy and the discretization.
A checkpoint trained over many noise levels need not use all of them at inference. Shortening a schedule changes approximation error; neither “DDPM must take 1,000 steps” nor “DDIM always needs 20” is a general rule. Solver order and network-function evaluations (NFEs) are different counts: some steps evaluate the model more than once.
Latent diffusion adds an autoencoder. Instead of corrupting pixels directly, encode the image, generate in latent coordinates and decode afterward. A 512×512×3 image has 786,432 scalar entries; a 64×64×4 latent has 16,384, a 48-fold reduction in scalar count. It is not a 48-fold compute guarantee: channels, network structure, attention, encoding and decoding determine cost. The autoencoder can discard detail that later denoising cannot restore faithfully.
Conditioning: classifier-free guidance
Train conditional and unconditional/null-condition behavior, often by dropping conditions on some training examples. At inference, a common noise-prediction convention is:
ε_guided = ε_uncond + w × (ε_cond − ε_uncond).
Here w=0 uses the unconditional prediction; w=1 uses the ordinary conditional prediction; w>1 extrapolates the conditioning difference. If ε_uncond=1, ε_cond=0.5 and w=3, the result is −0.5. It need not lie between the two predictions.
More guidance may improve a conditioning metric while reducing diversity or introducing artifacts. The useful scale depends on the model and convention. The two prediction branches also cost compute, even if they are batched in one call. Guidance changes the sampling process; it does not guarantee exact sampling from a simple powered version of the final conditional data distribution.
Flow matching: learn how to move a sample
Use different symbols to avoid reversing diffusion notation: y_0 is noise, y_1 is data, and τ runs from 0 to 1. In a simple linear conditional path:
y_τ=(1−τ)y_0+τy_1, with target velocity u=y_1−y_0.
Train a vector field v_θ(y_τ,τ,c) to predict that velocity from sampled path points. Sampling solves dy/dτ=v_θ(y,τ,c), an ordinary differential equation. τ is a dimensionless path coordinate, not seconds of physical motion.
For y_0=−1 and y_1=3, the conditional velocity is 4. Four Euler steps of Δτ=0.25 trace −1 → 0 → 1 → 2 → 3. This straight, constant-velocity toy case is exactly integrated by Euler. Real learned fields vary with state and time; averaging conditional paths does not guarantee globally straight sampling trajectories or one-step accuracy.
Flow matching permits several probability paths, including diffusion-related ones. Path choice, coupling between noise/data, time sampling, weighting and numerical solver still require decisions. Rectified-flow reflow and consistency training/distillation can improve few-step generation in particular setups; they do not guarantee one-step quality for an arbitrary checkpoint.
Keep architecture and acceleration in their own boxes
Convolutional GAN designs such as DCGAN add image-oriented architectural choices; style-based generators such as StyleGAN map latent inputs into controls applied at multiple layers. Style mixing, modulation/demodulation and alias-aware signal processing address different control or image-consistency problems. A style coordinate is not automatically a disentangled, independently editable real-world attribute.
For WGAN-GP, a typical penalty is λ(‖∇D(x_hat)‖₂−1)² at sampled interpolations between real and generated inputs. It regularizes the critic's input gradient where sampled, while leaving approximation and optimization limits. Its λ is a penalty weight, not the learning rate.
A UNet and a diffusion transformer are alternative backbones for predicting noise, scores or velocities. Their names do not specify the probability path or sampler. Rectified-flow reflow uses trajectories from a learned model to form a new coupling and retrain; improvement still needs measurement. Consistency models learn a mapping that agrees on points along a trajectory and may be trained directly or distilled from a teacher. Few-step models can also use other distillation objectives, including adversarial ones; a fast product name does not identify its training method.
The advanced connection: scores and probability flow
A score is s(x,t)=∇_x log p_t(x), the gradient of log density at noise level t. For a forward SDE with drift f and scalar, state-independent diffusion g(t), the reverse-time SDE uses drift f−g²s when integrated backward in t. A corresponding probability-flow ODE has drift f−0.5g²s.
With the exact score and appropriate regularity, they share time-marginal distributions, not identical sample paths. Learned score error and finite-step integration break exact equivalence. Flow matching trains a transport field directly for a chosen path; the connection is useful without claiming every diffusion and flow recipe is identical or that all conditional paths are straight.
Evaluate the distribution and the system
Inspect fidelity, diversity/coverage, prompt or spatial-condition adherence and memorization on independent data. FID summarizes feature-distribution differences and depends on representation and sample size; it cannot certify prompt correctness or expose every collapsed mode. Human preference and task-specific checks supply different evidence.
For deployment, report resolution, model/precision, sampler, NFE, guidance, batch and hardware with latency/throughput. A one-step sampler can still be expensive, and a lower FID is better under its defined protocol. Avoid private-architecture guesses from a product's visual style or marketing name.
Check yourself
Does halving the number of sampler steps necessarily halve end-to-end response time?
Solution: No. A step can contain multiple NFEs, guidance can add prediction branches, and text encoding, autoencoder decoding, queues and transfer remain. Measure the whole path at matched quality and load.
Where to go next
World Models uses generative prediction for environment futures. Vision-Language-Action Models uses flow/diffusion ideas for action chunks. Model Serving covers capacity and latency measurement.
References
- Generative Adversarial Networks: the original adversarial objective and idealized solution.
- Denoising Diffusion Probabilistic Models: Gaussian noising and denoising training.
- Score-Based Generative Modeling through SDEs: reverse-time dynamics and probability-flow ODEs.
- Flow Matching for Generative Modeling: simulation-free vector-field regression over chosen paths.