A diffusion model learns to generate data by first learning to undo a corruption it has seen many times. During training, real examples are gradually buried in noise, and a neural network learns how to step back from each noisier version toward the cleaner one. At generation time, the model starts from pure noise and applies that learned backward step repeatedly until structure emerges. The three foundational papers from 2020 that define this picture are Ho, Jain and Abbeel’s Denoising Diffusion Probabilistic Models, Song and coauthors’ score-based SDE framework, and Song, Meng and Ermon’s Denoising Diffusion Implicit Models. Together they explain why corrupting data is a workable recipe for generating it, and why the same idea can be described as a discrete chain or as a continuous-time process.
Why destroying data can teach a model to create it
The apparent paradox is that noise erases information, yet a model trained on noisy data can produce clean, new examples. The resolution is that the forward corruption is not the part being learned. It is fixed in advance and follows a rule the designer chooses. What the network learns is the reverse direction: given a noisy input and the noise level, which direction moves it back toward the data distribution.
As an Amazon Associate I earn from qualifying purchases.
Song and coauthors put the asymmetry in one sentence: “Creating noise from data is easy; creating data from noise is generative modeling.” The first half is trivial. Adding Gaussian noise to an image requires no knowledge of what images look like. The second half is the hard problem, and the whole diffusion approach is a strategy for breaking it into many small pieces, each of which is simple enough for a network to learn.
Free tools Windows power users keep installed
One-click scans. No signup required.
The forward process: corruption that needs no learning
In the discrete DDPM formulation, the forward process is a Markov chain. Starting from a data point x0, each step mixes in a small amount of Gaussian noise according to a schedule of variances β1 through βT. Each transition is a Gaussian centred on a slightly shrunken copy of the previous sample, with variance βt. After enough steps, xT is statistically close to an isotropic Gaussian, which is the simple prior the generator will later start from.
#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
Three details matter for a reader trying to build intuition:
- The schedule is a design choice. The noise levels and their rate of increase can be tuned, and no single schedule is mandatory. Different choices change how quickly structure is lost and how many steps the reverse process needs.
- Nothing in the forward pass is trained. The continuous-time treatment makes this explicit: the forward SDE does not depend on the data and has no trainable parameters.
- Noisy samples at any chosen time can be generated directly. Because the corruption has a closed form, training can jump to a random time step, add the corresponding noise, and ask the network to predict what it needs to recover.
The reverse process and the role of the score
Running the corruption backward is not a matter of subtracting the exact noise that was added. At generation time the model does not know the noise realization that produced a given sample; it only knows the distribution of plausible clean data at each level of corruption. The reverse dynamics therefore depend on the distributions of the intermediate noisy data, and these are summarised by the score.
The score at time t is the gradient of the log density of the noisy data distribution, ∇x log pt(x). It is a vector field that points toward regions where the noisy data is more probable. A network that estimates this field, or an equivalent target such as the noise that was added, can guide a sample from noise toward the data manifold. Diffusion models differ in exactly which quantity the network predicts, and the parameterizations are related but not interchangeable in every implementation.
Rank #2
Ho, Jain and Abbeel describe their training objective as a weighted variational bound, and they connect it to denoising score matching. In their simplified form, the network is trained to predict the noise that was mixed into a sample at a random time step. This is the sense in which a model “learns to denoise”: it is trained on many corrupted versions and penalised for mistakes in recovering the corruption.
DDPM: a discrete chain of learned reverse steps
Ho, Jain and Abbeel’s paper presents diffusion as a latent variable model inspired by nonequilibrium thermodynamics. Their abstract states: “We present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics.” In the discrete picture, the forward transitions are prescribed, and the reverse transitions are a sequence of learned Gaussian distributions. Sampling walks backward through those transitions from xT to x0.
This is the picture most readers meet first. Each generation step is a stochastic update: the model proposes the mean of the next, slightly cleaner sample, and fresh noise is injected at each step except the last. Because every step depends on the previous one, generation with a DDPM takes as many network evaluations as the number of steps in the chain.
Score-based SDEs: one continuous-time picture
Song and coauthors show that the discrete picture and score-based methods are not rival explanations of unrelated mechanisms. Their paper, Score-Based Generative Modeling through Stochastic Differential Equations, treats the noise levels as a continuum described by a stochastic differential equation. They state that DDPM and score-matching-with-Langevin-dynamics can be seen as discretizations of different SDE choices. For a reader, the practical consequence is that the same generative idea can be written with different time representations and different solvers.
The framework has three components:
- A forward SDE that gradually turns data into noise. It is specified in advance, with a drift term f(x, t) and a diffusion coefficient g(t) that are not learned.
- A reverse-time SDE that runs the corruption backward. Its drift is shifted by the time-dependent score, so the reverse process requires a learned estimate of ∇x log pt(x).
- A numerical sampler that discretises the reverse dynamics. Once a score network is trained, samples are produced by solving the equation numerically.
Reverse-time SDE sampling
The stochastic reverse process injects fresh randomness as it runs backward, in the same spirit as the discrete chain. Its sampling path is an ancestral-style trajectory where each step contains both a deterministic pull toward higher-probability data and a random component. This stochasticity is part of the mechanism, not an accident of implementation.
Predictor-corrector sampling
Song et al. also describe predictor-corrector methods. A predictor takes a step along the reverse dynamics, and a corrector then applies one or more score-based Langevin-style updates at the same noise level to bring the sample closer to the intermediate distribution. The approach trades extra score evaluations for more accurate samples, and the paper’s experiments explore that trade-off rather than declaring a single best setting.
Rank #4
The probability-flow ODE
The same paper derives a probability-flow ordinary differential equation. It has the same marginal distributions over time as the reverse SDE but contains no random term, so starting from a fixed noise sample it produces a deterministic trajectory. In the notation of the paper, the reverse-time SDE and the probability-flow ODE differ by a single coefficient on the score term. Conceptually, the SDE sampler is stochastic and recovers the distribution through randomness at each step, while the ODE is a deterministic map from noise to data that can be inverted, which is useful for tasks that need a consistent encoding of a given image.
DDIM: sampling faster without retraining the model
DDPM’s main practical drawback is that generation is slow. Each sample needs a long chain of evaluations, and Song, Meng and Ermon’s DDIM paper starts from exactly that complaint: DDPMs “require simulating a Markov chain for many steps to produce a sample.” Their answer keeps DDPM’s training procedure but changes the sampling process.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchDDIM defines a family of non-Markovian sampling processes that share the same trained model. Because the trained model is unchanged, a DDPM-trained network can be sampled along a shorter, skipped sequence of time steps. The authors report generation 10× to 50× faster in wall-clock time than DDPM sampling in their experiments, with a trade-off between computation and sample quality. That figure is specific to the paper’s datasets, architectures and step settings; it is not a universal guarantee for every diffusion model or every hardware setup.
Best Value
Two practical points follow. First, fewer steps usually mean less compute but can reduce fidelity, so the number of steps is a setting that needs to be tested for a given model. Second, DDIM’s gain comes from the sampler, not from a new training objective, so it is a change to how an existing model is run rather than a different model.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Comparing the three formulations
| Aspect | DDPM (Ho, Jain, Abbeel, 2020) | Score-SDE (Song et al., 2020) | DDIM (Song, Meng, Ermon, 2020) |
|---|---|---|---|
| Time representation | Discrete Markov steps | Continuous time, described by an SDE | Same trained DDPM model, discrete sampling steps |
| Forward process | Prescribed Markov chain of Gaussian noising steps | Prescribed SDE with no trainable parameters | Inherited from DDPM |
| Learned quantity | Reverse transitions, commonly parameterised through predicted noise | Time-dependent score estimate | Inherited from DDPM training |
| Sampling path | Ancestral reverse chain with injected noise | Reverse-time SDE, predictor-corrector, or probability-flow ODE | Non-Markovian sampling over a shorter step sequence |
| Cost note from source | Many sequential steps | Solver and corrector choices trade evaluations against accuracy; not stated as a single cost value | 10× to 50× wall-clock speedup reported in the paper’s experiments |
The table is a map of concepts, not a ranking. None of these papers establishes a universal winner among samplers, and the trade-offs depend on the model, data and compute budget.
Reading the 2020 benchmark numbers correctly
All three papers report quantitative results, and those numbers are easy to misread as current standings. Each figure below belongs to a specific dataset, model setup and date.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →| Result | Dataset and setting | Source and year | Status |
|---|---|---|---|
| Inception score 9.46; FID 3.17 | Unconditional CIFAR-10 | Ho, Jain, Abbeel, 2020 (DDPM abstract) | Historical result from that paper’s experiments |
| Sample quality described as similar to ProgressiveGAN | 256×256 LSUN | Ho, Jain, Abbeel, 2020 (authors’ comparison) | Authors’ reported comparison, not a current benchmark |
| 10× to 50× faster wall-clock sampling | Experiments in the DDIM paper relative to DDPM sampling | Song, Meng, Ermon, 2020 | Paper-specific; depends on its settings |
| Inception score 9.89; FID 2.20; likelihood 2.99 bits/dim | CIFAR-10 under the score-SDE paper’s experimental description | Song et al., 2020 | Historical result; compare only within the same paper’s setup |
Numbers from different papers should not be placed side by side as if they came from one controlled test. The 2020 papers establish conceptual foundations; they do not describe the latest models, the best current samplers, or modern text-to-image systems, which build on these ideas with different architectures and training choices.
Quick Recap
Common misreadings to avoid
- “The model memorises the noise.” The network is trained to predict or estimate a denoising target across many noise levels; generation does not reverse a stored noise sequence.
- “DDPM and score-based diffusion are different methods.” The score-SDE paper presents them as related, with DDPM appearing as a discretization of a particular SDE choice.
- “DDIM retrains the model to be faster.” DDIM keeps DDPM’s training procedure and changes how sampling is performed.
- “The probability-flow ODE is a different generative model.” It is a deterministic route within the same framework that shares the marginal distributions of the reverse SDE.
Where to read the primary papers
- Ho, Jain and Abbeel, Denoising Diffusion Probabilistic Models, NeurIPS 2020 proceedings abstract page.
- Song et al., Score-Based Generative Modeling through Stochastic Differential Equations, arXiv:2011.13456 (2020).
- Song, Meng and Ermon, Denoising Diffusion Implicit Models, arXiv:2010.02502 (2020).
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




