October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk8 min

Diffusion Models Explained: From Noise Corruption to Reverse Generation

Diffusion models learn to undo a fixed noise corruption and generate data by running that reversal from pure noise. Here is how DDPM, the score-based SDE framework and DDIM fit together.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A diffusion model learns to generate data by first learning to undo a corruption it has seen many times. During training, real examples are gradually buried in noise, and a neural network learns how to step back from each noisier version toward the cleaner one. At generation time, the model starts from pure noise and applies that learned backward step repeatedly until structure emerges. The three foundational papers from 2020 that define this picture are Ho, Jain and Abbeel’s Denoising Diffusion Probabilistic Models, Song and coauthors’ score-based SDE framework, and Song, Meng and Ermon’s Denoising Diffusion Implicit Models. Together they explain why corrupting data is a workable recipe for generating it, and why the same idea can be described as a discrete chain or as a continuous-time process.

Why destroying data can teach a model to create it

The apparent paradox is that noise erases information, yet a model trained on noisy data can produce clean, new examples. The resolution is that the forward corruption is not the part being learned. It is fixed in advance and follows a rule the designer chooses. What the network learns is the reverse direction: given a noisy input and the noise level, which direction moves it back toward the data distribution.

As an Amazon Associate I earn from qualifying purchases.

Song and coauthors put the asymmetry in one sentence: “Creating noise from data is easy; creating data from noise is generative modeling.” The first half is trivial. Adding Gaussian noise to an image requires no knowledge of what images look like. The second half is the hard problem, and the whole diffusion approach is a strategy for breaking it into many small pieces, each of which is simple enough for a network to learn.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The forward process: corruption that needs no learning

In the discrete DDPM formulation, the forward process is a Markov chain. Starting from a data point x0, each step mixes in a small amount of Gaussian noise according to a schedule of variances β1 through βT. Each transition is a Gaussian centred on a slightly shrunken copy of the previous sample, with variance βt. After enough steps, xT is statistically close to an isotropic Gaussian, which is the simple prior the generator will later start from.

#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Three details matter for a reader trying to build intuition:

  • The schedule is a design choice. The noise levels and their rate of increase can be tuned, and no single schedule is mandatory. Different choices change how quickly structure is lost and how many steps the reverse process needs.
  • Nothing in the forward pass is trained. The continuous-time treatment makes this explicit: the forward SDE does not depend on the data and has no trainable parameters.
  • Noisy samples at any chosen time can be generated directly. Because the corruption has a closed form, training can jump to a random time step, add the corresponding noise, and ask the network to predict what it needs to recover.

The reverse process and the role of the score

Running the corruption backward is not a matter of subtracting the exact noise that was added. At generation time the model does not know the noise realization that produced a given sample; it only knows the distribution of plausible clean data at each level of corruption. The reverse dynamics therefore depend on the distributions of the intermediate noisy data, and these are summarised by the score.

The score at time t is the gradient of the log density of the noisy data distribution, ∇x log pt(x). It is a vector field that points toward regions where the noisy data is more probable. A network that estimates this field, or an equivalent target such as the noise that was added, can guide a sample from noise toward the data manifold. Diffusion models differ in exactly which quantity the network predicts, and the parameterizations are related but not interchangeable in every implementation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Ho, Jain and Abbeel describe their training objective as a weighted variational bound, and they connect it to denoising score matching. In their simplified form, the network is trained to predict the noise that was mixed into a sample at a random time step. This is the sense in which a model “learns to denoise”: it is trained on many corrupted versions and penalised for mistakes in recovering the corruption.

DDPM: a discrete chain of learned reverse steps

Ho, Jain and Abbeel’s paper presents diffusion as a latent variable model inspired by nonequilibrium thermodynamics. Their abstract states: “We present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics.” In the discrete picture, the forward transitions are prescribed, and the reverse transitions are a sequence of learned Gaussian distributions. Sampling walks backward through those transitions from xT to x0.

This is the picture most readers meet first. Each generation step is a stochastic update: the model proposes the mean of the next, slightly cleaner sample, and fresh noise is injected at each step except the last. Because every step depends on the previous one, generation with a DDPM takes as many network evaluations as the number of steps in the chain.

Score-based SDEs: one continuous-time picture

Song and coauthors show that the discrete picture and score-based methods are not rival explanations of unrelated mechanisms. Their paper, Score-Based Generative Modeling through Stochastic Differential Equations, treats the noise levels as a continuum described by a stochastic differential equation. They state that DDPM and score-matching-with-Langevin-dynamics can be seen as discretizations of different SDE choices. For a reader, the practical consequence is that the same generative idea can be written with different time representations and different solvers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The framework has three components:

  1. A forward SDE that gradually turns data into noise. It is specified in advance, with a drift term f(x, t) and a diffusion coefficient g(t) that are not learned.
  2. A reverse-time SDE that runs the corruption backward. Its drift is shifted by the time-dependent score, so the reverse process requires a learned estimate of ∇x log pt(x).
  3. A numerical sampler that discretises the reverse dynamics. Once a score network is trained, samples are produced by solving the equation numerically.

Reverse-time SDE sampling

The stochastic reverse process injects fresh randomness as it runs backward, in the same spirit as the discrete chain. Its sampling path is an ancestral-style trajectory where each step contains both a deterministic pull toward higher-probability data and a random component. This stochasticity is part of the mechanism, not an accident of implementation.

Predictor-corrector sampling

Song et al. also describe predictor-corrector methods. A predictor takes a step along the reverse dynamics, and a corrector then applies one or more score-based Langevin-style updates at the same noise level to bring the sample closer to the intermediate distribution. The approach trades extra score evaluations for more accurate samples, and the paper’s experiments explore that trade-off rather than declaring a single best setting.

The probability-flow ODE

The same paper derives a probability-flow ordinary differential equation. It has the same marginal distributions over time as the reverse SDE but contains no random term, so starting from a fixed noise sample it produces a deterministic trajectory. In the notation of the paper, the reverse-time SDE and the probability-flow ODE differ by a single coefficient on the score term. Conceptually, the SDE sampler is stochastic and recovers the distribution through randomness at each step, while the ODE is a deterministic map from noise to data that can be inverted, which is useful for tasks that need a consistent encoding of a given image.

DDIM: sampling faster without retraining the model

DDPM’s main practical drawback is that generation is slow. Each sample needs a long chain of evaluations, and Song, Meng and Ermon’s DDIM paper starts from exactly that complaint: DDPMs “require simulating a Markov chain for many steps to produce a sample.” Their answer keeps DDPM’s training procedure but changes the sampling process.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

DDIM defines a family of non-Markovian sampling processes that share the same trained model. Because the trained model is unchanged, a DDPM-trained network can be sampled along a shorter, skipped sequence of time steps. The authors report generation 10× to 50× faster in wall-clock time than DDPM sampling in their experiments, with a trade-off between computation and sample quality. That figure is specific to the paper’s datasets, architectures and step settings; it is not a universal guarantee for every diffusion model or every hardware setup.

Two practical points follow. First, fewer steps usually mean less compute but can reduce fidelity, so the number of steps is a setting that needs to be tested for a given model. Second, DDIM’s gain comes from the sampler, not from a new training objective, so it is a change to how an existing model is run rather than a different model.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Comparing the three formulations

Aspect DDPM (Ho, Jain, Abbeel, 2020) Score-SDE (Song et al., 2020) DDIM (Song, Meng, Ermon, 2020)
Time representation Discrete Markov steps Continuous time, described by an SDE Same trained DDPM model, discrete sampling steps
Forward process Prescribed Markov chain of Gaussian noising steps Prescribed SDE with no trainable parameters Inherited from DDPM
Learned quantity Reverse transitions, commonly parameterised through predicted noise Time-dependent score estimate Inherited from DDPM training
Sampling path Ancestral reverse chain with injected noise Reverse-time SDE, predictor-corrector, or probability-flow ODE Non-Markovian sampling over a shorter step sequence
Cost note from source Many sequential steps Solver and corrector choices trade evaluations against accuracy; not stated as a single cost value 10× to 50× wall-clock speedup reported in the paper’s experiments

The table is a map of concepts, not a ranking. None of these papers establishes a universal winner among samplers, and the trade-offs depend on the model, data and compute budget.

Reading the 2020 benchmark numbers correctly

All three papers report quantitative results, and those numbers are easy to misread as current standings. Each figure below belongs to a specific dataset, model setup and date.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Result Dataset and setting Source and year Status
Inception score 9.46; FID 3.17 Unconditional CIFAR-10 Ho, Jain, Abbeel, 2020 (DDPM abstract) Historical result from that paper’s experiments
Sample quality described as similar to ProgressiveGAN 256×256 LSUN Ho, Jain, Abbeel, 2020 (authors’ comparison) Authors’ reported comparison, not a current benchmark
10× to 50× faster wall-clock sampling Experiments in the DDIM paper relative to DDPM sampling Song, Meng, Ermon, 2020 Paper-specific; depends on its settings
Inception score 9.89; FID 2.20; likelihood 2.99 bits/dim CIFAR-10 under the score-SDE paper’s experimental description Song et al., 2020 Historical result; compare only within the same paper’s setup

Numbers from different papers should not be placed side by side as if they came from one controlled test. The 2020 papers establish conceptual foundations; they do not describe the latest models, the best current samplers, or modern text-to-image systems, which build on these ideas with different architectures and training choices.

Common misreadings to avoid

  • “The model memorises the noise.” The network is trained to predict or estimate a denoising target across many noise levels; generation does not reverse a stored noise sequence.
  • “DDPM and score-based diffusion are different methods.” The score-SDE paper presents them as related, with DDPM appearing as a discretization of a particular SDE choice.
  • “DDIM retrains the model to be faster.” DDIM keeps DDPM’s training procedure and changes how sampling is performed.
  • “The probability-flow ODE is a different generative model.” It is a deterministic route within the same framework that shares the marginal distributions of the reverse SDE.

Where to read the primary papers

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.