Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Repeatedly applying a simple moving average changes its equal weights into a symmetric, increasingly bell-shaped pattern. The reason is convolution: the effective weight at each lag counts how many combinations of window positions reach that lag. Those normalized weights are also the probability distribution of a sum of independent discrete uniform variables, so the Central Limit Theorem explains why their standardized shape approaches a Gaussian. That result describes the filter’s weights—not, by itself, the distribution of the data being smoothed.
From equal weights to a bell-shaped filter
A length-m trailing moving average is
y[t] = (x[t] + x[t−1] + ··· + x[t−m+1]) / m.
Each of the m observations receives weight 1/m. Apply the same average again and observations near the middle of the combined span contribute through more overlapping windows than observations near its ends. The resulting weights are therefore unequal.
For a three-point average, the unnormalized coefficient patterns are:
#1 Best Overall
- This guide is a perfect overview for the topics covered in introductory statistics courses.
- One pass:
(1, 1, 1), normalized by 3. - Two passes:
(1, 2, 3, 2, 1), normalized by 9. - Three passes:
(1, 3, 6, 7, 6, 3, 1), normalized by 27. - Four passes:
(1, 4, 10, 16, 19, 16, 10, 4, 1), normalized by 81.
These are “natural weights” only in the sense that repeated equal-weight averaging produces them. The phrase does not mean they are universally optimal or inherently better than other smoothing weights.
Convolution generates the weights
Represent one pass by the finite kernel
h[j] = 1/m for j = 0, …, m−1, and zero otherwise. One pass convolves the input series with this kernel; r passes convolve it with the same kernel r times. If * denotes convolution, the composite kernel is h*r.
A compact way to calculate its coefficients is with a generating polynomial. The unnormalized kernel for one pass is 1 + z + ··· + zm−1. After r passes:
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →(1 + z + ··· + zm−1)r.
The coefficient of zj, divided by mr, is the composite weight at lag j:
wr,m(j) = [zj](1 + z + ··· + zm−1)r / mr, for 0 ≤ j ≤ r(m−1).
The support has r(m−1)+1 positions. All weights are nonnegative, sum to one, and are symmetric around the center: w(j) = w(r(m−1)−j). The coefficients count the number of ways to choose r integers from 0 through m−1 whose sum is j.
The probability distribution inside the filter
Let U1, …, Ur be independent random variables, each equally likely to be any integer from 0 to m−1, and let Sr = U1 + ··· + Ur. Then the composite filter weight is exactly
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
wr,m(j) = P(Sr = j).
This is a mathematical interpretation of the kernel; it does not assume that the observed time-series values are independent. The distribution of each index choice has mean (m−1)/2 and variance (m2−1)/12. Independence gives the sum’s mean and variance:
Rank #2
E[Sr] = r(m−1)/2, Var(Sr) = r(m2−1)/12.
The mean is the center of the weights. In a causal, trailing implementation it is also the filter’s delay in samples. A centered offline filter can align that center with the target observation instead, although it then uses future observations relative to that target.
Why the Central Limit Theorem applies to the weights
The ordinary Central Limit Theorem says that a sum of many independent, identically distributed variables with finite variance approaches a normal distribution after centering and scaling. Here, each Ui is bounded and has finite variance, so
(Sr − r(m−1)/2) / √(r(m2−1)/12) ⇒ N(0,1) as r → ∞.
Recommended Free Tools
Consequently, the normalized discrete weights become approximately Gaussian in shape as the number of passes grows. They remain an exact finite, discrete distribution at any finite number of passes; for small r, the Gaussian description may be poor, especially near the support edges. Use the exact coefficients when the iteration count is small.
The continuous counterpart begins with a uniform density on an interval. Repeated convolution produces piecewise-polynomial densities in the Irwin–Hall family; the standardized shapes approach a normal density as the number of convolutions grows. Repeated convolution of box functions is also related to cardinal B-spline kernels, though that connection does not make every moving-average filter a general B-spline construction. For broader spline terminology, see Wolfram MathWorld’s B-spline reference.
Why convolution and the CLT fit together
Convolution in the original domain corresponds to multiplication in the characteristic-function domain. For a discrete uniform variable,
φU(t) = (1/m) Σj=0m−1 eijt = ei(m−1)t/2 sin(mt/2)/(m sin(t/2)),
and independence gives φSr(t) = φU(t)r. Near zero, the centered and variance-scaled characteristic function approaches e−t²/2, the characteristic function of the standard normal. This product property is the analytic mechanism behind the familiar convolution-based explanation of the CLT. See the Encyclopedia of Mathematics entry on characteristic functions for the general identity.
Rank #3
What smoothing does to independent noise
Suppose the input consists of independent noise samples, each with variance σ², and a normalized filter produces Y = ΣjwjXj. Then
Var(Y) = σ² Σjwj².
A single equal-weight window has variance σ²/m. After repeated passes, use the squared composite weights, not σ²/mr. Repeating the operation does not create mr independent observations: its kernel spans only r(m−1)+1 distinct input positions, with unequal weights.
A useful summary is the effective sample size, defined for normalized weights as Neff = 1 / Σwj². For many passes, the Gaussian approximation to the weights gives
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchΣwr,m(j)² ≈ √(3 / (πr(m²−1))), and therefore Neff ≈ √(πr(m²−1)/3).
This approximation is asymptotic, not an exact formula for every window and pass count. It highlights a practical point: as passes increase, effective sample size grows roughly as √r, not linearly with r. Smoothing reduces variation, but with diminishing returns.
Real observations are often dependent. In that case the output variance is instead
Var(Y) = Σj,kwjwk Cov(Xt−j, Xt−k).
Ignoring those covariances can misstate noise reduction. Outputs from overlapping windows are correlated even if the original samples are independent, so treating every smoothed output as a new independent observation can also produce misleading standard errors.
Free tools Windows power users keep installed
One-click scans. No signup required.
Frequency response: the same operation in another domain
For the causal length-m average, the frequency response is
Rank #4
Hm(ω) = (1/m)Σj=0m−1e−ijω = e−i(m−1)ω/2 sin(mω/2)/(m sin(ω/2)).
After r passes it is Hm,r(ω) = Hm(ω)r. The filter retains low-frequency variation and attenuates higher-frequency variation; iteration strengthens that low-pass effect. Frequencies where a single-pass response is zero remain zero after further passes. The exponential phase term encodes delay for a causal filter.
This is the Fourier-domain counterpart of the probability interpretation: convolution multiplies frequency responses just as it multiplies characteristic functions for sums of independent variables. More smoothing can remove meaningful short-lived events along with noise, flatten peaks, and blur transitions.
Computing the exact weights
Repeated convolution is usually the clearest way to generate a kernel. For example, this Python function uses NumPy:
import numpy as np
def iterated_moving_average_weights(window, passes):
if window < 1 or passes < 1:
raise ValueError("window and passes must be positive integers")
weights = np.ones(window, dtype=float) / window
for _ in range(passes - 1):
weights = np.convolve(weights, np.ones(window) / window)
return weights
print(iterated_moving_average_weights(3, 3))
The result is approximately [1, 3, 6, 7, 6, 3, 1] / 27. For a two-point average, the coefficients after r passes are exactly binomial:
wr,2(j) = 2−r binom(r,j), for j = 0, …, r.
For large windows or many passes, repeated array convolution may be inefficient; polynomial multiplication or FFT-based convolution can be preferable. A direct coefficient formula, useful for exact derivations, is
wr,m(j) = (1/mr) Σk=0⌊j/m⌋ (−1)k binom(r,k) binom(j−mk+r−1,r−1),
where invalid binomial-coefficient terms are treated as zero.
Best Value
Indexing, boundaries, and related meanings
The formulas above use a causal lag convention with indices 0 through m−1. A centered odd-length moving average can be symmetric around the target and have zero phase delay in offline use. An even-length window has its center between two sample positions, so alignment requires an explicit convention. Convolution and correlation also differ in how they orient a kernel; state the operation and index direction when implementing or comparing results.
For a finite data record, the interior formula does not determine what happens at its ends. Software must drop incomplete windows, pad with zeros, repeat edge values, reflect, wrap, or renormalize the available weights. Each choice changes boundary values, so specify it rather than assuming the full kernel applies unchanged at the first and last observations.
“Moving average” can also refer to different things. A smoother is an operation applied to observed data; a finite impulse-response filter is its signal-processing description. An MA(q) process is a stochastic model represented as a finite linear combination of white-noise terms. These are related linear-filter concepts but not interchangeable meanings. See the Encyclopedia of Mathematics overview of moving-average processes for the model terminology.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Where the bell-curve analogy stops
- It does not make arbitrary data Gaussian. The CLT above describes the distribution represented by the filter weights. Whether a weighted output is approximately normal depends on the input distribution and assumptions about its dependence and tails.
- Dependence needs separate treatment. A CLT for dependent time-series observations requires additional conditions. The shape of a deterministic kernel does not establish those conditions.
- Heavy tails can change the limit. The ordinary Gaussian CLT relies on finite variance. Some infinite-variance settings have non-Gaussian stable limits instead.
- Small pass counts are not a Gaussian limit. The exact kernel may remain visibly flat, triangular, or otherwise far from a normal curve.
- Unequal-weight sums need conditions too. In triangular-array settings, a useful sufficient negligibility condition is that no one term dominate, for example
maxj|wn,j| / √(Σjwn,j²) → 0, alongside suitable assumptions on the inputs. - Smoothing can mislead about time structure. It may lag turning points, blur breaks, flatten local peaks, and make a series look more stable while inducing serial correlation.
For further context on conditions for normal convergence, see the Encyclopedia of Mathematics entry on the Central Limit Theorem.
Choosing a smoother
Increasing the window width m broadens each averaging operation and changes the frequency-response zeros. Increasing the pass count r makes the composite kernel more bell-shaped, broadens its support, increases causal delay, and produces diminishing noise-reduction returns. Neither choice is universally best. Consider the sampling interval, expected duration of real events, tolerable delay, noise spectrum, and whether the goal is visualization, denoising, feature extraction, or forecasting.
Other methods may suit particular goals better: a Gaussian filter directly specifies a Gaussian-shaped kernel; a Savitzky–Golay filter can preserve local polynomial shape; exponential smoothing uses more recent observations more heavily; a median filter resists isolated impulse noise; LOESS offers local regression; and state-space or Kalman methods encode an explicit model. Iterated moving averages are a transparent choice when their simple, symmetric, finite kernel matches the task—not because the CLT makes them optimal.
The key chain is: one moving average is a uniform convolution kernel; repeated passes convolve that kernel with itself; its coefficients are probabilities for a sum of discrete uniforms; and the standardized kernel approaches a normal shape by the CLT. That explains the bell curve without confusing a filter’s weights with the statistical behavior of the data.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

