Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
If an iterative calculation converges by bouncing above and below its answer, average each estimate with the one immediately before it. This inexpensive post-processing step can cancel part of a shrinking, alternating error. It is not a universal speedup: test it against the original sequence, because smoother estimates are not necessarily more accurate ones.
The trick: average neighboring estimates
Suppose an algorithm produces estimates f1, f2, … of a limit f. Replace each pair of neighbors with their arithmetic mean:
gk = (fk + fk−1)/2
This does not change the algorithm generating the estimates. It transforms the sequence it has already produced, so the underlying computation may still require the same iterations or function evaluations.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Why averaging can help
Write the signed error as Ek = f − fk. The averaged estimate’s error is:
#1 Best Overall
f − gk = (Ek + Ek−1)/2.
If successive errors have opposite signs and similar magnitudes, their sum is small. For example, a sequence of the form fk = f + (−1)kak, where positive ak decreases gradually, alternates around the limit. Neighboring estimates then tend to straddle it, and averaging cancels much of the alternating component.
The key condition is not simply that the values wiggle. The error should change sign in a reasonably regular way while shrinking. Monotone convergence, growing oscillations, or random noise do not provide the same reason to expect improvement.
Worked example: the alternating series for log 2
The partial sums of the alternating harmonic series converge to log(2):
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Rank #2
- This guide is a perfect overview for the topics covered in introductory statistics courses.
Sn = 1 − 1/2 + 1/3 − 1/4 + ··· + (−1)n+1/n.
This is a classic example of slow convergence with alternating truncation error. A first smoothed estimate is An = (Sn + Sn−1)/2. Repeating the operation gives a higher-order smoothed sequence.
The following values illustrate the difference. “Two passes” means averaging twice, using the recurrence below. The errors are absolute differences from log(2) ≈ 0.69314718056.
Rank #3
| n | Raw Sn | Raw error | One pass | One-pass error | Two passes | Two-pass error |
|---|---|---|---|---|---|---|
| 3 | 0.8333333333 | 0.1401861528 | 0.6666666667 | 0.0264805139 | 0.7083333333 | 0.0151861528 |
| 5 | 0.7833333333 | 0.0901861528 | 0.6733333333 | 0.0198138472 | 0.6991666667 | 0.0060194861 |
| 10 | 0.6456349206 | 0.0475122599 | 0.6981726190 | 0.0050254385 | 0.6949236111 | 0.0017764305 |
| 20 | 0.6687714032 | 0.0243757774 | 0.6931606444 | 0.0000134638 | 0.6932658294 | 0.0001186488 |
Here, averaging substantially reduces the error at the listed iteration counts, but the second pass is not always better than the first. That is exactly why the comparison should be measured rather than inferred from how smooth a sequence looks.
Free tools Windows power users keep installed
One-click scans. No signup required.
Repeated averaging
Let gk(0) = fk. After each pass, average neighboring values of the previous pass:
gk(m) = (gk(m−1) + gk−1(m−1))/2.
After m passes, this equals a binomially weighted average of m+1 consecutive original estimates:
Rank #4
gk(m) = 2−m Σj=0m C(m,j) fk−j.
Each additional pass broadens the averaging window and increases lag. It can cancel more of a smooth alternating component, but also discards more of the sequence’s short-term detail. There is no general rule that more passes mean faster or more accurate convergence.
Python implementation
def smooth_once(values):
return [
0.5 * (values[i] + values[i - 1])
for i in range(1, len(values))
]
def repeated_smoothing(values, passes):
result = list(values)
for _ in range(passes):
result = smooth_once(result)
return result
Each pass shortens the returned list by one value, because it needs a neighboring pair. Equivalently, an m-pass estimate requires m+1 consecutive original estimates. For a single streaming pass, retain the previous value and average it with the new one; repeated passes require stored intermediate sequences or the binomial formula.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesHow to test whether it actually improves convergence
- Keep the raw sequence. Do not overwrite the estimates produced by the original algorithm.
- Check the pattern. If you know the reference value, inspect signed errors as well as absolute errors. Otherwise, examine a meaningful residual or validation measure; visual zigzags alone are not proof of alternating error.
- Compare at equal cost. Measure raw and smoothed accuracy after the same number of underlying iterations or function evaluations. Also compare iterations to a fixed tolerance if the application has a stopping threshold.
- Use an independent measure when needed. A smoothed training curve can look better without improving held-out predictions or the actual objective.
- Choose the smallest useful number of passes. More smoothing adds lag and can hide changes or instability.
“Faster convergence” can mean lower error at a fixed iteration, fewer iterations to a target tolerance, fewer expensive evaluations, or lower wall-clock time. Averaging may help the first two while leaving the cost of generating the original sequence unchanged. Report the measure that matters for your task rather than treating a smoother plot as a speed benchmark.
Best Value
Where it may fit—and where it may not
The idea can be tried on numerical series, fixed-point iterations, root-finding sequences, or other iterative calculations when estimates exhibit decaying, sign-alternating error. It can also be applied component by component to vector-valued estimates. Whether that makes sense depends on what the components represent.
In optimization, averaging parameter vectors is not the same as averaging objective values or predictions. An averaged parameter vector may perform worse than either neighboring model, and it may violate constraints. Simple averaging is also distinct from momentum, Nesterov acceleration, iterate averaging, exponential moving averages, Richardson extrapolation, Aitken’s Δ² process, and Anderson acceleration; those methods have different definitions and assumptions. This technique alone is not a general improvement to gradient descent.
- Monotone convergence: Neighboring estimates approach from the same side, so there is no alternating error to cancel. Averaging may change the estimate but has no general advantage.
- Stochastic noise: Alternation caused by noisy minibatches or measurements may be smoothed, but that does not prove the underlying method converges faster. Check an independent validation metric or repeated runs.
- Persistent or growing oscillation: Smoothing can conceal instability rather than fix it. Inspect raw iterates, residuals, and stopping criteria too.
- Constraints and nonlinear representations: Arithmetic means can leave a feasible region or mishandle angles, positive parameters, integer states, or normalized quantities. Use a constraint-preserving representation or method where appropriate; projection changes the procedure and should be assessed separately.
- Rapidly changing or delayed output: Averaging uses older values and may lag behind the current state. It can be unsuitable when immediate response matters or when the algorithm’s guarantee depends on its latest iterate.
The central idea and its caveat—that gains depend on the particular sequence—are also described in Vincent Granville’s 2020 explanation. The useful interpretation is error cancellation for a favorable pattern, not a guaranteed dramatic speedup.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

