What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Regularization accepts a controlled amount of bias to make a model less sensitive to quirks in its training data. Ordinary least squares can fit a sample very closely, but when the data are noisy or predictors overlap, its fitted coefficients may change sharply from one sample to another. Ridge and lasso restrict coefficient sizes in different ways: ridge usually shrinks coefficients without removing predictors, while lasso can shrink some coefficients exactly to zero.

Start with the prediction problem

Suppose an outcome is generated by Y = f(X) + ε: there is an underlying relationship, plus noise that cannot be predicted from the available features. A regression algorithm sees one finite training sample and uses it to estimate that relationship. The practical goal is not merely to fit that sample; it is to predict well for new observations drawn from the population.

Ordinary least squares (OLS) chooses coefficients to minimize the sum of squared residuals:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

β̂OLS = arg minβ ||y − Xβ||²2

That objective rewards a close fit to the training data. It does not directly discourage very large coefficients, sensitivity to small data changes, or reliance on unstable combinations of redundant predictors. As a result, low training error does not guarantee low test error.

#1 Best Overall

Imagine fitting a flexible curve through noisy points. It can follow both the underlying pattern and accidental bumps in the sample. A less flexible fit may miss some real detail, but it may also ignore the bumps that will not recur in new data. Regularization formalizes that restraint by limiting the coefficient values the model can choose.

Bias and variance describe repeated samples

Bias and variance are properties of a learning procedure across hypothetical training sets, not just labels for one fitted model. Imagine repeatedly drawing a new sample from the same population and fitting the same procedure each time.

  • Bias is systematic error: the procedure’s average prediction differs from the true relationship. A straight-line model used for a strongly curved relationship, for example, may consistently miss the curve.
  • Variance is instability: predictions differ across the models fitted to those different samples. A high-variance procedure may fit one sample well but change substantially when a few observations change.
  • Irreducible noise is random variation in the outcome that even a perfect estimate of the relationship cannot remove.

For squared-error prediction at a fixed input x, the expected test error decomposes as:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

E[(Y − f̂(x))²] = (E[f̂(x)] − f(x))² + E[(f̂(x) − E[f̂(x)])²] + Var(ε)

The terms are squared bias, variance, and irreducible noise. The expectations over f̂ refer to repeated training sets from the same population. This exact decomposition applies to the specified squared-loss setting; it is not a promise that every real-world validation curve will have a smooth U shape. A useful illustration of the decomposition is available in scikit-learn’s bias–variance example.

Why a biased estimate can predict better

Suppose two predictors contain nearly the same information. OLS may assign one a large positive coefficient and the other a large negative coefficient. Those values can partially cancel in the training sample and produce a good fit, yet small changes to the data may cause the two estimates to swing dramatically. The predictions—and especially the individual coefficients—can then be unstable.

Regularization pulls estimates toward less extreme values. That makes them biased relative to OLS, but it can reduce their variance enough to lower total test mean-squared error. A small, consistent miss can be better for prediction than an estimate that is correct on average across samples but highly erratic from one sample to the next.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
Statistics Laminate Reference Chart: Parameters, Variables, Intervals, Proportions (Quickstudy: Academic )
  • This guide is a perfect overview for the topics covered in introductory statistics courses.

This is a trade-off, not a guarantee of improvement. A little regularization may help when variance is high; too much can introduce enough bias to underfit. The best balance depends on the data and should be evaluated on held-out data.

Regularization as a restriction

Ridge and lasso are often written in penalized form: minimize the training error plus a penalty for coefficient size. They can also be understood as minimizing the training error subject to a limit on coefficient size:

  • Penalized form: minimize RSS + λP(β).
  • Constrained form: minimize RSS, subject to P(β) ≤ t.

These are two views of the same trade-off: stronger regularization means a tighter restriction. Under standard convexity conditions, changing the penalty strength traces corresponding solutions to the constrained problem. The precise mapping between λ and t depends on the data and objective scaling; it is not generally as simple as t = 1/λ.

Ridge: keep the predictors, tame the coefficients

Ridge regression minimizes squared error plus a squared-coefficient penalty:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

β̂ridge = arg minβ [||y − Xβ||²2 + λ Σj βj²]

The penalty is the L₂ norm squared. Large coefficients are expensive, so ridge asks whether nearly the same fit can be achieved with more moderate values. In the standard unconstrained formulation, ridge generally shrinks coefficients continuously toward zero rather than setting them exactly to zero. It therefore tends to retain all predictors, which can be useful when many features each contribute a little signal.

Geometry: with two coefficients, the constrained ridge region is a circle: Σβj² ≤ t. OLS residual contours are ellipses, and the ridge solution is where the smallest fitting ellipse first touches the allowed region. A circle has no corners aligned with the axes, so the touching point is not usually on an axis; neither coefficient is typically exactly zero.

Why ridge helps with collinearity: when predictors are highly correlated or nearly redundant, some coefficient combinations are poorly determined by the data. Ridge discourages extreme combinations and improves numerical conditioning. More precisely, if X = UDVᵀ is a singular-value decomposition, the shrinkage factor in direction k is dk²/(dk² + λ). Directions with small singular values—the directions the data determine least reliably—are shrunk more strongly than well-supported directions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is also a Bayesian interpretation: ridge corresponds to a maximum-a-posteriori estimate under a zero-centered Gaussian prior on coefficients, with the relationship between prior variance, noise variance, and penalty strength depending on the formulation. This is one way to describe the model’s preference for moderate coefficients, not proof that the coefficients are truly near zero.

Lasso: shrink coefficients and allow exact zeros

Lasso minimizes squared error plus an absolute-value penalty:

β̂lasso = arg minβ [||y − Xβ||²2 + λ Σj|βj|]

The penalty is the L₁ norm. Lasso both shrinks coefficients and can set them exactly to zero, yielding a sparse model. This combination of shrinkage and variable selection was the motivation of the original lasso method (Tibshirani’s paper).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Geometry: the constrained lasso region in two dimensions is a diamond, Σ|βj| ≤ t. Its corners lie on the coordinate axes. An ellipse of equal fit often first touches the diamond at one of those corners, where a coefficient is zero. This explains why exact zeros are common, though the picture describes a tendency rather than a guarantee for every correlated or degenerate dataset.

Threshold intuition: with orthogonal predictors, the lasso estimate has a soft-thresholding form:

β̂j = sign(zj)(|zj| − λ)+

Here zj is the corresponding unregularized estimate and (a)+ = max(a, 0). If its magnitude is below the threshold, the coefficient becomes zero; if it exceeds the threshold, it is reduced toward zero. This cleanly shows how lasso combines selection with shrinkage.

A zero lasso coefficient means that, under the chosen data, preprocessing, loss, and penalty, the fitted model assigns that feature no contribution. It does not prove the feature is causally irrelevant or unrelated to the outcome.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why the two penalties behave differently

Geometry is one explanation; the penalties’ derivatives give another. Away from zero, the derivative of the ridge penalty is proportional to βj, so its pull toward zero weakens as the coefficient becomes small. The lasso penalty’s derivative is proportional to sign(βj), so its pull has roughly constant magnitude on either side of zero. At zero, the absolute-value penalty has a sharp, nondifferentiable point. This helps small lasso coefficients reach and remain at zero.

Correlation is a practical difference. Ridge often spreads weight across correlated predictors, which can make the overall prediction more stable. Lasso may keep one predictor and discard a highly correlated substitute; which one it retains can change across samples. The elastic-net method was developed in part to combine sparsity with more useful behavior for correlated predictors (Zou and Hastie’s paper).

Question Ridge Lasso
Penalty L₂: squared coefficient magnitudes L₁: absolute coefficient magnitudes
What happens to coefficients? Usually all shrink continuously Shrink; some can become exactly zero
Correlated predictors Often shares weight and stabilizes estimates May select one and suppress substitutes
Typical use Stable prediction when many features may contribute Sparse prediction or feature screening when sparsity is plausible
Important caution Usually does not give a compact feature list Selected features can be unstable; coefficients are biased

Elastic net: a compromise for sparsity and correlation

Elastic net combines both penalties, for example as RSS + λ₁||β||₁ + λ₂||β||₂². The L₁ component encourages exact zeros; the L₂ component discourages extreme coefficients and can help correlated predictors behave more coherently. It is a sensible candidate when a compact model is useful but predictors arrive in correlated groups. It still requires validation and does not guarantee that selected features are scientifically or causally correct.

What stronger regularization changes

In general, increasing the penalty shrinks coefficients more. With λ = 0, the penalized fit reduces to OLS; with very large regularization, slopes approach zero, leaving an intercept-only model when one is included and unpenalized. The usual pattern is less coefficient variance and more bias, but the actual test-error curve can be noisy, flat, or otherwise irregular.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Penalty strength Training fit Usual effect Risk to watch
Zero Best or near-best training fit Little shrinkage High variance in unstable settings
Small Slightly worse Some variance reduction May help, but not on every dataset
Moderate Worse than OLS Smaller coefficients; lasso may be sparse Could omit useful signal
Very large Poor Slopes approach zero High bias and underfitting
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to choose the regularization strength

The penalty strength is a hyperparameter: it should normally be selected using validation data or cross-validation, not by choosing whichever value gives the lowest training error. A practical workflow is:

  1. Reserve a final test set. Do not use it to choose the penalty.
  2. Build preprocessing into the training workflow. Standardize predictors using means and scales calculated only from the relevant training fold.
  3. Try a grid of candidate strengths for ridge, lasso, or both, using the same preprocessing and validation procedure.
  4. Select using cross-validation on the training portion, with a metric appropriate to the prediction task.
  5. Refit on the full training portion at the selected strength, then evaluate once on the untouched test set.
  6. Check selection stability across resamples if the identity of lasso-selected features matters.

For a simpler model with performance close to the cross-validation minimum, one option is the one-standard-error rule: choose the strongest penalty whose score is within one standard error of the best observed score. It is a heuristic, not a theorem.

In scikit-learn, the parameter is commonly called alpha, and estimators such as RidgeCV and LassoCV support cross-validation-based selection. The linear-model guide documents the relevant estimators and conventions. Do not transfer an alpha or lambda value blindly between libraries: objectives can differ by factors such as n or 1/2, and software may use different parameter names or scalings.

Standardization is part of the model choice

Penalties act on coefficient magnitudes, so feature units matter. If one predictor is measured in dollars and another in thousands of dollars, the same restriction on their coefficients does not correspond to the same restriction on their contributions. Standardizing predictors puts them on a more comparable scale:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

xijscaled = (xij − μj)/sj

Estimate μj and sj from training data only, including within each cross-validation fold, to prevent leakage. Most standard implementations handle the intercept separately, but check the behavior of your chosen library. If coefficient magnitudes are to be compared, use a common scale and remember that penalization itself biases estimates toward zero.

Which method should you try?

  • Try ridge first when prediction is the priority, many predictors may each contain signal, predictors are correlated, or discarding features seems risky. It is often a stable choice when the number of features is large relative to the number of observations.
  • Try lasso when a genuinely sparse signal is plausible and a smaller active feature set is operationally useful. Treat it as a selection method, not as proof that the chosen variables are uniquely important.
  • Try elastic net when you want sparsity but predictors are correlated or grouped, so selecting a single arbitrary representative would be undesirable.

Compare candidates under the same data splits and preprocessing. If prediction is the only goal, judge them by held-out predictive performance. If feature selection is part of the goal, also examine whether selections persist across folds or resamples.

What regularization does not tell you

Regularization is a way to control a model’s flexibility; it does not repair a flawed evaluation design or establish a scientific explanation. A selected coefficient is not evidence of causation, and a zero coefficient does not establish irrelevance. Lasso selection can depend on sample variation, feature scaling, correlations, and the selected penalty. Penalized coefficients are deliberately shrunk, so they are not automatically unbiased effect estimates.

Nor does regularization fix data leakage. Features containing future information, post-outcome measurements, duplicate records, or preprocessing performed using validation data can still make a model’s evaluation misleading. Ridge and lasso also do not guarantee improved accuracy: the right penalty depends on the problem, and excessive shrinkage can underfit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A compact mental model

OLS trusts the observed sample to choose the coefficients that fit it best. Ridge says, “use the available predictors, but avoid extreme coefficients.” Lasso says, “shrink coefficients and allow some to be removed entirely.” Elastic net combines those preferences. All three manage the same underlying tension: fitting the current sample closely versus building predictions that remain stable on new data.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.