Recommended Free Tools
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
There is no single best replacement for R-squared. Use adjusted R² to summarize comparable ordinary least-squares models while accounting for predictor count; cross-validated MAE or RMSE to compare predictions; AIC, AICc, or BIC to compare compatible likelihood-based models; and a specifically named pseudo-R² for many generalized linear models. For consequential decisions, pair a score with a meaningful baseline, validation that reflects deployment, and checks of errors in real units.
What R-squared measures—and what it does not
For ordinary regression, R² is commonly defined as:
R² = 1 − SSE / SST = 1 − [Σ(yᵢ − ŷᵢ)² / Σ(yᵢ − ȳ)²]
SSE is the sum of squared residuals: the differences between observed values and model predictions. SST is the sum of squared differences between observed values and their sample mean. In ordinary least-squares regression with an intercept, R² is commonly between 0 and 1. On held-out data—or for models without an intercept—it can be negative.
#1 Best Overall
R² compares the model’s squared errors with a baseline that predicts the sample mean. In-sample, it is often described as the proportion of variation accounted for by the fitted model. It is unitless, but it is still tied to squared-error loss, so large misses count disproportionately. Most importantly, an R² calculated on training data describes in-sample fit, not necessarily performance on new cases.
“Explains 80% of the variance” does not mean predictions are within 20% of the correct value, that the model is causally valid, or that its errors are acceptable in practice. A high R² can coexist with leakage, poor calibration, or bad performance on new data. A low R² does not automatically make a model useless: in noisy settings, predictions can still be useful in the original units. R² is not inherently useless, either; it is informative when the question is about in-sample variation and its assumptions and limits are understood. A discussion of how R² compares with other measures appears in this [methodological paper](https://pmc.ncbi.nlm.nih.gov/articles/PMC8279135/).
Quick comparison
| Measure | Main question | Best suited to | Better direction | Main limitation |
|---|---|---|---|---|
| Adjusted R² | How does in-sample fit look after accounting for predictor count? | Comparable ordinary least-squares models | Higher | Still in-sample; does not validate predictions |
| Test or cross-validated R² | How does squared-error performance compare with a stated baseline on new data? | Regression prediction | Higher | Depends on baseline and validation design; can be negative |
| MAE | How large is the typical absolute miss? | Interpretable error in target units | Lower | Can underemphasize rare, very large errors |
| RMSE | How large are squared-error-weighted misses? | Tasks where large errors deserve extra penalty | Lower | Sensitive to outliers; measured in target units |
| AIC, AICc, BIC | Which compatible likelihood-based candidate balances fit and complexity? | Model selection within a coherent likelihood framework | Lower | Not an error in target units or a guarantee of predictive accuracy |
| Pseudo-R² | How does a specified likelihood-based model compare with a reference? | Some logistic and other generalized models | Usually higher | Definitions differ; not ordinary R² |
| MASE or other scaled loss | How does forecast error compare with an appropriate benchmark? | Time-series forecasting | Lower | Meaning depends on the chosen benchmark |
Adjusted R-squared: a complexity-aware summary, not a cure for overfitting
A common adjusted R² formula is:
Adjusted R² = 1 − (1 − R²) × (n − 1) / (n − p − 1)
Here, n is the number of observations and p is the number of predictors. Unlike ordinary R², adjusted R² applies a penalty for adding predictors. It can fall when a new predictor contributes too little relative to the sample size and model complexity.
Advantages: It preserves a familiar fit summary, accounts for the number of predictors, and can help compare ordinary least-squares models fitted to the same response and observations. Limitations: It remains an in-sample statistic. It does not directly measure new-data accuracy, and its penalty does not prevent overfitting. Comparisons are questionable when models use different rows, outcome transformations, weights, or definitions. It is not a universal score for generalized linear, mixed, nonlinear, or machine-learning models.
Use adjusted R² as a complexity-aware descriptive measure for comparable linear models, not as a substitute for validation. For a discussion of R² variants and adjusted R², see this [statistical reference](https://hbiostat.org/bib/r2).
For prediction, compare MAE and RMSE on held-out data
When the goal is prediction, calculate errors on data that were not used to fit the model, or estimate them through a suitable cross-validation design. Two widely used choices are:
MAE = (1/n) Σ|yᵢ − ŷᵢ|RMSE = √[(1/n) Σ(yᵢ − ŷᵢ)²]
Both are expressed in the target’s units. If the target is dollars, the error is in dollars; if it is hours, the error is in hours. Neither is scale-free, so values for targets with different units or very different ranges should not be compared directly.
MAE: a clear measure of absolute error
MAE averages the absolute difference between predictions and observations. It is often easy to explain as the average size of a prediction miss and is less sensitive to extreme errors than RMSE. It can be a good primary measure when each unit of error has roughly equal cost.
MAE does not show whether the model tends to overpredict or underpredict; pair it with a bias measure such as mean error. It also gives less additional weight to very large misses. If a rare but severe error matters greatly, MAE alone can hide that risk. MAE and RMSE answer different questions, so neither is universally better.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchRMSE: extra weight for large misses
RMSE takes the square root of the average squared error. Because errors are squared before averaging, a few large misses can strongly affect it. Choose it when large errors deserve substantially more penalty or when squared loss matches the operational objective.
Rank #3
That sensitivity is also a drawback: an outlier can dominate the score. RMSE does not show which cases failed, whether error is biased, or whether a particular group receives much worse predictions. Report it alongside MAE when both typical performance and the cost of large misses matter. R’s documentation lists common absolute- and squared-error cost functions, including RMSE-related measures, in its [cross-validation reference](https://search.r-project.org/CRAN/refmans/cv/html/cost-functions.html).
Percentage and scaled errors: handle denominators carefully
Percentage and scaled measures can be useful when relative performance matters, but their denominators determine what gets emphasized.
- MAPE averages the absolute error divided by the actual value, usually expressed as a percentage. It is undefined when an actual value is zero and can become extreme near zero. It also gives small actual values disproportionate influence, so it is not a neutral, universally comparable score.
- sMAPE is intended to moderate some MAPE problems, but multiple formulas are in use and denominator issues remain. State the exact formula and implementation rather than treating the name as unambiguous.
- WAPE divides total absolute error by the total absolute actual value. It can suit aggregate reporting, but may conceal poor performance in low-volume segments and becomes unstable when the denominator is small.
- MASE scales forecast error against a naive in-sample benchmark. It is often useful for comparing forecasts across series, provided the benchmark is appropriate. A poor benchmark can make the resulting score misleading.
For zero-heavy, intermittent, or negative-valued data, consider MAE, RMSE, or a loss tied to the real operational cost instead. When reporting a percentage or scaled metric, explain how zeros, missing values, negative values, and the benchmark were handled. Research describes how MAPE’s denominator changes the effective weighting of errors; see this [analysis of percentage-error measures](https://arxiv.org/abs/1605.02541).
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →AIC, AICc, and BIC: for likelihood-based model selection
AIC and BIC combine a measure of likelihood fit with a penalty for model complexity. Common forms are:
AIC = 2k − 2 log LBIC = k log(n) − 2 log L
Here, L is the maximized likelihood, k is the number of estimated parameters, and n is the sample size. AICc is a small-sample correction to AIC and is particularly worth considering when the sample is small relative to the number of parameters. For all three, lower is preferred among the candidate models being compared.
These criteria are useful when the candidates are fit under compatible likelihood assumptions and to the same observations and response definition. AIC and BIC can choose different candidates because they penalize complexity differently; BIC commonly applies the stronger penalty. Neither score has a direct interpretation as “percentage explained” or typical prediction error, and a lower value does not establish acceptable real-world accuracy.
Rank #4
Do not compare values casually across different datasets, response scales, missing-row handling, or incompatible likelihood conventions. Software can differ in conventions for likelihood constants, parameter counts, weights, and variance parameters, so keep comparisons within a compatible model set. Use AIC, AICc, or BIC for relative model selection—not to communicate error in business units. SAS provides a [reference to common regression and information-criterion formulas](https://support.sas.com/documentation/cdl/en/etsug/68148/HTML/default/etsug_tffordet_sect048.htm); R’s model-performance documentation lists AIC, AICc, BIC, R², adjusted R², and error measures together [here](https://search.r-project.org/CRAN/refmans/performance/html/model_performance.lm.html).
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsLog likelihood, deviance, and pseudo-R² for generalized models
For binary, count, and other non-Gaussian outcomes, ordinary least-squares R² is generally not the default fit measure. Log likelihood measures how well a model assigns probability to the observed data under its assumed distribution. Deviance compares the fitted model with a likelihood-based reference, often a saturated model. These quantities are natural for generalized linear models and can support nested-model comparisons, likelihood-ratio tests, and information criteria. They are less intuitive than errors in original units, depend on the distribution and likelihood convention, and do not by themselves establish better business outcomes or predictive performance.
Some software reports pseudo-R² values, including McFadden, Cox–Snell, and Nagelkerke. They provide compact summaries relative to a reference model, but their definitions and scales differ. They are not generally interpretable as the percentage of variance explained, and there is no universal threshold for a “good” value. Name the exact statistic and report it alongside measures that match the task. IBM’s documentation explains several [pseudo-R² measures](https://www.ibm.com/docs/en/spss-statistics/cd?topic=model-pseudo-r-squared-measures) and treats them as distinct statistics.
For binary predictions, consider log loss or Brier score when probability quality matters, and check calibration. ROC-AUC measures ranking discrimination, not whether predicted probabilities are well calibrated or whether errors are acceptable at a chosen threshold. For imbalanced outcomes, also consider prevalence, precision-recall performance, and the costs of false positives and false negatives. Scikit-learn’s [model-evaluation guide](https://scikit-learn.org/stable/modules/model_evaluation.html) documents regression, classification, and probability-scoring measures, including Brier score.
Out-of-sample R² and validation that matches use
Held-out R² evaluates squared prediction error relative to a stated baseline. A common form is:
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Test R² = 1 − Σ(yᵢ − ŷᵢ)² / Σ(yᵢ − baselineᵢ)²
The baseline matters. It could be the training-set mean for an ordinary regression problem, a seasonal-naive forecast for time series, or the existing operational method. A negative value means the model performed worse than that baseline under squared-error loss—not necessarily that the computation is wrong. Always state the baseline. Out-of-sample R² can reveal overfitting that training R² hides, but it still uses squared errors and depends on the split. A recent paper discusses out-of-sample R² and validation-based estimation [here](https://arxiv.org/abs/2302.05131).
Choose the validation design to represent how the model will be used:
- Ordinary independent observations: Use a held-out test set or cross-validation. Repeated or nested cross-validation can help when data are limited or model tuning is extensive. Keep feature selection and preprocessing inside each training fold.
- Time series: Use chronological splits, rolling-origin evaluation, or blocked validation. Random k-fold splitting can train on the future and test on the past, leaking information that would not be available at deployment. Compare against a sensible time-series benchmark.
- Grouped observations: If rows belong to the same person, patient, store, device, or company, split by group when deployment means predicting for new groups. Random row splits can leak entity-specific information across training and validation.
Cross-validation estimates performance only to the extent its splits and preprocessing reflect deployment. Small or noisy validation sets can produce unstable rankings; report variation across folds or an uncertainty interval rather than declaring a winner from a tiny difference.
Free tools Windows power users keep installed
One-click scans. No signup required.
Why one number is not enough: diagnostics and uncertainty
A mean score can conceal systematic failure. Pair metrics with checks suited to the model and decision:
- Plot residuals against fitted values to look for nonlinear patterns or changing error spread.
- Check residual autocorrelation when observations are ordered in time. Use Q–Q plots when distributional assumptions matter.
- Inspect leverage, influential observations, and outliers. Determine whether extreme cases are data problems, rare but important events, or signs of model misspecification.
- Report bias and errors by target magnitude, time period, and important subgroups. A favorable overall average can hide poor performance for a group or high-value cases.
- For probabilistic predictions, assess calibration. For prediction intervals, check coverage as well as width.
Heteroscedasticity can make a single average error misleading, especially if high-value cases have much larger errors. Consider reporting errors by magnitude or using a justified weighted loss. MAE is less sensitive to outliers than RMSE, but neither reveals why a case failed. Diagnostics can point toward a different transformation, functional form, robust method, weighting scheme, or model family.
Which measure should you use?
| Your question | Start with | Also report or check |
|---|---|---|
| How much in-sample variation is associated with predictors? | R² | Adjusted R², residual plots, and a clear statement that it is in-sample |
| Did added predictors justify their complexity? | Adjusted R² for comparable OLS models | AICc/BIC where applicable and validation performance |
| Which model predicts new observations best? | Cross-validated or test-set MAE or RMSE | Baseline, validation uncertainty, and a secondary error measure |
| Are large mistakes particularly costly? | RMSE or a domain-specific squared/weighted loss | MAE and tail-error analysis |
| What is the typical error in practical units? | MAE | Bias and subgroup error |
| Does relative forecast error matter? | MASE, WAPE, or carefully qualified MAPE | Benchmark choice, zero policy, and absolute error |
| Which likelihood-based candidates should be preferred? | AIC, AICc, or BIC | Comparable likelihoods, same observations, and validation |
| Is the outcome binary or otherwise non-Gaussian? | Log loss, deviance, or an appropriate proper score | Calibration, discrimination, decision costs, and a named pseudo-R² if useful |
| Are predictions for a time series? | Rolling or blocked validation with MAE, RMSE, or MASE | Seasonal-naive baseline and interval coverage |
| Do different models use different targets, transformations, or rows? | Do not compare headline scores directly | Evaluate predictions on a common outcome scale and comparable observations |
A practical reporting template
For a model comparison, report:
- Baseline: for example, a training-mean predictor, seasonal-naive forecast, or current operational method.
- Validation design: holdout split, cross-validation folds, grouped split, or rolling-origin dates. Explain how preprocessing and feature selection were kept within training data.
- Primary loss: choose the measure that reflects the cost of errors, such as MAE, RMSE, or a domain-specific weighted loss.
- Secondary measure: add a complementary measure, such as the other of MAE/RMSE, or a clearly defined R² or information criterion when relevant.
- Uncertainty and diagnostics: give fold-to-fold variation or an interval, plus bias, subgroup errors, and relevant residual or calibration checks.
For example: “We compared the candidate models with five grouped validation folds, using MAE as the primary metric because errors have approximately equal per-unit cost. We report RMSE to show sensitivity to large misses, compare both against the existing baseline, and summarize fold variability and subgroup errors.” This describes a defensible comparison without pretending one score answers every question.
Do not declare a model best solely because it has the highest training R², the lowest training RMSE, or the lowest AIC. Make sure candidates were evaluated on comparable observations and target scales. A model fitted to log-transformed outcomes, for instance, should be evaluated on the common original outcome scale if that is where decisions are made, with retransformation effects considered. Predefine the primary metric where possible; selecting and reporting only the measure that favors a preferred model can distort the comparison.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Bottom line
Keep R² when its narrow job—summarizing in-sample variation—is useful. For prediction, prioritize validated error in meaningful units; for complexity-aware linear summaries, consider adjusted R²; for likelihood-based model selection, use compatible AICc or BIC; and for generalized models, name the exact pseudo-R² rather than equating it with ordinary R². The strongest assessment is usually a bundle: a relevant baseline, deployment-matched validation, a primary loss aligned with the real decision, uncertainty, and diagnostics.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

