Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Data can be accurate and still give a false impression. The distortion may begin with who was measured, continue through the choice of denominator or average, and end with a chart or headline that suggests more than the evidence supports. To assess a numerical claim, reconstruct what was measured, for whom, over what period, and how the result was turned into a conclusion.
What it means to lie with data
Data do not have intent. People can fabricate observations, falsify or suppress results, or present genuine numbers in a way that encourages an unsupported conclusion. The last case is not automatically fraud: misleading presentation can come from poor methods, software defaults, carelessness, or pressure to simplify a complicated result.
Darrell Huff’s How to Lie with Statistics, first published in 1954, made readers alert to biased samples, selective averages, distorted graphs, and causal overreach. The same concerns remain relevant, alongside modern issues such as missing data, model selection, and machine-learning evaluation. The National Academies’ Reference Manual on Scientific Evidence, Fourth Edition and its statistical reference material discuss how summaries, rates, graphs, and variability affect interpretation.
A useful way to inspect a claim is to follow the data’s path: collection, definition, inclusion and exclusion, summary, comparison, visualization, and interpretation. A problem at any point can change what readers take away.
#1 Best Overall
- Wiley
- Language: english
- Book - storytelling with data: a data visualization guide for business professionals
Start with the sample and the definition
Before asking whether a calculation is right, ask whether the observations can answer the question being asked. A poll of a company’s current customers does not automatically describe all consumers. A voluntary survey may attract people with unusually strong views, and people who do not respond may differ from those who do. A large sample does not fix a systematic selection problem.
- Convenience and self-selected samples: Participants are easy to reach or choose to take part, so they may not represent the wider population.
- Nonresponse and survivorship: People who do not answer, leave a study, stop using a service, or fail may disappear from the final results.
- Small or clustered samples: A small group is more vulnerable to chance extremes. Several observations from one household or organization may not be independent evidence from several unrelated cases.
- Population mismatch: A result about respondents, registered users, or eligible patients may be described as if it covered everyone.
Then inspect the definition. Does an event mean confirmed, suspected, reported, or estimated? Is the unit a person, transaction, device, visit, or account? Are repeat events counted? Where does the measurement period begin and end? Two apparently conflicting statistics may both be correct if they use different definitions, populations, denominators, or periods.
The National Academies’ statistics and research-methods reference explains how sample design shapes inference from a study group to a broader population. The key question is not simply how many observations there are, but who was eligible, who was included, and who the conclusion is meant to describe.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Check the denominator and the baseline
A percentage is a numerator divided by a denominator. Without both, it is hard to know what the number means. A service receiving 100 complaints has a 10% complaint rate among 1,000 customers, but a 0.1% rate among 100,000 customers. The total complaints are unchanged; the rate answers a different question.
- What is the numerator, and what exactly counts?
- What is the denominator, and who or what is eligible to be counted?
- Do numerator and denominator cover the same population and period?
- Were inactive users, incomplete surveys, or other groups excluded?
- Would another reasonable denominator change the conclusion?
Rates and totals can tell different legitimate stories. A city with more residents may have more incidents but a lower per-person rate. A small subgroup may have a high rate while contributing few events overall. A rate can fall even as the number of events rises if the population grows faster.
Always recover the baseline when a claim gives a percentage change. If a rate rises from 1% to 2%, that is a 1-percentage-point increase and a 100% relative increase. Both are mathematically correct; neither is interpretable without the underlying rate and its denominator. The same discipline applies to claims such as “doubled”: success rising from 1 in 1,000 to 2 in 1,000 is a doubling, but the absolute change is one additional success per 1,000.
Ask what an average hides
“Average” is ambiguous. The mean adds values and divides by their count; the median is the middle value after sorting; the mode is the most frequent value. A weighted average gives some observations more influence, while a trimmed or adjusted average changes or removes certain values. None is inherently the honest one: the right summary depends on the distribution and the question.
Free tools Windows power users keep installed
One-click scans. No signup required.
Consider five salaries: $30,000, $32,000, $35,000, $38,000, and $500,000. The mean is $127,000, while the median is $35,000. Both are calculated correctly, but the mean is pulled upward by the highest salary. For skewed income data, reporting the median, distribution, and relevant percentiles alongside the mean can make the picture clearer.
An overall average can also conceal variation among groups. Retention may be high for one customer cohort and low for another; an institution’s aggregate result may reflect different case mixes. In Simpson’s paradox, an aggregate trend can reverse when the data are separated into relevant groups, often because group sizes or compositions differ. Examine subgroup results and group sizes where appropriate, and explain any adjustment rather than treating it as a neutral correction.
Look for selective comparisons and missing cases
Cherry-picking occurs when a presentation highlights favorable dates, geographies, demographics, metrics, product versions, survey questions, or successful examples while leaving out relevant alternatives. Starting a time series at an unusually favorable point can make a later rise appear stronger. Reporting one outcome from many tested outcomes can make chance look like a discovery.
Selection is not automatically improper. A focused analysis can use a limited group or period if the inclusion rule is clear and relevant. Concern rises when the rule changes after results are seen, exclusions are unexplained, or only supportive evidence is shown. Look for the full relevant time series, stated criteria, all material outcomes, and contrary findings.
Missing data deserve the same scrutiny. If a study starts with 1,000 participants but reports outcomes for only 600 who completed follow-up, ask what happened to the other 400 and whether their outcomes may differ. The same issue appears when only customers who completed a satisfaction survey, successful users who stayed, or records with complete fields remain in an analysis.
Trying many subgroups, outcomes, time windows, correlations, or models and reporting only the favorable result is another form of selection. Pre-registration, holdout data, replication, and appropriate adjustments for multiple comparisons can help distinguish a predicted test from an exploratory finding. A result discovered while exploring should be treated as a lead to check, not as though it had been specified in advance.
Read the chart’s visual encoding
Charts change how values look, even when the underlying numbers do not. A review of visual communication describes how scales, color, and graphical conventions affect interpretation; see the review of visual data communication.
- Truncated bar axes: Bars are read by length, so starting the axis well above zero can exaggerate differences. For example, values of 100 and 110 look modestly different on a zero baseline but dramatically different if the axis starts at 95.
- Unequal scales and aspect ratios: A line’s apparent steepness can change when the axis range or chart shape changes. Check labeled limits and units.
- Area, volume, and perspective: Circles, icons, and 3D columns can magnify differences if viewers compare area or volume rather than a clearly scaled length.
- Dual axes: Two independently scaled y-axes can make unrelated series appear to move together.
- Unequal time spacing or omitted observations: Equal gaps between irregular dates, missing zero values, or omitted earlier periods can distort a trend.
- Color scales: A strong gradient applied to a narrow range can make small numerical differences appear categorical or alarming. Read the legend and the actual values.
- Cumulative totals and smoothing: Cumulative counts commonly rise as observations are added; that alone does not show acceleration. Moving averages and trend lines can mute variability, so the window or fitting method matters.
A nonzero axis is not automatically deceptive. A line chart may use one to show small fluctuations; logarithmic scales, broken axes, or relative measures can also be appropriate for a clear purpose. The scale, transformation, and limits should be labeled, and raw values or a complementary view should be available when needed. The point is to judge whether the visual encoding supports the stated comparison, not to apply a slogan mechanically.
Separate association, prediction, and cause
Two variables moving together establishes an association, not by itself a causal link. Ice-cream sales and drownings can both rise in hot weather; temperature and seasonal activity are plausible factors behind the pattern. Other possibilities include a third variable influencing both, causality running in the opposite direction, or selection and measurement choices creating an apparent relationship.
Best Value
Ask whether the claim is descriptive (what was observed), predictive (what a model expects), or causal (what would change if an intervention occurred). A causal conclusion needs a credible design for ruling out alternative explanations, such as randomization or another appropriate identification strategy. Correlation can still be useful evidence; it simply cannot carry the causal claim alone.
The same distinction matters for predictive and machine-learning models. Training performance is not evidence of equivalent performance on new cases. Overall accuracy can conceal poor results for a minority class or subgroup; precision, recall, false-positive rates, and false-negative rates answer different questions. A forecast is not an observation, feature importance does not establish cause, and a sophisticated model can reproduce historical bias.
Make uncertainty part of the claim
A number with many decimal places is not necessarily precise. Sampling variation, measurement error, model assumptions, missing records, and forecast uncertainty all affect what can reasonably be concluded. Confidence intervals and margins of error describe aspects of uncertainty under specified methods; a confidence interval is not simply a guarantee that a fixed parameter has a stated probability of lying within this particular calculated interval.
Recommended Free Tools
Statistical significance is not the same as practical importance. A large sample can make a tiny difference statistically detectable. Conversely, a result that is not statistically significant does not prove there is no effect; the data may be too imprecise to distinguish one from zero. Even a narrow interval cannot repair biased collection or an unsuitable sample.
Regression to the mean is another reason to be cautious about dramatic before-and-after stories: unusually high or low observations tend, in part, to be followed by values closer to the average even without an intervention. Compare against an appropriate control or longer record where possible, and distinguish what the data establish from what remains uncertain.
A seven-part audit for any numerical claim
- Source: Who collected the data, when, and for what purpose? Is the original source or dataset available, and does the source have an interest in the conclusion?
- Scope: What population, geography, unit, and period are covered? Does the headline generalize beyond them?
- Base: What are the numerator, denominator, and baseline? Is the change absolute, relative, or both?
- Shape: Is the chart type suitable? Are its axes, intervals, colors, and scale transformations clear?
- Spread: What variability and uncertainty exist? Are outliers, subgroups, missing data, and attrition visible?
- Cause: Is the statement descriptive, predictive, or causal? What competing explanations or study-design limits matter?
- Counterevidence: Which periods, outcomes, groups, or studies are absent? Would a reasonable alternative definition or analysis change the takeaway?
For reproducibility, a trustworthy analysis should make it possible to find the source, collection date, field definitions, inclusion rules, cleaning and transformation steps, calculation method, and dataset version. Code or supplementary tables can help others verify the work. A chart lacking this context is difficult to audit, though lack of documentation alone does not prove it is false.
How to present data more honestly
- Pair relative percentages with the baseline and absolute change.
- State numerator, denominator, population, unit, and time period for rates.
- Show a full relevant time series and explain any selected window or exclusions.
- Choose a summary that matches the distribution; include median, spread, or subgroup results when the mean alone would hide important structure.
- Label scales and transformations, provide values, and avoid decorative area or volume that implies unsupported magnitude.
- Show uncertainty and distinguish exploratory findings from confirmatory tests.
- Use association language unless the design supports a causal claim; describe predictive limits and subgroup performance for models.
- Document provenance and transformations so another reader can trace the result.
Good data literacy is neither automatic trust nor automatic suspicion. It means checking whether the conclusion follows from the observations, definitions, comparisons, and uncertainty actually shown.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

