Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Testing many hypotheses creates more opportunities for a chance result to look significant. The right response depends on which tests form one analysis family, how costly any false positive would be, and whether the chosen correction’s assumptions fit the data. Familywise error rate (FWER) methods limit the chance of at least one false rejection; false discovery rate (FDR) methods instead limit the expected share of false findings among rejected hypotheses.
Why multiple testing increases false positives
Every statistical test has a chance of rejecting a null hypothesis that is actually true. When researchers test many hypotheses, outcomes, subgroups, or analytical choices, they create more opportunities for at least one low p-value to arise by chance. The overall chance depends on both the number of tests and how they are related; it cannot be inferred from the per-test significance level alone.
Multiplicity can arise from testing several outcomes, producing several p-values, repeatedly checking data as it accumulates, or running unplanned analyses after seeing results. A review by Streiner describes these as common forms of the multiple-testing problem (2015 review abstract).
Define the analysis family before choosing a correction
An analysis family is the set of hypotheses whose results could be used to support the same scientific claim or decision. It is not necessarily every test in a paper, nor is it automatically just the tests in one table. If readers or decision-makers could select whichever favorable result emerges from a group of outcomes or analyses, those tests may belong to the same family.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
Researchers can define separate families when the scientific questions are genuinely distinct, but should explain that rationale rather than splitting tests simply to avoid adjustment. Identify primary hypotheses in advance, distinguish them from exploratory work, and account for multiple outcomes, interim looks, and post hoc tests when deciding what claims the results can support.
FWER and FDR control different risks
| Target | What it controls | When it can fit |
|---|---|---|
| Familywise error rate (FWER) | The probability of one or more false rejections within a defined family. | When even one false positive in that family could lead to a consequential claim or decision. |
| False discovery rate (FDR) | The expected proportion of false discoveries among the hypotheses rejected. | When examining a broader set of candidates and a controlled share of false findings is an acceptable discovery trade-off. |
These are not interchangeable guarantees. FWER is concerned with whether any false positive occurs in the family; FDR concerns the expected fraction of false findings among the results declared significant. Choose the error target to match the consequences of the claim, not simply because one correction is familiar.
Bonferroni, Holm, and Benjamini–Hochberg
Bonferroni and Holm for FWER
Bonferroni is a simple FWER-oriented procedure. Holm’s step-down procedure is another FWER option, testing ordered p-values sequentially. Both aim to limit the probability of one or more false rejections in the specified family. FWER procedures can be conservative and reduce power, so their stricter protection may come at the cost of missing real effects.
Benjamini–Hochberg for FDR
The Benjamini–Hochberg (BH) procedure targets FDR rather than FWER. In their 1995 paper, Yoav Benjamini and Yosef Hochberg wrote: “A different approach to problems of multiple significance testing is presented. It calls for controlling the expected proportion of falsely rejected hypotheses — the false discovery rate.” Their paper presents FDR as an alternative when that is the desired criterion and reports the potential for greater power than FWER control (Benjamini and Hochberg, 1995).
Rank #3
The original BH result establishes FDR control for independent test statistics. Do not assume that guarantee carries over unchanged to every dependent testing setup. Dependence-aware methods and resampling approaches are available, but the appropriate guarantee depends on the method and its assumptions (Benjamini, 2010).
A practical workflow for controlling multiplicity
- Define the claim and family before examining results. List the hypotheses and outcomes that could support the same claim. Explain any decision to treat distinct questions as separate families.
- Mark confirmatory and exploratory work. Identify prespecified primary hypotheses, and plan how exploratory analyses, repeated looks, and post hoc tests will be reported.
- Choose the error target. Use FWER when the chance of any false positive in the family must be kept low; consider FDR when the relevant concern is the expected share of false discoveries among a set of reported findings.
- Match a procedure to the design. Choose a method whose assumptions fit the number and dependence of tests. State the target level and procedure, and consider whether its conservativeness and power trade-off suit the decision.
- Report the complete result set. Provide effect estimates and uncertainty alongside adjusted results, and disclose the outcomes, analyses, interim looks, and post hoc work that were performed.
What a correction cannot fix
A multiplicity correction addresses a specified error target under its assumptions. It does not repair biased measurement, poor study design, selective reporting, p-hacking, or an exaggerated interpretation of effect size. Nor does adjustment turn a result chosen after looking at the data into a prespecified confirmatory finding. Transparent reporting and a defensible analysis plan remain essential; as Streiner notes, whether and how to correct has been debated, making the rationale for the chosen family and target important (Streiner, 2015).
Rank #4
Methods can also be tailored to specialized designs. For example, a review of functional neuroimaging compares Bonferroni, random-field, and permutation approaches for FWER control; results from such methods depend on the setting and assumptions (comparative review).
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




