Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

A p-value and a critical value are not competing measures. They are two ways to make a decision in the same hypothesis test: compare the p-value with the chosen significance level, or compare the test statistic with a cutoff. When the test, direction, and assumptions match, both methods normally lead to the same conclusion.

The terms to keep straight

In a hypothesis test, the analyst starts with a null hypothesis (H0) and an alternative hypothesis (HA). A test statistic summarizes the sample in a form that can be compared with a reference distribution under the null. The chosen significance level, written as α, sets the decision threshold for the test procedure. The p-value is a probability; the critical value is a cutoff on the test-statistic scale.

Term What it is What you compare it with
Test statistic A value calculated from the sample, such as z, t, χ², or F A critical value or a reference distribution
Significance level (α) A probability threshold selected for the testing procedure The p-value
Critical value A boundary defining a rejection region under the null distribution The observed test statistic
P-value A tail probability calculated under the null model α

For a conventional test, the decision rules are:

  • P-value method: reject H0 if p ≤ α.
  • Critical-value method: reject H0 if the observed statistic falls in the rejection region.

Do not compare a p-value with a critical value. A p-value is compared with α; a test statistic is compared with a critical value. NIST describes the two procedures as analogous ways to define the same rejection decision.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What a p-value means

A p-value is the probability, assuming the null hypothesis and the test model are true, of obtaining a result at least as extreme as the observed test statistic. “At least as extreme” depends on the alternative hypothesis and the test’s definition. It means a more extreme result in the right tail for a right-tailed test, in the left tail for a left-tailed test, or in the relevant tails for a two-tailed test. NIST’s definition is conditional on the null hypothesis.

#1 Best Overall

A p-value is not the probability that the null hypothesis is true, nor the probability that the alternative is true. It does not show that a result occurred “by chance alone,” quantify the size of an effect, or say whether an effect matters in practice. The American Statistical Association’s guidance on p-values urges interpretation in the context of study design, analysis choices, and other evidence.

What a critical value means

A critical value is the boundary between the rejection region and the rest of the test-statistic distribution. Its location depends on the null distribution, α, the tail direction, and—where relevant—the degrees of freedom. NIST defines a critical value as a cutoff associated with the rejection region.

For standard-normal z-tests, common cutoffs are:

Test direction α Reject when
Right-tailed 0.05 z > 1.645
Left-tailed 0.05 z < −1.645
Two-tailed 0.05 z < −1.96 or z > 1.96
Two-tailed 0.01 |z| > 2.576

These are not universal cutoffs. A one-sample mean test with unknown population standard deviation usually uses a t distribution with n − 1 degrees of freedom, while χ² and F tests use their own distributions and degrees of freedom. See NIST’s guidance on the one-sample t-test.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why the two methods usually agree

Choosing α defines how much probability lies in the rejection region under the null. The edge of that region is the critical value. The p-value measures the tail probability beyond the observed statistic. Thus, for a correctly specified test, a statistic in the rejection region has a p-value no greater than α—and a p-value no greater than α corresponds to a statistic in that region:

Rank #2
Sale
Statistics Laminate Reference Chart: Parameters, Variables, Intervals, Proportions (Quickstudy: Academic )
  • This guide is a perfect overview for the topics covered in introductory statistics courses.

Statistic in rejection region ⇔ p ≤ α

This equivalence assumes both approaches use the same test statistic, null distribution, alternative, tail convention, and assumptions. Discrete tests can have conservative rejection regions or p-values that do not correspond to a perfectly sharp continuous cutoff; exact and approximate procedures can also differ. In those cases, follow the definition of the particular test rather than assuming every implementation maps identically.

Worked example: right-tailed z-test

Suppose an analyst tests whether a population mean is greater than 100:

H0: μ = 100
HA: μ > 100
α = 0.05; observed z = 2.10

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Critical-value method: The right-tail critical value for a standard-normal test at α = 0.05 is 1.645. Since 2.10 > 1.645, reject H0.

P-value method: The right-tail probability for z = 2.10 is about 0.0179. Since 0.0179 < 0.05, reject H0.

Both methods say the data provide statistically significant evidence, at the 5% level, in favor of μ > 100 under the test model. They do not establish that the alternative is true or that any difference is practically important.

One-tailed and two-tailed tests must match

The alternative hypothesis determines which results count as evidence against the null:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Right-tailed: HA: θ > θ0. Large positive statistics support the alternative.
  • Left-tailed: HA: θ < θ0. Large negative statistics support it.
  • Two-tailed: HA: θ ≠ θ0. Extreme results in either direction count against the null.

For a symmetric standard-normal test at α = 0.05, a two-tailed test allocates 0.025 to each tail, giving cutoffs of about −1.96 and 1.96. A right-tailed test at the same α puts the full 0.05 in the right tail, giving a cutoff of 1.645. A two-sided p-value must not be paired with a one-sided critical cutoff. Decide the test direction before examining the results; switching from two-tailed to one-tailed after seeing the observed direction can invalidate the intended error rate.

Which approach should you use?

Neither approach is generally more accurate when both are correctly specified. Choose based on the task:

  • Use a p-value when reporting results or when readers need to see how the result compares with different thresholds. It is commonly supplied by statistical software and communicates more than a simple threshold-crossing decision.
  • Use a critical-value rule when a test, examination, protocol, quality-control procedure, or other formal process specifies a fixed rejection boundary in advance.

In research reporting, give the test statistic and degrees of freedom where applicable, the p-value, and the prespecified α. Also report the estimated effect and an uncertainty interval. A p-value of 0.049 and one of 0.001 both pass a 0.05 threshold, but they are different results; neither alone conveys effect size or practical importance.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the decision does—and does not—say

α is the planned Type I error rate: the probability of rejecting a true null under the assumptions of the procedure. Common choices include 0.10, 0.05, and 0.01, but 0.05 is a convention, not a natural boundary. NIST notes that the choice of α is partly conventional.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If the test does not cross the rejection threshold, say “fail to reject the null hypothesis,” not “prove” or “accept” it. A nonsignificant result may reflect no meaningful effect, but it may also result from high variability, a small sample, low power, or a poorly fitting model. To assess whether effects of practical importance remain plausible, look at the estimate and its confidence interval.

For many matched two-sided tests, rejecting H0: θ = θ0 at α = 0.05 corresponds to a 95% confidence interval that excludes θ0; failing to reject corresponds to an interval that includes it. The procedures and assumptions must match. A frequentist 95% interval does not mean there is a 95% probability that the fixed parameter lies inside this particular interval. See NIST’s discussion of the test–interval correspondence.

Statistical significance also is not practical significance. A tiny effect can yield a small p-value in a very large sample; an important effect can fail to reach a threshold in a small or noisy study. Interpret the magnitude and precision of the estimate, the consequences of the effect, and the quality of the design. NIST distinguishes practical and statistical significance.

Common mistakes

Mistake Correct interpretation
Comparing the p-value with the critical value Compare p with α, or compare the test statistic with its critical value.
Calling α the critical value α is a probability; a critical value is on the test-statistic scale.
Calling the p-value the chance that H0 is true It is calculated assuming the null and model are true.
Interpreting “fail to reject” as proof of no effect The test did not provide sufficient evidence for rejection; it did not prove the null.
Treating significance as importance Assess the estimated effect, uncertainty, and real-world context.
Using a generic cutoff without distribution or degrees of freedom Identify the test distribution and its degrees of freedom before finding a critical value.
Treating p = 0.049 and p = 0.051 as scientifically opposite A strict 0.05 rule gives different binary decisions, but the values are not separated by a meaningful evidential cliff. Be cautious with rounded values near the threshold.

If many hypotheses are tested, or results are repeatedly checked as data arrive, a nominal p-value may not preserve the intended overall false-positive rate. Multiple-comparison adjustments or sequential-testing methods may be needed for the inferential goal. P-values also depend on the model, sampling design, test, and analysis choices, so values from unrelated tests are not automatically comparable. The ASA statement emphasizes the importance of knowing how many analyses were conducted and how results were selected for reporting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A reporting template

“We tested H0: [parameter = value] against a [left-/right-/two-sided] alternative using a [test name]. The test statistic was [value] ([degrees of freedom, if applicable]), yielding p = [value]. At the prespecified α = [value], we [reject/fail to reject] H0. The estimated effect was [estimate], with [confidence interval].”

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.