The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Use scipy.stats.chisquare when you have counts for one categorical variable and want to compare them with expected frequencies (a goodness-of-fit test). Use scipy.stats.chi2_contingency when you have a table of counts for two or more categorical variables and want to test whether they are independent. Both return a test statistic and a p-value. The contingency function also returns the degrees of freedom and the expected-frequency table.
Which function fits your data?
chisquare |
chi2_contingency |
|
|---|---|---|
| Question | Do observed counts in one variable differ from expected frequencies? | Are the row and column variables independent? |
| Input | 1-D observed counts, plus optional expected counts | 2-D (or higher) table of observed counts |
| Expected values | You supply them with f_exp. If omitted, categories are assumed equally likely. |
Derived from the table margins under independence |
| Returns | statistic, p-value | statistic, p-value, degrees of freedom, expected frequencies |
In both cases the input must be counts of observations per category. Do not pass raw continuous measurements, percentages, or rates as if they were counts. Bin continuous data first, and only if binning makes sense for your question.
Goodness-of-fit with chisquare
import numpy as np
from scipy.stats import chisquare
observed = np.array([16, 18, 16, 14, 12, 12])
expected = np.array([16, 16, 16, 16, 16, 8])
result = chisquare(observed, f_exp=expected)
print(result.statistic, result.pvalue)
Both arrays must line up category by category, and both sum to 88 here. Working this by hand, the statistic is 3.5 on 5 degrees of freedom (6 categories minus 1). That gives a p-value of roughly 0.62, so these counts show no evidence of departing from the expected frequencies. Run the code to see the exact value.
The null hypothesis in SciPy’s documentation is that the observations are sampled independently from a categorical distribution with the expected frequencies you gave.
#1 Best Overall
Equal-probability shortcut
For a die or any “all categories equally likely” hypothesis, call chisquare(observed) with no f_exp.
Expected proportions, not counts
If your hypothesis is a set of proportions, multiply them by the sample size to get expected counts. For example, expected = np.array([0.5, 0.3, 0.2]) * observed.sum(). This also guarantees the totals match. The Pearson p-value is only accurate when observed and expected totals agree, and SciPy’s sum_check option guards against a mismatch.
Rank #2
When parameters were estimated from the data
If you fitted distribution parameters to produce the expected counts, the default degrees of freedom (categories minus 1) are too many. The ddof argument adjusts them. SciPy documents k - 1 - p for the efficient maximum-likelihood case, where p is the number of estimated parameters. It also warns that in some situations the asymptotic distribution is not chi-square at all. Treat such models with care.
Test of independence with chi2_contingency
import numpy as np
from scipy.stats import chi2_contingency
table = np.array([[10, 10, 20],
[20, 20, 20]])
res = chi2_contingency(table)
print(res.statistic, res.pvalue)
print(res.dof)
print(res.expected_freq)
Rows and columns are the categories of the two variables, and each cell is a count. Here the margins imply expected counts of 12, 12, 16 in the first row and 18, 18, 24 in the second. By hand, the statistic is about 2.78 with 2 degrees of freedom, and the p-value is about 0.25. That is no evidence that the two variables are associated in this table.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →SciPy describes the function as a test for the independence of different categories of a population.
Building the table from raw data
If you have one row per observation, cross-tabulate first. With pandas: table = pd.crosstab(df["group"], df["outcome"]), then pass table (or table.to_numpy()) to chi2_contingency.
Options inside chi2_contingency
- Yates’ continuity correction (
correction=True, the default). It applies only when the degrees of freedom equal 1, which means a 2×2 table. It moves each observed count 0.5 toward its expected count, which makes the result more conservative. Setcorrection=Falsefor the uncorrected Pearson statistic. - Statistic choice (
lambda_). The default is Pearson’s chi-square. Other values select members of the Cressie-Read power-divergence family, such as the log-likelihood-ratio (G-test) statistic.chisquareacceptslambda_as well. - Resampling p-values (
method). In the SciPy 1.18.0 documentation, permutation and Monte Carlo p-values are supported only for a two-way table withcorrection=Falseand the defaultlambda_. The Monte Carlo configuration draws tables withscipy.stats.random_table. These options can help when counts are sparse, but check your installed SciPy version’s documentation, because this behavior is version-sensitive.
Check the assumptions before trusting the p-value
- Expected counts. The chi-square p-value is an asymptotic approximation. SciPy cites “at least 5” for observed and expected cell frequencies as an often-quoted guideline, and warns that small counts can invalidate the test. It is a diagnostic rule of thumb, not a guarantee in either direction. For contingency tests, inspect
res.expected_freqand look at the smallest value. - Independent observations. Each observation should fall into one category and be sampled independently. Repeated measures on the same subjects, or paired before/after data, violate this.
- Sparse tables. If expected counts are too small, switch to a method suited to your design. SciPy points to Fisher’s exact test (
scipy.stats.fisher_exact) for 2×2 tables and to exact alternatives such as Barnard’s test (scipy.stats.barnard_exact). Which is appropriate depends on how the data were collected, for example whether the margins were fixed in advance. Alternatively, merge categories where that is meaningful, or use the resampling option above.
Interpreting the result
A small p-value means the observed counts would be unlikely if the null hypothesis were true. It does not tell you the following:
- Which cells drive the result. The test is two-sided and omnibus. Compare
tablewithres.expected_freq, or compute standardized residuals(table - res.expected_freq) / np.sqrt(res.expected_freq)to see where counts are higher or lower than independence predicts. - Direction or practical size. With large samples, trivial differences become “significant”. Add an effect size. SciPy provides
scipy.stats.contingency.association, which computes Cramér’s V, among other measures:
from scipy.stats.contingency import association
v = association(table, method="cramer")
print(v)
A large p-value is not proof of independence either. It means these data do not provide evidence against it, which is a weaker claim, especially with small samples.
Quick Recap
Best Value
What to report
- The test type (goodness-of-fit or independence) and the observed counts, or a clear reference to the table.
- For goodness-of-fit: the expected proportions or counts, and whether any parameters were estimated (and the
ddofused). - For independence: the expected-count check.
- The statistic, degrees of freedom, and p-value.
- Any continuity correction, alternative
lambda_, or resampling method. - An effect size such as Cramér’s V, rather than the p-value alone as a measure of strength.
Common mistakes
- Using
chisquareon a two-way table. It treats the array as separate observed values rather than testing independence, so pass cross-tabulations tochi2_contingency. - Passing percentages or proportions as observed counts. The statistic depends on sample size, so it will be wrong.
- Supplying expected counts that do not sum to the observed total.
- Ignoring the automatic Yates correction on 2×2 tables when comparing against results from another tool that does not apply it.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




