Chi-square or Fisher's exact test: which one should you use?
When the chi-square approximation holds, where the expected-count-of-5 rule comes from and what it says, and when to use Fisher's exact or McNemar's test.
Updated .
Both tests ask the same question of a table of counts: are the row and column variables associated? They differ in how the p-value is obtained. The chi-square test uses a large-sample approximation; Fisher’s exact test computes the p-value from an exact conditional distribution. When the table is well filled, they agree closely and the choice hardly matters. When some cells are expected to be nearly empty, the approximation can be poor, and that is the situation the famous “expected count of 5” rule is trying to catch. If the two variables were measured on the same units (before and after, or two raters on the same items), neither test is correct, and you need McNemar’s test.
What the chi-square test actually tests
For an r x c table, Pearson’s statistic sums (observed - expected)² / expected over every cell, where each expected count is row total x column total / grand total. That is the count you would expect if the two variables were independent, given the margins you observed. The null hypothesis is independence (or, when one margin was fixed by design, equal proportions across groups). Under that null, and with enough data, the statistic follows approximately a chi-square distribution with (r - 1)(c - 1) degrees of freedom. The degrees of freedom are Fisher’s correction of Pearson’s original count (Fisher, 1922).
The key word is approximately: the statistic is built from discrete counts, the chi-square distribution is continuous, and the approximation can be noticeably wrong when some expected counts are small.
Where “expected count of at least 5” comes from
The rule usually taught is “do not use chi-square if any expected count is below 5”. Its usual source is Cochran (1954), and what Cochran actually recommended is more specific than the classroom version: stricter for small samples, more lenient for large samples and larger tables:
- For a 2 x 2 table, use Fisher’s exact test when the total N is below 20, and also when N is between 20 and 40 and the smallest expected count is below 5. With N above 40, the (continuity-corrected) chi-square was acceptable.
- For larger tables, the approximation was considered adequate if no expected count is below 1 and no more than about 20% of expected counts are below 5.
Three points get lost in retelling. First, the rule concerns expected counts, not observed ones. An observed zero is not by itself a problem; a table whose margins make a cell’s expected count 0.8 is. Second, Cochran described the threshold of 5 as a somewhat arbitrary choice, not a derived boundary (Campbell, 2007, notes this directly). Third, later simulation work suggests the traditional rule is conservative. Campbell (2007) compared seven two-sided tests for 2 x 2 tables and found the best policy was the “N - 1” chi-square (Pearson’s statistic multiplied by (N - 1)/N) whenever every expected count is at least 1, falling back to Fisher’s test only below that.
The defensible position: the approximation is fine for well-filled tables, gets shaky as expected counts fall toward 1 to 5, and when in doubt an exact test costs little.
What Fisher’s exact test conditions on
Fisher’s test treats both sets of margins as fixed. Given the row and column totals, the count in any one cell of a 2 x 2 table follows a hypergeometric distribution under the null, so the probability of every possible table with those margins can be computed exactly. The two-sided p-value in R sums the probabilities of all tables that are no more probable than the one observed.
“Exact” means the p-value is not an approximation. It does not mean the test is ideal. Because only a limited number of tables share the observed margins, the achievable p-values are discrete, and the test’s real type I error rate is usually below the nominal 5%, sometimes well below it. Conditioning on both margins is also hard to justify when they were not fixed by the design (a survey where group sizes and outcome counts were both free to vary), which is the usual case in practice. That is the reason Campbell’s recommendation prefers the N - 1 chi-square when expected counts permit.
Fisher’s test also gives you an effect size: R reports the conditional maximum-likelihood odds ratio with an exact confidence interval. It is not the same number as the simple cross-product odds ratio, as the example below shows.
Larger tables
Both tests extend beyond 2 x 2. The chi-square test works unchanged with (r - 1)(c - 1) degrees of freedom. The exact test generalises to r x c tables (often called the Fisher-Freeman-Halton test), although it can become slow for large tables with large counts, in which case a Monte Carlo p-value is a practical substitute. With more than one degree of freedom, a significant result says only that some association exists somewhere in the table. Look at the standardized residuals or the proportions to see where, and treat any follow-up cell-by-cell tests as multiple comparisons.
A worked example: newts and fish in a pond survey
Synthetic data: 30 ponds, 12 with predatory fish and 18 without, each surveyed for the presence of a newt species.
The observed table:
| absent | present | |
|---|---|---|
| fish | 9 | 3 |
| no fish | 2 | 16 |
The expected counts are 4.4, 7.6, 6.6 and 11.4. One of four cells (25%) is below 5 and none is below 1. By the strict classroom rule, chi-square is out. By Cochran’s own 2 x 2 guidance (N = 30, smallest expected count below 5), Fisher’s test is advised. By Campbell’s criterion (all expected counts at least 1), the N - 1 chi-square is acceptable. R warns that the approximation may be incorrect for this table.
The results:
- Pearson chi-square without correction: X² = 12.66, df = 1, p = 0.00037.
- N - 1 chi-square: 12.66 x 29/30 = 12.23, p = 0.00047.
- Yates-corrected chi-square: X² = 10.05, p = 0.0015.
- Fisher’s exact test: p = 0.0012; conditional odds ratio 20.5, 95% CI 2.6 to 291.
Every version reaches the same conclusion, which is typical: the choice matters most near a threshold. The more useful output is the effect size. The simple cross-product odds ratio is 9 x 16 / (3 x 2) = 24, while Fisher’s conditional estimate is 20.5, and the confidence interval runs from about 2.6 to 291. With 30 ponds the direction is clear and the magnitude is poorly pinned down. Report the interval.
Paired binary data need McNemar’s test
Suppose the same 30 ponds are surveyed again the following year. A chi-square or Fisher test on a “year x presence” table would treat 60 pond-years as independent, which they are not. The right table cross-classifies each pond by its status in both years.
Three ponds lacked newts both years, 16 had them both years, 8 gained them and 3 lost them. Concordant ponds carry no information about change. McNemar’s test (McNemar, 1947) asks whether the two kinds of discordant pond are equally likely, which is the same as asking whether the proportion with newts is the same in both years. With continuity correction the statistic is (|8 - 3| - 1)² / 11 = 1.45, p = 0.23. With only 11 discordant pairs the exact binomial version is the better choice, and it gives p = 0.23 as well. The paired odds ratio is 8/3 = 2.7, and the honest summary is that 11 changes are too few to say whether occupancy rose.
Practical checklist
- Were the two variables measured on the same units? Use McNemar (or its exact binomial form) for paired binary data.
- Compute the expected counts, not just the observed ones.
- Well-filled table: chi-square. Some expected counts small, or you simply prefer not to argue about it: Fisher’s exact test.
- Report an effect size with a confidence interval (odds ratio, risk ratio or risk difference for 2 x 2; Cramér’s V or the table of proportions for larger tables).
- If you need to adjust for other variables, a contingency-table test is no longer enough; see logistic regression.
How Celeus handles this
Celeus checks the expected counts for you and recommends Fisher’s exact test when the table is too sparse for chi-square. For a 2 x 2 table it reports the risk ratio, risk difference and odds ratio with 95% confidence intervals, and when two binary columns look like a paired measurement it suggests McNemar’s test instead.
Sources
- Fisher RA (1922). On the interpretation of chi-square from contingency tables, and the calculation of P. Journal of the Royal Statistical Society 85(1):87-94. https://doi.org/10.2307/2340521
- Cochran WG (1954). Some methods for strengthening the common chi-square tests. Biometrics 10(4):417-451. https://doi.org/10.2307/3001616
- Campbell I (2007). Chi-squared and Fisher-Irwin tests of two-by-two tables with small sample recommendations. Statistics in Medicine 26(19):3661-3675. https://doi.org/10.1002/sim.2832
- McNemar Q (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12(2):153-157. https://doi.org/10.1007/BF02295996
Try it on your own data
Upload a dataset and Celeus suggests a method, checks its assumptions, runs it in R, and gives you the script to rerun it.
Try it in Celeus