Multiple comparisons: Bonferroni, Holm, or Benjamini-Hochberg?
Family-wise error versus false discovery rate, how Bonferroni, Holm and Benjamini-Hochberg differ, and which post hoc test follows ANOVA or Kruskal-Wallis.
Updated .
Every test run at the 0.05 level carries a 5% chance of a false positive when its null hypothesis is true. Run one test and that is the error rate you signed up for. Run 20 independent tests of true nulls and the chance of at least one false positive is 1 - 0.95²⁰, about 64%. Multiple-comparison procedures exist to put a bound back on that. The hard part is not the arithmetic, which is simple. The hard part is deciding what you want to bound, and over which set of tests.
A screen that shows the problem
Synthetic data from a gene-expression screen: 2,000 genes measured in 8 control and 8 treated samples. The first 100 genes are truly shifted by 2 standard deviations; the other 1,900 are not changed at all. Each gene gets a Welch t-test.
| Rule (at 0.05) | Genes called | Truly changed | False positives |
|---|---|---|---|
| No adjustment | 182 | 97 | 85 |
| Bonferroni | 6 | 6 | 0 |
| Holm | 7 | 7 | 0 |
| Benjamini-Hochberg | 45 | 44 | 1 |
Without adjustment, 85 of the 182 “hits” are noise (1,900 x 0.05 = 95 were expected), so almost half the list is wrong. Bonferroni and Holm make no false calls but find only 6 or 7 of the 100 real changes. Benjamini-Hochberg finds 45, of which one is false. None of these rows is “correct”. They answer different questions.
Two different error rates
Family-wise error rate (FWER) is the probability of making even one false rejection across the whole family of tests. Controlling it at 5% means that, whatever mix of true and false nulls you have, the chance of any false positive at all is at most 5%. That is the right target when a single false claim is costly: a small set of confirmatory hypotheses, each of which will be reported as a finding on its own.
False discovery rate (FDR) is the expected proportion of false rejections among the rejections you make (Benjamini and Hochberg, 1995). Controlling it at 5% accepts that some calls will be wrong, but keeps them to about 5% of the list on average. That is the right target for screening, where the output is a candidate list that will be followed up, and where losing most of the real signal is the bigger cost. In the screen above, 1 false call out of 45 is about 2%. That is one run, so it neither confirms nor breaks a guarantee about the average, but it is the kind of list the target describes.
Three procedures
Bonferroni
Multiply each p-value by the number of tests m (capping at 1), or equivalently test each at α/m. With 2,000 genes that is a per-test threshold of 0.000025. It controls FWER under any dependence between the tests, and it is the most conservative option here.
Holm
Sort the p-values from smallest to largest. Compare the smallest to α/m, the next to α/(m - 1), and so on, stopping at the first one that fails; everything before it is rejected (Holm, 1979). Holm controls FWER under exactly the same conditions as Bonferroni and always rejects at least as many hypotheses. The practical conclusion: if you would use Bonferroni, use Holm instead. There is no situation in which Bonferroni rejects something Holm does not. The difference is often small, as here (7 versus 6), because the gain comes from the later steps.
Benjamini-Hochberg
Sort the p-values, find the largest rank k for which p(k) ≤ (k/m) x q, and reject the k smallest. Here q is the target FDR. The procedure controls FDR when the tests are independent, and Benjamini and Yekutieli (2001) showed it still holds under a form of positive dependence between the tests. BH-adjusted p-values are often read as “the smallest FDR at which this test would be called”.
A BH-adjusted value is not the probability that the particular gene is a false positive. It is a property of the list you would get by cutting at that level.
Defining the family is the real decision
No procedure can tell you which tests belong together. Some working rules:
- Tests that answer one question form a family. All pairwise comparisons after an omnibus test, all genes in a screen, all outcomes feeding a single claim of “the intervention worked”.
- Pre-specified primary analyses are usually not adjusted against exploratory ones. A single primary outcome, stated before data collection, is tested at the nominal level; secondary and exploratory results are labelled as such, with or without their own correction.
- Report the number of tests you ran, not just the ones that came out. A correction applied to the 5 tests you chose to show, after running 40, controls nothing.
- Correction does not rescue a fishing expedition. It makes the error rate honest for the tests you declare. It does not make post hoc hypotheses confirmatory.
Confidence intervals need the same care. If you report intervals for all pairwise differences and draw conclusions from which ones exclude zero, use simultaneous intervals (such as Tukey’s) rather than a set of unadjusted 95% intervals.
Post hoc tests after ANOVA and Kruskal-Wallis
A significant one-way ANOVA or Kruskal-Wallis test says the groups are not all the same. It does not say which ones differ. The pairwise follow-up is itself a family of k(k - 1)/2 comparisons, and the standard procedures build the correction in.
- Tukey’s HSD after a one-way ANOVA with roughly equal variances. It gives simultaneous confidence intervals for every pairwise mean difference and controls FWER across all of them, using the pooled variance. It handles unequal group sizes (the Tukey-Kramer form, which R uses automatically).
- Games-Howell when variances differ, the natural partner to Welch’s ANOVA. Each pair gets its own variance estimate and degrees of freedom.
- Dunn’s test after Kruskal-Wallis (Dunn, 1964). It compares mean ranks from the joint ranking used by the omnibus test, which keeps the follow-up consistent with it, and is paired with a p-value adjustment such as Holm or Benjamini-Hochberg. Running separate Mann-Whitney tests for each pair re-ranks the data for every pair and does not match the omnibus test.
Tukey’s procedure controls FWER on its own, so it does not strictly require a significant F-test first. Requiring one anyway makes the combined procedure more conservative.
Example: one gene across four conditions
A synthetic follow-up of one gene: expression in a control and three drug conditions, 8 samples each.
The ANOVA gives F(3, 28) = 12.15, p = 0.00003. Tukey’s HSD then shows:
| Comparison | Difference | 95% simultaneous CI | Adjusted p |
|---|---|---|---|
| drugB - ctrl | 1.63 | 0.46 to 2.81 | 0.004 |
| drugC - ctrl | 1.75 | 0.57 to 2.93 | 0.002 |
| drugA - ctrl | -0.28 | -1.46 to 0.90 | 0.92 |
| drugC - drugB | 0.11 | -1.07 to 1.29 | 0.99 |
(The two remaining pairs, drugB and drugC against drugA, have adjusted p below 0.001.) Drugs B and C raise expression relative to control and to drug A. There is no evidence that B and C differ from each other, and the interval for that difference, -1.07 to 1.29, is too wide to claim they are equivalent. Holm-adjusted pairwise t-tests lead to the same conclusions here, with adjusted p of 0.002 for drugB versus control and 0.0015 for drugC versus control. Tukey adds the simultaneous intervals, which are what you should report.
A short decision guide
- A handful of confirmatory hypotheses: Holm.
- All pairwise comparisons after ANOVA: Tukey (or Games-Howell with unequal variances).
- Pairwise comparisons after Kruskal-Wallis: Dunn’s test with Holm.
- Many exploratory tests producing a candidate list: Benjamini-Hochberg, reported as an FDR.
- Always: say which procedure, over how many tests, and report effect sizes with intervals alongside adjusted p-values.
How Celeus handles this
After an overall test across three or more groups, Celeus runs the matching post hoc comparisons for you (Tukey’s HSD, Games-Howell or Dunn’s test), each with an effect size and its confidence interval. When you combine several analyses into one report, it also tells you which primary results stay below 0.05 after a Holm adjustment across the report, while deciding which analyses form a family stays your call.
Sources
- Holm S (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics 6(2):65-70. https://www.jstor.org/stable/4615733
- Benjamini Y, Hochberg Y (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society Series B 57(1):289-300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x
- Benjamini Y, Yekutieli D (2001). The control of the false discovery rate in multiple testing under dependency. Annals of Statistics 29(4):1165-1188. https://doi.org/10.1214/aos/1013699998
- Dunn OJ (1964). Multiple comparisons using rank sums. Technometrics 6(3):241-252. https://doi.org/10.1080/00401706.1964.10490181
Try it on your own data
Upload a dataset and Celeus suggests a method, checks its assumptions, runs it in R, and gives you the script to rerun it.
Try it in Celeus