T-test or Mann-Whitney U test: which should you use?
Welch's t-test, Student's t-test and the Mann-Whitney U test answer different questions. What each one tests, and why Welch is a sensible default.
Updated .
The usual advice is “use a t-test if the data are normal and Mann-Whitney if they are not”. That framing hides the more important difference: the tests ask different questions. Welch’s t-test compares means. The Mann-Whitney U test (also called the Wilcoxon rank-sum test) asks whether a value drawn from one group tends to be larger than a value drawn from the other. Pick the test whose question you want answered, then check whether its assumptions are reasonable for your data.
For independent groups on a continuous outcome, a good default is Welch’s t-test, reported as a difference in means with a confidence interval. Use Mann-Whitney when a rank-based comparison is what you actually care about, as with ordinal scores, or when a small sample is badly skewed and a comparison of means is not meaningful anyway.
Three tests, three sets of assumptions
All three assume the two groups are independent: different people, plots or animals in each group, one value per unit. If the same units were measured twice, see paired t-test vs Wilcoxon signed-rank instead.
Student’s t-test
Null hypothesis: the two population means are equal. It assumes the outcome is normally distributed within each group and that the two groups have the same variance. It pools the two sample variances into one estimate. When the variances differ and the group sizes differ too, that pooled estimate is wrong in a predictable direction: if the smaller group has the larger variance, Student’s test rejects too often; if the larger group has the larger variance, it becomes too conservative.
Welch’s t-test
Same null hypothesis, equal means. It drops the equal-variance assumption: each group’s variance is estimated separately and the degrees of freedom are adjusted (Welch, 1947). When the variances really are equal, Welch loses very little power compared with Student’s test. When they are not, it keeps the false-positive rate close to the nominal level where Student’s test does not. That trade is why Delacre, Lakens and Leys (2017) argue it should be the default, and why R’s t-test function already runs Welch unless you explicitly assume equal variances.
Testing for equal variances first (with Levene’s test, say) and then choosing between Student and Welch is not a good fix. The pre-test has low power in small samples and adds its own error to the procedure. Using Welch from the start avoids the problem.
Normality matters less than it seems. The t-test depends on the sampling distribution of the mean, which approaches normality as samples grow even when the raw data are skewed. With small samples and strong skew, the confidence interval can be off, and this is where a rank test earns its place.
Mann-Whitney U test
The statistic U counts, over every pair of one observation from each group, how often the first group’s value is larger (ties count one half). Divide U by the number of pairs and you get an estimate of
P(X > Y) + 0.5 x P(X = Y),
sometimes called the probability of superiority. The test’s exact null hypothesis is that both groups come from the same distribution, and it has power against alternatives where that probability differs from 0.5, which is what Mann and Whitney (1947) set out to test: whether one variable is “stochastically larger” than the other.
It does not assume normality. It does not, however, test whether the medians are equal, and it does not escape the unequal-variance problem. Two points are worth being precise about:
- It is a test of medians only under a shift model, where the two distributions have the same shape and spread and differ only in location. Then the shift equals the difference in medians (and in means). Outside that model, two groups with identical medians can produce a significant Mann-Whitney result, and two groups with different medians can fail to (Divine et al., 2018).
- Different spreads can create significance on their own. If the groups differ in shape or variance, the test’s null distribution no longer fits, so it can reject more often than its nominal rate even when neither group tends to be larger. Fay and Proschan (2010) set out the different hypotheses the same Mann-Whitney decision rule can be read as testing, and which assumptions each reading needs.
Worked example: reaction times after sleep restriction
A psychology lab compares mean reaction times (ms, one value per participant) between 40 control participants and 25 who were sleep restricted. Reaction times are right-skewed, and the restricted group is more variable. The data are synthetic, generated in R 4.6.1.
The control group had mean 512.7 ms (SD 101.7, median 506). The restricted group had mean 586.0 ms (SD 149.9, median 582).
| Test | Estimate (restricted minus control) | 95% CI | p |
|---|---|---|---|
| Welch’s t | 73.4 ms difference in means | 4.5 to 142.2 | 0.037 |
| Student’s t | 73.4 ms difference in means | 11.1 to 135.7 | 0.022 |
| Mann-Whitney | 75 ms Hodges-Lehmann shift | 15 to 149 | 0.018 |
Welch’s test gives t = -2.16 on 37.9 degrees of freedom. Student’s test, which pools the variances, gives t = -2.35 on 63 degrees of freedom and a narrower interval. The narrower interval is not extra precision. The smaller group is the more variable one here, which is exactly the case where pooling understates the standard error. Welch’s interval is the honest one.
The Mann-Whitney test used an exact null distribution, which R 4.6.1 does by default when both groups have fewer than 50 values (a few rounded reaction times are tied, and this version of R handles ties exactly; older versions switch to a normal approximation and give a slightly different p-value). Its estimate is the Hodges-Lehmann shift, the median of all 1,000 pairwise differences between groups. R reported the interval at an achieved level of 95.1%, since exact rank-based intervals can only hit certain confidence levels. The more direct summary of what the test examines is the probability of superiority: across all restricted-control pairs, the restricted participant was slower in 67.5% of them (ties counted as half).
All three agree on direction. They should not be read as three attempts at the same number. The t-tests estimate a difference in means; the Mann-Whitney estimate is a shift that equals the difference in medians only if the two distributions have the same shape, and here the spreads clearly differ.
Same medians, significant Mann-Whitney
To see the shape sensitivity directly, we drew 1,000 values from a symmetric normal distribution and 1,000 from a right-skewed exponential distribution shifted so that its population median is also zero.
The sample medians were 0.073 and -0.024. The Mann-Whitney test still gave W = 536,678 and p = 0.0045, because values from the skewed distribution exceed values from the normal one in 53.7% of pairs. The result is correct for the question the test asks, P(b > a) differs from 0.5. It is wrong only if you report it as “the medians differ”.
Which one to report
- You want a difference in means, with a confidence interval: Welch’s t-test. With moderate to large samples this holds up well even for skewed data.
- The outcome is ordinal, or you want to know which group tends to score higher: Mann-Whitney, reported with the probability of superiority or a rank-biserial correlation, not described as a comparison of medians.
- You need a difference in medians specifically: neither test gives you that without extra assumptions. Estimate the median difference directly, for example with a bootstrap confidence interval or quantile regression.
- Small sample, heavy skew: consider whether the mean is the right summary at all. A log transformation often makes a ratio of geometric means the natural quantity.
Avoid running both and reporting whichever has the smaller p-value. Decide from the question before looking at results. If an outcome has values below a detection limit, neither plain test is right; see data below the detection limit.
How Celeus handles this
Celeus uses Welch’s t-test by default when you compare two groups on a continuous outcome, reports the mean difference with its 95% confidence interval, and checks normality and equal variances for you. If the normality check fails, it offers Mann-Whitney as a re-run and describes that result as which group tends to have larger values, not as a comparison of means.
Sources
- Welch BL (1947). The generalization of 'Student's' problem when several different population variances are involved. Biometrika 34(1-2):28-35. https://doi.org/10.1093/biomet/34.1-2.28
- Mann HB, Whitney DR (1947). On a test of whether one of two random variables is stochastically larger than the other. Annals of Mathematical Statistics 18(1):50-60. https://doi.org/10.1214/aoms/1177730491
- Delacre M, Lakens D, Leys C (2017). Why psychologists should by default use Welch's t-test instead of Student's t-test. International Review of Social Psychology 30(1):92-101. https://doi.org/10.5334/irsp.82
- Fay MP, Proschan MA (2010). Wilcoxon-Mann-Whitney or t-test? On assumptions for hypothesis tests and multiple interpretations of decision rules. Statistics Surveys 4:1-39. https://doi.org/10.1214/09-SS051
- Divine GW, Norton HJ, Baron AE, Juarez-Colunga E (2018). The Wilcoxon-Mann-Whitney procedure fails as a test of medians. The American Statistician 72(3):278-286. https://doi.org/10.1080/00031305.2017.1305291
Try it on your own data
Upload a dataset and Celeus suggests a method, checks its assumptions, runs it in R, and gives you the script to rerun it.
Try it in Celeus