Paired t-test or Wilcoxon signed-rank test?
Both tests analyse the paired differences, not the raw columns. What each assumes about those differences, how zeros and ties are handled, and what to report.
Updated .
When each unit is measured twice (before and after an intervention, left and right eye, two assays on the same sample), the analysis is about one column of numbers: the within-unit differences. Both the paired t-test and the Wilcoxon signed-rank test reduce the data to those differences first. Everything that matters for choosing between them, including the assumptions people usually check on the wrong variables, is a property of the differences.
What each test does with the differences
Call the difference for unit i d_i = after_i - before_i.
The paired t-test is a one-sample t-test on the d_i. Its null hypothesis is that the population mean difference is zero. It estimates that mean difference and gives a confidence interval for it. Its assumptions are that the pairs are independent of each other and that the differences are approximately normally distributed, or that there are enough pairs for the sampling distribution of the mean difference to be close to normal. That is the whole list. It makes no assumption about the before and after columns being normal, and none about their variances being equal. The underlying distribution is the one Student (1908) derived for the mean of a small sample.
The Wilcoxon signed-rank test ranks the absolute differences |d_i| from smallest to largest, then sums the ranks belonging to the positive differences (Wilcoxon, 1945). Its null hypothesis is that the distribution of the differences is symmetric about zero. It assumes the pairs are independent, and that the differences are measured on a scale where their sizes can be meaningfully ordered, since it ranks magnitudes and not just signs. If the differences are symmetric, rejecting the null means their median (which then equals their mean) is not zero. If they are clearly asymmetric, the test can reject even when the median difference is zero, so “the median difference is non-zero” is a claim that needs the symmetry assumption.
The natural estimate that goes with the signed-rank test is the Hodges-Lehmann estimator: the median of all pairwise averages (d_i + d_j) / 2, including each difference with itself (Hodges and Lehmann, 1963). R labels it the “(pseudo)median”. It equals the median of the differences only when their distribution is symmetric.
If the differences are only reliable in sign (who improved and who got worse, with no trust in the size), neither test fits and the sign test is the honest choice.
Worked example: a teacher training workshop
An education researcher gives 30 teachers a 40-point assessment of assessment literacy before and after a workshop. Half are novices and half are experienced, so baseline scores cluster in two groups. The data are synthetic, generated in R 4.6.1.
Look at the right column
The before scores are bimodal, one cluster per experience level. A Shapiro-Wilk test on them gives p = 0.005, and someone checking “normality of the data” would reach for a rank test. The differences tell a different story. They range from -2 to +6 points, with most teachers gaining 2 or 3, and the Shapiro-Wilk p-value for them is 0.27. The bimodality belongs to the teachers, not to the change, and the pairing removes it.
A normality test is a weak tool for this decision either way: it has little power with few pairs and flags trivial departures with many. A dot plot or histogram of the 30 differences is more informative than any p-value about them. The example is here to show which variable to look at.
Results
| Paired t-test | Wilcoxon signed-rank | |
|---|---|---|
| Estimate | mean difference 2.2 points | Hodges-Lehmann estimate 2.5 points |
| Confidence interval | 95% CI 1.50 to 2.90 | 97.4% CI 1.5 to 3.0 |
| Test statistic | t = 6.46, df = 29 | V = 436 |
| p | < 0.0001 | < 0.0001 |
Both tests say the same thing: scores went up by roughly two to three points on a 40-point scale, and the interval excludes zero comfortably. The Wilcoxon interval is reported at 97.4% rather than 95% because the test statistic is discrete, so only certain confidence levels are exactly attainable, and R reports the level it achieved.
Note that the Hodges-Lehmann estimate (2.5) is not the sample median of the differences (2.0). It is the median of the pairwise averages of the differences, a different estimator from the sample median, so the two need not agree, particularly on integer scores where many differences tie.
For contrast, analysing the same scores as two independent groups (a Welch t-test on after versus before) gives t = 1.07, p = 0.29, and a 95% CI from -1.92 to 6.32. Ignoring the pairing throws away the fact that each teacher is compared with themselves, and the large differences between teachers swamp a consistent two-point gain.
Zeros and ties
Three teachers scored exactly the same both times, and the differences are integers, so many share a value. Both situations need a decision in the signed-rank test.
A zero difference has no sign. The conventional procedure drops zeros and ranks the rest. Pratt (1959) proposed ranking the zeros along with everything else and then leaving them out of the sum, which he argued behaves more sensibly. Software differs on this, and so can two code paths in the same package: in R 4.6.1, the exact calculation used for these results ranks the zeros (V = 436), while the large-sample approximation removes them before ranking (which would give V = 364 for these data). Earlier versions of R could not compute an exact p-value with ties or zeros and fell back to the normal approximation with a warning. None of this changes the conclusion here, but it does change the reported statistic, which is one reason to record the software version with the result.
The paired t-test has no special handling: a zero is just a difference of zero.
Effect sizes that match the design
The raw mean difference, in the units of the scale, is usually the most useful effect size: “2.2 points on a 40-point test” is something a reader can judge. If you need a standardized measure, be explicit about which one. For these data, dividing the mean difference by the standard deviation of the differences (1.86) gives d_z = 1.18. That figure depends on how strongly before and after correlate, so it is not comparable with a Cohen’s d from a between-groups design. Lakens (2013) sets out the variants (d_z, and versions standardized by the average of the before and after standard deviations) and when each is appropriate. For the signed-rank test, the matched-pairs rank-biserial correlation is the usual companion.
Choosing between them
- The differences look roughly symmetric and the mean change is what you want to report: paired t-test. With a reasonable number of pairs, moderate skew in the differences is not a problem.
- Few pairs and clearly heavy-tailed differences, or one or two extreme changes you have no grounds to exclude: Wilcoxon signed-rank, reporting the Hodges-Lehmann estimate and its interval.
- Strongly asymmetric differences: be careful with both. The mean and the median of the differences are then different quantities; decide which one answers your question, and consider a transformation if the outcome is a ratio-type measurement.
- Three or more time points: neither. Use a repeated-measures model (see choosing a statistical test).
For independent groups rather than paired measurements, the corresponding comparison is t-test vs Mann-Whitney.
How Celeus handles this
When two columns look like the same measurement taken twice, such as before and after, Celeus proposes a paired t-test and checks normality on the paired differences, not on either column. If that check fails, it offers the Wilcoxon signed-rank test as a re-run, reported with the Hodges-Lehmann estimate and its confidence interval.
Sources
- Student (1908). The probable error of a mean. Biometrika 6(1):1-25. https://doi.org/10.2307/2331554
- Wilcoxon F (1945). Individual comparisons by ranking methods. Biometrics Bulletin 1(6):80-83. https://doi.org/10.2307/3001968
- Pratt JW (1959). Remarks on zeros and ties in the Wilcoxon signed rank procedures. Journal of the American Statistical Association 54(287):655-667. https://doi.org/10.1080/01621459.1959.10501526
- Hodges JL, Lehmann EL (1963). Estimates of location based on rank tests. Annals of Mathematical Statistics 34(2):598-611. https://doi.org/10.1214/aoms/1177704172
- Lakens D (2013). Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Frontiers in Psychology 4:863. https://doi.org/10.3389/fpsyg.2013.00863
Try it on your own data
Upload a dataset and Celeus suggests a method, checks its assumptions, runs it in R, and gives you the script to rerun it.
Try it in Celeus