Celeus

How to choose a statistical test for your data

Pick a test from four facts about your study: the question, the outcome type, how observations are linked, and how many groups. With a decision table.

Updated .

Most flowcharts for choosing a test start with “is your data normally distributed?”. That is the wrong first question. The test follows from how the study was designed and what you want to estimate. The shape of the data matters later, and less than people think. Settle four things, in this order, and the choice usually makes itself.

Four questions, in order

1. What quantity do you want to report?

Write the sentence you hope to put in the results section before you pick a method. “Mean soil nitrate was 2.4 mg/kg lower under cover crops” calls for a difference in means with a confidence interval. “The odds of a correct response were 1.8 times higher in the cued condition” calls for an odds ratio, which means logistic regression. “Half of the seedlings survived past day 40” is a survival estimate. Each target quantity has a family of methods built to estimate it, and a test that estimates something else will answer a question you did not ask.

This also forces the difference between comparing groups (is there a difference in the outcome between conditions?) and describing a relationship (how does the outcome change as a continuous predictor changes?). The first leads to t-tests and their relatives; the second to correlation and regression.

2. What type is the outcome?

  • Continuous: reaction time, biomass, a test score treated as a measurement.
  • Binary: survived or not, correct or incorrect.
  • Ordinal: a single Likert item, a severity grade. The categories are ordered but the gaps between them are not known to be equal.
  • Count: number of nests per transect, errors per session.
  • Time to an event, possibly censored: time until germination, when some seeds had not germinated when the study ended.

A total score summed over many Likert items usually behaves well enough to treat as continuous. A single five-point item usually does not.

3. How are the observations linked?

This is the question that most often goes wrong, and it matters more than any distributional assumption. Observations are independent when each unit (participant, plot, animal) contributes one value. They are paired when each unit is measured twice, or when units come in natural matched pairs (twins, left and right leaves, two plots at one site). They are repeated or clustered when a unit is measured three or more times, or when units are nested inside larger units (pupils in classrooms, quadrats in sites).

Linked observations carry shared information. A test that treats them as independent gets the standard error wrong, sometimes very wrong, as the example below shows.

4. How many groups or conditions?

Two groups, or three or more. With three or more, an overall test is usually followed by pairwise comparisons, which brings in the problem of multiple comparisons.

Decision table

The first column is the method to reach for by default. The second is what to consider, and why. It is not a mechanical rule.

Outcome and design Default choice Consider instead
Continuous, 2 independent groups Welch’s t-test Mann-Whitney U, if a rank-based question fits better (guide)
Continuous, 2 paired measurements Paired t-test Wilcoxon signed-rank (guide)
Continuous, 3+ independent groups Welch’s ANOVA Classic one-way ANOVA if spreads are similar; Kruskal-Wallis for ranks
Continuous, 3+ measurements per unit, or clustered units Linear mixed model Repeated-measures ANOVA for complete, balanced data
Continuous outcome and continuous predictor Linear regression, or Pearson correlation Spearman correlation for a monotonic but non-linear relationship
Binary, 2+ independent groups Chi-square test Fisher’s exact test when expected counts are small (guide)
Binary, paired McNemar’s test
Binary, with covariates to adjust for Logistic regression (guide)
Ordinal, 2 independent groups Mann-Whitney U Ordinal logistic regression, especially with covariates
Counts Poisson regression Negative binomial regression when variance exceeds the mean
Time to event, with censoring Kaplan-Meier curves, Cox regression (guide)

Any of the regression models can replace the simpler tests in their row once you need to adjust for other variables. Student’s t-test is a linear regression with one binary predictor, and a chi-square test on a 2 x 2 table asks the same question as a logistic regression with one binary predictor.

Worked example: the design decides, not the distribution

A soil scientist runs a field trial at 12 sites. Each site has two plots, one with a winter cover crop and one without, and residual soil nitrate (mg/kg) is measured in both. Sites differ a lot in their baseline nitrate. The data are synthetic, generated in R 4.6.1.

Treating the 24 plots as two independent groups gives a mean difference of -2.38 mg/kg, t = -0.85 on about 22 degrees of freedom, p = 0.41, and a 95% confidence interval from -8.18 to 3.43 mg/kg. Read naively, the cover crop made no detectable difference.

The paired analysis uses the same 24 numbers. The mean difference is identical, -2.38 mg/kg, but now t = -7.99 on 11 degrees of freedom, p < 0.0001, and the 95% confidence interval runs from -3.03 to -1.72 mg/kg.

Nothing about the distribution changed. What changed is that the paired test removes the site-to-site variation, which in this trial is far larger than the treatment effect (the two plots at a site correlate at 0.99). The independent analysis buries a reduction that shows up at all 12 sites (2.4 mg/kg on average) under differences between sites. The reverse mistake also happens: treating clustered data as independent, for example analysing every pupil as if the class they sit in did not matter, usually makes standard errors too small and p-values too optimistic.

What the table does not decide for you

Assumption pre-tests are a poor gatekeeper. Running a Shapiro-Wilk or Levene test first and choosing the analysis based on its p-value sounds careful, but when the pre-test and the main test use the same observations, the error rates of the whole procedure are no longer the ones you think you have. After a systematic simulation study, Rasch and colleagues recommended skipping the pre-tests and using the Welch test as the standard two-group test (Rasch et al., 2011). A normality test has little power in small samples, where normality matters most, and flags trivial departures in large ones, where it matters least.

Large samples forgive a lot. For comparing means, the t-test and linear regression depend on the sampling distribution of the mean, not on the raw data being normal. With moderate to large samples that distribution is close to normal even for skewed outcomes (Lumley et al., 2002). Switching to a rank test in a large study because a histogram looks skewed changes the question being asked, often without anyone noticing (Fagerland, 2012).

Rank tests answer different questions. The Mann-Whitney U test is not simply “the t-test for non-normal data”. It compares how often a value from one group exceeds a value from the other. That can be a better question, especially for ordinal outcomes, but it is a different one.

Report estimates, not just p-values. Whatever test you choose, report the effect (a difference in means, an odds ratio, a hazard ratio) with its confidence interval. A p-value alone says nothing about how large the effect is or how precisely it was estimated, and many common readings of p-values are wrong (Greenland et al., 2016).

Decide before you look. If you try three tests and report the one with the smallest p-value, the p-value no longer means what it claims. Choose the method from the design, write it down, and treat any other analysis as a sensitivity check.

How Celeus handles this

When you upload a dataset, Celeus proposes suitable analyses from the structure of your columns, including a paired comparison when two columns look like the same measurement taken twice, and grades the alternatives against checks run on your data, so that, for instance, Fisher’s exact test is preferred when expected counts are too small for chi-square. You can see how its results compare with independent computations on the validation page.

Sources

  1. Rasch D, Kubinger KD, Moder K (2011). The two-sample t test: pre-testing its assumptions does not pay off. Statistical Papers 52(1):219-231. https://doi.org/10.1007/s00362-009-0224-x
  2. Lumley T, Diehr P, Emerson S, Chen L (2002). The importance of the normality assumption in large public health data sets. Annual Review of Public Health 23:151-169. https://doi.org/10.1146/annurev.publhealth.23.100901.140546
  3. Fagerland MW (2012). t-tests, non-parametric tests, and large studies: a paradox of statistical practice? BMC Medical Research Methodology 12:78. https://doi.org/10.1186/1471-2288-12-78
  4. Greenland S, Senn SJ, Rothman KJ, Carlin JB, Poole C, Goodman SN, Altman DG (2016). Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. European Journal of Epidemiology 31(4):337-350. https://doi.org/10.1007/s10654-016-0149-3

Try it on your own data

Upload a dataset and Celeus suggests a method, checks its assumptions, runs it in R, and gives you the script to rerun it.

Try it in Celeus