What makes a statistical analysis reproducible?
Seeds, pinned package versions, the exact script, session information and a decision record: what reproducibility needs and how it differs from replication.
Updated .
A reviewer asks why your adjusted estimate is 1.84 when the table in the supplement says 1.87. You open the project folder from eighteen months ago and find analysis_final.R, analysis_final_v2.R, and a spreadsheet with a column you do not remember recoding. You rerun the script that looks most recent and get 1.81.
Nothing fraudulent happened. But you can no longer show how the published number was produced, and that is the problem reproducibility is meant to solve.
Reproducible is not the same as replicable
The two words are used loosely and sometimes in opposite senses across fields, which causes real confusion. The National Academies (2019) settled on definitions that are now widely used:
- Reproducibility means getting consistent results from the same data with the same code, methods and conditions of analysis.
- Replicability means getting consistent results in a new study that asks the same question and collects its own data.
Goodman, Fanelli and Ioannidis (2016) split the idea further into three parts. Methods reproducibility is whether enough detail is given for someone else to repeat the procedures exactly. Results reproducibility is what most people mean by replication: an independent study reaching the same result. Inferential reproducibility is whether others would draw qualitatively similar conclusions, from a reanalysis of the original study or from an independent replication, which is where disagreements about priors, multiplicity and interpretation live.
This guide is about the first kind, the computational one. It is the minimum. An analysis that cannot be reproduced from its own data cannot be meaningfully checked, and Peng (2011) makes the point that when full independent replication is not possible, reproducibility can serve as a minimum standard for judging a claim.
Reproducibility does not make a result correct. A reproducible analysis with the wrong model gives the wrong answer every time. What it does is make the analysis inspectable, so that errors can be found.
The five things you need to rerun an analysis
Think of it as a recipe that someone else, or you in two years, can follow without asking any questions.
1. The exact data that went into the model
Not the raw export, and not “the cleaned file”. The specific dataset that the final model saw, after exclusions, recoding and derived variables. If cleaning happened in a spreadsheet by hand, that step is not reproducible, and Sandve and colleagues (2013) list avoiding manual data manipulation as one of their ten rules for exactly this reason. Put the cleaning in code.
2. The script, and only one of it
The script that produced the reported numbers, run top to bottom in a fresh session. Not a script plus some commands typed into the console afterwards. A good test: restart the session, run the file, and compare every number to the manuscript.
3. The random seed
Bootstrap confidence intervals, permutation tests, multiple imputation, cross-validation and many optimisers use random numbers. Without a fixed seed, the reported estimate can change in its later decimal places on every run (how many depends on the number of resamples), and occasionally a p-value will cross 0.05 in one run and not the next. Setting a seed does not make the analysis less valid; it makes the specific draw you reported recoverable.
One subtlety: R’s random number generator has changed between versions (the default sampling algorithm changed in R 3.6.0, for example). The same seed on a different R version is not guaranteed to give the same draws, which is one reason the next item matters.
4. The software versions
Statistical packages change their defaults, fix bugs, and alter numerical algorithms between releases. An analysis run with one version of a mixed-model package can give slightly different estimates, or a convergence warning, under another. Sandve and colleagues put it plainly: archive the exact versions of all external programs used.
Two records cover this. Session information captures the R version, operating system and every attached package with its version. A lockfile pins those versions so someone else can install them.
Keep in mind that even with identical package versions, results can differ in the last few digits across operating systems or maths libraries, because floating-point arithmetic is not always performed in the same order. Agreement to many significant figures is what a reproducible analysis delivers in practice; exact bit-level identity across machines is a stronger claim that has to be checked, not assumed.
5. A record of the decisions
This is the part most often missing. Code shows what was run but not why. Why was the Welch test used instead of Student’s? Why were three outliers excluded? Why was the model adjusted for age but not site? Were other models tried first?
An audit trail records each analytic decision, the assumption checks that informed it, and any changes from the pre-specified plan. It is also what separates a reproducible analysis from a reproducible fishing expedition: if twenty models were fitted and one was reported, rerunning that one script will faithfully reproduce a misleading result.
What does not count
- “Available on request.” Requests often go unanswered years later, when the analyst has moved on.
- A description in the methods section alone. “Data were analysed using mixed models in R” does not specify the random-effects structure, the optimiser, the handling of missing data or the software version.
- Point-and-click output with no log. Many menu-driven tools can export the syntax they ran. If yours can, save it; if not, the analysis cannot be rerun without re-clicking from memory.
- A script that depends on objects in your workspace. If it only works after you have run something else first, it is not the script that produced the result.
Why reviewers care
Peer review increasingly asks for code and data, and for good reason. A reviewer with the script can check that the model in the code matches the model in the text, that exclusions were applied as described, and that the reported interval is the one the code computes. These checks catch real errors: mislabelled groups, a reversed reference level, a filter applied in the wrong order.
It also protects you. When a question arrives after publication, a complete record lets you answer it in an afternoon instead of reconstructing an analysis from fragments. And if a result does turn out to be wrong, being able to show exactly where the error entered is the difference between a correction and a credibility problem.
A short checklist before you submit
- Restart the session and run the analysis script from start to finish.
- Confirm every number in the manuscript matches the output.
- Save the seed, the session information and the package versions alongside the results.
- Write down every decision that was not pre-specified, and why it was made.
- Store the analysis dataset, the script and the output together, under version control if possible.
How Celeus handles this
Celeus fixes the seed for every analysis and gives you a reproducible package with the analysis plan, the script, the results, the exact software versions and an audit log of each step, including the assumption checks. Rerunning it under the recorded versions is designed to give the same results, and the validation report shows how Celeus results compare with independent computations.
Sources
- Goodman SN, Fanelli D, Ioannidis JPA (2016). What does research reproducibility mean? Science Translational Medicine 8(341):341ps12. https://doi.org/10.1126/scitranslmed.aaf5027
- Peng RD (2011). Reproducible research in computational science. Science 334(6060):1226-1227. https://doi.org/10.1126/science.1213847
- Sandve GK, Nekrutenko A, Taylor J, Hovig E (2013). Ten simple rules for reproducible computational research. PLoS Computational Biology 9(10):e1003285. https://doi.org/10.1371/journal.pcbi.1003285
- National Academies of Sciences, Engineering, and Medicine (2019). Reproducibility and Replicability in Science. Washington, DC: The National Academies Press. https://doi.org/10.17226/25303
Try it on your own data
Upload a dataset and Celeus suggests a method, checks its assumptions, runs it in R, and gives you the script to rerun it.
Try it in Celeus