Skip to content
Celeus

Determinism and reproducibility

A statistical result is only as useful as your ability to show where it came from. Celeus is built so that every number on a results page can be traced to an exact dataset version, method, setting and software version, and produced again.

  • Runs in R, on pinned package versions. The engine runs a fixed, recorded set of R packages. The package versions behind a run are listed in its reproducible package, together with the lockfile they come from.
  • Uses a fixed seed. Any step that involves randomness, such as multiple imputation or a resampling method, starts from a recorded seed, shown on the results page.
  • Uses an exact data version. A dataset is never changed in place. Cleaning and de-identifying create new versions, and each run records the version it used.
  • Writes the exact script. The package contains an R script that re-runs the analysis with the recorded settings.
  • Keeps an audit log. Every step of the run is recorded, including each prompt sent to the AI.

Run the same analysis again on the same engine, with the same data version and settings, and you get the same numbers. Celeus’s automated checks test this before a change to the engine ships: a corpus of analyses is re-run from their packages and must reproduce every recorded number exactly, and the same analysis run in two separate processes must produce an identical package. That shows a result is repeatable, not that it is correct. Correctness is what the validation report covers, by comparing Celeus’s numbers with published reference values and independent computations.

  • AI-written text. The interpretation and manuscript draft are written by an AI model and can be worded differently each time. They are marked as AI-written, they never supply the numbers, and the prompts that produced them are kept in the audit log.
  • A different software environment. The promise is about the same engine and package versions. Re-running the script on another computer with other versions of R or its packages can give slightly different results, which is why the package records the exact versions.
  • A different dataset version or different settings. Changing the data, the population filter, the missing-data strategy or the test produces a new result. Both runs are kept.

A reviewer who asks “how exactly did you get this?” can be given the package. It answers the question completely: the data version, the method and why it was chosen, every setting, the seed, the software, and the checks and flags that came with the result.