Celeus

Statistics you can defend in front of a reviewer.

Upload a dataset from any field. Celeus suggests the right statistical method and tells you why, runs it deterministically in R, and interprets the result - so the number in your paper is the number your data actually supports.

Paired t-test

Data: Student (1908), Biometrika 6(1):1-25

The same p-value computed by Celeus and by an independent R computation
Celeusp = 0.00283289
Independent R computationp = 0.00283289matches Celeus

Same number, computed two independent ways. Every run ships its exact R script, seed and package versions. See all checks

See a run from upload to package

A synthetic dataset is uploaded, its identifier columns are replaced with pseudonymous codes, a method is suggested and run in R, and the reproducible package is downloaded.
Transcript

Celeus: from a dataset to a reproducible result.

We drop in a synthetic dataset: forty rows, five columns.

Celeus scans every column first. Two of them, patient name and MRN, look like direct identifiers.

The review recommends replacing both columns with coded stand-ins. We apply the recommendation as proposed.

That creates a new, de-identified version. The original upload is kept, and a fresh scan finds no identifiers.

Next, Celeus ranks suitable analyses from the de-identified profile, without calling an AI model. To compare the outcome between the two arms, it recommends Welch's t-test.

The test, outcome and group come pre-filled. We add a research question, which is saved with the results.

The analysis runs in R, the statistical programming language, with a fixed seed. The wait is shortened in this video.

The result: the treatment group's mean is 6.7 points higher, with a 95 percent confidence interval of 4.25 to 9.15, and p below 0.001.

Below it is a plain-language summary written without AI, and a warning: with forty rows, only large effects can be detected reliably.

An AI interpretation is added on top, clearly labelled. Only de-identified metadata reaches the model, and the computed numbers remain the ground truth.

Here is the exact R code that reproduces this result, with its seed and the engine version.

We check the results and confirm them right here, next to the download. The confirmation is saved with the run, and the reproducible package unlocks: the data, the script and the pinned package versions, together.

How it works

  1. 1

    Upload

    CSV, TSV or Excel on every plan; SPSS, Stata and SAS on paid plans. Or add one of the public sample datasets to try it first.

  2. 2

    Method chosen and explained

    Celeus suggests a method for your question and your data, grades the alternatives it considered, and checks the assumptions. If you pick something else, that choice is recorded.

  3. 3

    Deterministic run

    Seeded, in R, on pinned package versions. Run it again on the same engine and you get the same numbers.

  4. 4

    Interpretation and flags

    A plain-language reading of the result, with failed assumptions and other problems flagged where you will see them.

  5. 5

    Reproducible package

    The plan, the exact R script, the seed, session info, the results and the audit log, in one download.

Works with any research data

The subject of your data does not matter. What matters is its shape and your question.

How to choose a method: the guides

Example research questions and a method that fits each
Question
Clinical researchDo patients on the new treatment stay event-free longer?Kaplan-Meier and Cox regression
PsychologyDid scores change after the intervention?Paired t-test or Wilcoxon signed-rank
EcologyDo species counts differ between sites?Poisson or negative-binomial regression
EducationDo outcomes differ across three schools?One-way ANOVA or Kruskal-Wallis
Public healthIs an exposure associated with the outcome?Chi-squared test or logistic regression
Lab scienceDo two assays agree?Bland-Altman agreement

What reaches the AI, and what never does

Every path from your data to the AI model passes a guard that refuses data tables. Every upload is scanned for identifiers, and columns flagged as identifiers, free text or dates are excluded from the AI by default. You decide what happens to each flagged column.

Reaches the AI

  • Column names, types and variable labels stored in your file
  • Summary statistics, including each numeric column's minimum and maximum
  • Category names, with counts of 11 or fewer not listed
  • Aggregate results of your analyses
  • Text you write: chat messages, research questions, notes and codebook descriptions

Never reaches the AI

  • Your uploaded file
  • Data rows or data tables

Text you write is sent as you wrote it, so keep identifiers out of it. That matters most for patient records, and just as much for survey responses, student records or anything else collected under a confidentiality promise.How Celeus handles your data

More than a test runner

Chat with your data

Ask about your dataset and results in plain language. The assistant works from the de-identified profile and aggregate results, never from data rows.

Draft methods and results

Turn a run into Methods, Results and a structured abstract. Every number comes from the run itself, not from the model. Export to Word, LaTeX or Markdown.

Share with collaborators

Send a time-limited, read-only link to a report snapshot, or give a named researcher in another workspace view or run access to a project.

Sensitivity analysis

Compare the complete-case result with multiple imputation for the same analysis, and see whether the two agree.

Censored and below-detection data

Values below a detection limit are modelled as censored (Peto-Peto test, Tobit-style regression), never replaced with a guessed value.

REDCap data dictionaries

Import a REDCap data dictionary to label your variables and categories. You review each change before it is applied.

36statistical methods, checked in the open

Every plan gets every method. On well-known public datasets, each result is checked against an independent computation, and a subset against values published in the literature. The report is generated, not written by hand.

Read the validation report

Comparing groups

  • Welch's t-test
  • Student's t-test
  • Paired t-test
  • Mann-Whitney U test
  • Wilcoxon signed-rank test
  • One-way ANOVA
  • Welch's ANOVA
  • Kruskal-Wallis test
  • Repeated-measures ANOVA (Greenhouse-Geisser)
  • ANCOVA (analysis of covariance)
  • Two-way ANOVA (factorial)

Association and correlation

  • Chi-squared test of association
  • Fisher's exact test
  • Pearson correlation
  • Spearman correlation
  • McNemar's test

Regression

  • Linear regression
  • Logistic regression
  • Firth penalised logistic regression
  • Poisson regression
  • Negative-binomial regression
  • Ordinal logistic regression
  • Multinomial logistic regression

Mixed-effects models

  • Linear mixed-effects model (random intercept)
  • Mixed-effects logistic regression (random intercept)
  • Mixed-effects Poisson regression (random intercept)

Survival and censored data

  • Peto-Peto generalized Wilcoxon test
  • Censored regression (Tobit-style)
  • Kaplan-Meier (log-rank test)
  • Cox proportional-hazards regression
  • Weibull accelerated failure time regression

Diagnostic accuracy and agreement

  • ROC curve / AUC
  • Diagnostic accuracy (2x2)
  • Bland-Altman agreement
  • Compare two ROC curves (DeLong)

Descriptive

  • Descriptive table (Table 1)

Pricing

One researcher or a whole lab - the per-person price stays the same.

Free

See it work on your own data.

$0/mo

Start free
  • 10 analysis runs a month
  • ~15 AI actions a month
  • 100 MB storage, 1 person
  • All 36 statistical methods
  • CSV, TSV and Excel import
  • Identifier scan and de-identification
  • Reproducible package on every run

Pro

One researcher's full toolkit.

$19/mo

for one researcher
Start Pro
  • 500 analysis runs a month
  • ~300 AI actions a month
  • 5 GB storage
  • Shareable report links
  • Multiple reports per project
  • SPSS, Stata and SAS import
  • Advanced data preparation
  • Assistant without a dataset
  • Share a project with a collaborator
  • Everything in Free

Built for labs

Team

A lab. Same price per person, plus the lab layer.

$19/mo per seat

per seat, minimum 3 seats, self-serve up to 25
Start a team
  • Per seat: 500 runs and ~150 AI actions a month, 2.5 GB storage, pooled across the lab
  • Shared workspace - the lab owns the data
  • Admin-managed membership and billing
  • Everything in Pro

Enterprise

The org wrapper around Team.

Contact us

for more than 25 seats, custom pricing
Contact us
  • Everything in Team
  • Need more than 25 seats or custom terms? Contact us.

Need more tokens or runs without changing plans? Top-up packs: AI pack $15 for 1M tokens, Analysis pack $12 for 200 runs.

Out of storage? A storage top-up adds 25 GB added permanently, on any plan - no plan change and no seat change needed. Buy it from Settings > Billing; the price is shown before you pay.

Every plan above starts by signing in. If that does not let you through, contact support and we will follow up by email.

Celeus is for research questions on research data. It does not make diagnosis or treatment decisions about individual patients.