Skip to content
Celeus

What we send to the AI and what we never send

Celeus uses an AI model to suggest methods, interpret results, draft text and answer chat. The statistics themselves are computed by Celeus’s own engine and never by the AI. This page lists what the model receives. For the identifier scan and how the guard works, see Security.

What reaches the AI:

  • Column names, types and variable labels stored in your file
  • Summary statistics, including each numeric column's minimum and maximum
  • Category names, with counts of 11 or fewer not listed
  • Aggregate results of your analyses
  • Text you write: chat messages, research questions, notes and codebook descriptions

What never reaches the AI:

  • Your uploaded file
  • Data rows or data tables

The rest of this page says exactly what each of those means.

The AI never works from your data. It works from a profile Celeus builds from it, which contains:

  • The size of the dataset - the number of rows and columns.
  • For every column: its name, its type (numeric, text, date and so on), the role Celeus inferred for it (for example continuous, binary or categorical), and how many values are missing, as a count and a percentage.
  • For numeric columns: the mean, standard deviation, median, minimum, maximum and the two quartiles, each rounded to 4 significant digits. This applies to every numeric column the AI is allowed to see, including number-coded ones.
  • For category columns with a limited number of distinct values: the category names, each with its count. Counts of 11 or fewer are not listed; the profile marks them as withheld instead. The row count, the missing-value count and the other categories’ counts are still sent, so a single withheld count can be worked out from them.
  • For text columns that look like identifiers or free text (too many distinct values to be categories) and for date columns: only the number of distinct values. Their values are never listed. A number-typed column is summarized as a numeric column even if it holds identifiers, so exclude such a column from the AI.
  • Variable labels stored in your file, such as SPSS or Stata variable labels or REDCap field labels.
  • Codebook text you write - labels, descriptions, units and value labels - only once the identifier scan of that text comes back clear. Until then it is held back, and Celeus tells you so. The column type you declare in a codebook, and counts of values the codebook marks as missing or out of range, are sent either way.
  • For a column you have declared as having detection limits: instead of the numeric summary above, the limits themselves, how many values fall below or above them, and quartiles that account for the limits. Small counts are hidden the same way, and the range of detected values is withheld when few values were detected.

When the AI suggests a method, it also sees the candidate analyses Celeus’s own advisor ranked from this profile - test names and the column names they would use.

When the AI interprets or discusses an analysis, or drafts text about it, it sees that analysis’s summary results: the test, the number of observations, test statistics, p-values, estimates, confidence intervals, effect sizes, group means by group label, post-hoc comparisons, regression coefficients by term, assumption-check outcomes, power figures, how missing data were handled, and the warning flags.

These are computed over the rows in each group, and a group can be small. Hiding small counts applies to the dataset profile, not to analysis results. Results can include category names and column names, for example “mean in the placebo group”.

Text you type is sent as you wrote it: your research question, the research context you give a project or dataset, the study details and instructions you give when drafting a report, your chat messages, and the names of your datasets (by default, the name of the file you uploaded). Keep identifiers and patient details out of it.

The identifier scan excludes every column it flags from the AI as soon as you upload. You can also exclude any column yourself, or include a flagged column again after review.

An excluded column is left out of the profile the AI works from, and the chat assistant cannot run or propose an analysis on it. At present, a new column you derive from an excluded one (for example by binning it) is not excluded automatically - exclude it too.

The chat assistant is told which columns the identifier scan flagged, by name, so that it can warn you about them. It is not given their values.

If you choose to run an analysis on an excluded column yourself, that analysis’s results are what the AI interpretation reads - so it sees the column’s name and its aggregate results, never its data rows. If the excluded column is used as a group, its category names also reach the AI, as the labels of the group results and post-hoc comparisons.

A workspace owner can choose to keep identified data. In that case the AI features work only from a de-identified version of each dataset, and ask you to create one before they run.

  • Your uploaded file.
  • Data rows or data tables. Every path from your data to the model passes a guard that refuses a data table, and each run’s reproducible package includes an audit log with the prompts sent to the model for that run.

The AI provider and its terms are listed in the privacy policy.