Skip to content
Celeus

Handle missing data

Missing values only matter for the columns an analysis uses. Celeus handles them the way you choose on the configure screen and records the choice in the reproducible package.

The dataset profile shows the share of missing values for every column. The Data check panel also points out missingness patterns that look structural or informative, which are worth understanding before you choose a strategy.

On the configure screen, open Advanced settings and pick a Missing-data strategy. Celeus offers only the strategies that are valid for the chosen test, and shows the pros, cons and reporting guidance for each before you run.

  • Complete case (listwise deletion) is the default and is available for every test. Rows with a missing value in any analysis column are left out. It is simple and transparent, but it loses statistical power and is biased unless the data are missing completely at random.
  • Mean imputation and median imputation fill each missing numeric value with the column mean or median. They keep every row, but they understate uncertainty, so standard errors and p-values come out too optimistic. Use them with care and disclose them.
  • Multiple imputation fills the gaps several times from a model of the other variables and pools the results. It is valid under the weaker, more realistic assumption that data are missing at random, and its standard errors reflect the uncertainty. Celeus offers it for linear regression, logistic regression, Cox proportional-hazards regression and Weibull accelerated failure time regression.

Some methods, such as survival curves, mixed models and diagnostic accuracy, offer complete case only, because filling in values would distort exactly what they estimate.

If missing values are recorded as codes (for example -99), add them as Missing-value codes in the dataset’s codebook first, so Celeus treats them as missing instead of as real readings.

When a complete-case analysis drops many rows, the results page shows a warning flag with the percentage dropped. Where the test supports multiple imputation, the flag offers:

  • Re-run with multiple imputation, which submits the same analysis with that strategy.
  • Compare complete case with multiple imputation, which runs both and shows the two results side by side, so you can see whether your conclusion depends on the choice.

Both runs stay in the record. Choose the strategy on the merits, not on which result you prefer.

Whatever you choose, report it: the number of rows analyzed, how many were excluded or imputed, and the assumption the strategy makes. The strategy card on the configure screen gives the reporting note for each option, and the methods document in the reproducible package states what the run actually did.