Celeus

How to analyse values below the detection limit

Why replacing non-detects with half the detection limit biases results, and how Peto-Peto tests and censored regression use them correctly.

Updated .

A lab report comes back with “<0.5 µg/L” in a third of the rows. The instrument did detect something, or nothing, but it could not quantify it below 0.5. The spreadsheet now has a column that is part numbers, part inequalities, and every standard test wants numbers.

The most common fix is to replace each “<0.5” with 0.25 (half the limit), or 0, or 0.5, and carry on. It is quick, it appears in countless published papers, and it is the wrong answer. This guide explains why and what to do instead.

A non-detect is an interval, not a value

“<0.5” is a real measurement. It says the true concentration lies somewhere between zero and 0.5. Statisticians call this left-censored: the value is known to be below a threshold, but not where below it. It is the mirror image of right-censoring in time-to-event analysis, where a unit is known to have survived at least a certain time.

Substitution throws away that interval and writes in a single invented number. Helsel (2006) calls this “fabricating data”, and the charge is precise. All the non-detects below one limit get the same value, so:

  • The spread is distorted, usually shrunk. A block of identical values typically understates the variance, which makes standard errors too small and p-values too optimistic. Substituting zero, or mixing several limits, can inflate it instead, and you cannot tell in advance which you got.
  • The centre moves. Whether it moves up or down depends on the arbitrary factor you chose (0, 1/2, 1/sqrt(2), or 1 times the limit), and the estimated effect of a covariate can shift in either direction.
  • Multiple limits create false patterns. If one lab reports down to 0.5 and another only to 2.0, substituting half the limit writes 0.25 for one set of samples and 1.0 for the other. That difference is pure artefact, and a regression will happily attribute it to whatever variable is confounded with the lab.

None of these errors shrink as the sample grows. More data just makes you more confident in a biased number.

Two situations that need different tools

One detection limit, no detects below it

If every sample shares the same limit, all the non-detects are below every detected value. Their ranks are known: they tie for the lowest positions. A rank test such as the Mann-Whitney U (or Kruskal-Wallis for three or more groups) handles this exactly, with the non-detects entered as a tied block at the bottom. No substitution, no distributional assumption.

Several limits, overlapping with detected values

This is the common real-world case: samples run in different batches, at different dilutions or in different labs, each with its own reporting limit. Now a “<2.0” might be larger or smaller than a detected 1.3, so ranks are no longer determined. Here you need methods built for censored data.

Comparing groups: the Peto-Peto test

The Peto-Peto test (Peto and Peto 1972) is a generalised Wilcoxon test for censored data. It was developed for right-censored survival times, and it transfers to left-censored concentrations by a simple trick: subtract every value from a constant larger than the maximum. Low concentrations become long “times”, and a “<0.5” becomes a unit still at risk past a point, which is ordinary right-censoring. Helsel (2012) covers this flipping approach in detail for environmental data.

Worked example: a contaminant in well water

The data are synthetic. Eighty wells are sampled, 40 upstream and 40 downstream of a site. Half the samples go to a lab with a detection limit of 0.5 µg/L, half to one with a reporting limit of 2.0 µg/L. True concentrations are lognormal, twice as high downstream on the geometric scale, and lower in deeper wells.

Of the 80 samples, 26 were non-detects: 12 upstream (30%) and 14 downstream (35%). The downstream site has slightly more non-detects despite higher true concentrations, partly because its wells happen to be deeper on average in this sample (31 m against 27 m). Raw detection frequency is a poor summary when limits and covariates vary.

The Peto-Peto test on site alone gives chi-squared = 2.6 on 1 df, p = 0.11. It does not adjust for depth, and depth explains a good deal of the variation here, so the unadjusted comparison is noisy.

Regression: censored (Tobit-type) models

To adjust for covariates, fit a regression by maximum likelihood where each detected value contributes its density and each non-detect contributes the probability of falling below its own limit. Tobin (1958) introduced this idea in economics for outcomes piled up at a boundary; the environmental literature calls the same model censored regression or censored maximum likelihood. Each non-detect enters the model as an interval: anywhere below its own limit.

For concentrations, a lognormal distribution is the usual choice: it keeps values positive and handles right skew. Coefficients are then on the log scale, so exponentiating gives ratios of geometric means.

Term Ratio 95% CI
Downstream vs upstream 1.63 1.08 to 2.47
Well depth (per metre) 0.976 0.961 to 0.990

Downstream wells have an estimated 63% higher geometric mean concentration than upstream wells at the same depth. Each additional metre of depth lowers it by about 2.4%. The fitted log-scale standard deviation is 0.87.

How much does substitution actually cost?

One dataset cannot show bias; any single estimate can land anywhere. So the well scenario above was simulated 2,000 times, fitting both the censored lognormal model and ordinary least squares on log concentrations after substituting half the limit.

Quantity True value Censored model (median) LOD/2 substitution (median)
Downstream / upstream ratio 2.00 2.00 1.74
Log-scale SD 1.00 0.97 0.89

Substitution understated the site effect by roughly 13% and the spread by about 11%, at the median across repetitions, and collecting more wells would not fix either. The censored model recovered both. The size and direction of substitution bias depend on the censoring pattern and the factor chosen, which is the core problem: you cannot know in advance which way it will push your result.

Assumptions and limits

  • Censored regression assumes the chosen distribution fits the underlying values, including the part you cannot see. Check it against the detected values with a probability plot, and prefer lognormal only when every value and limit is positive.
  • All of these methods assume the limit is known for each non-detect and that whether a value falls below its limit depends only on the value, not on something else about the sample.
  • With a very high proportion of non-detects there is little information about the distribution below the limits. Treating the outcome as detected versus not detected, and analysing it as binary, is often the more honest analysis.
  • Substitution still has one legitimate use: as a sensitivity analysis, to show a reader that a conclusion does not depend on it. It should never be the primary result.

How Celeus handles this

Celeus keeps each non-detect as “below its limit” instead of inventing a value, and compares groups with a rank test or the Peto-Peto test and adjusts for covariates with censored regression, using each sample’s own limit. Substitution is never proposed; it is available only as a sensitivity analysis you ask for.

Sources

  1. Helsel DR (2006). Fabricating data: how substituting values for nondetects can ruin results, and what can be done about it. Chemosphere 65(11):2434-2439. https://doi.org/10.1016/j.chemosphere.2006.04.051
  2. Helsel DR (2012). Statistics for Censored Environmental Data Using Minitab and R, 2nd ed. Hoboken, NJ: Wiley. https://doi.org/10.1002/9781118162729
  3. Peto R, Peto J (1972). Asymptotically efficient rank invariant test procedures. Journal of the Royal Statistical Society: Series A (General) 135(2):185-207. https://doi.org/10.2307/2344317
  4. Tobin J (1958). Estimation of relationships for limited dependent variables. Econometrica 26(1):24-36. https://doi.org/10.2307/1907382

Try it on your own data

Upload a dataset and Celeus suggests a method, checks its assumptions, runs it in R, and gives you the script to rerun it.

Try it in Celeus