Kaplan-Meier and Cox regression: analysing time-to-event data
How censoring works, what Kaplan-Meier curves and the log-rank test tell you, how to read a Cox hazard ratio, and when a Weibull AFT model fits better.
Updated .
Time-to-event data turn up far outside medicine. An engineer runs pumps until they fail, an ecologist waits for seedlings to flower, an HR analyst tracks how long new hires stay. The name “survival analysis” is historical; the methods apply whenever the outcome is the time until something happens and you do not get to see that time for every unit.
Why you cannot just average the times
Suppose a reliability test runs for 2,000 hours and some pumps have not failed when it stops. You know each of those pumps lasted at least 2,000 hours, but not how long. That observation is right-censored. Other pumps may be pulled from the test early for reasons that have nothing to do with wear, such as a rig being needed elsewhere. They are censored too, at the time they were removed.
Dropping censored units throws away the longest-lived ones, and treating the censoring time as a failure time records failures that never happened. Both bias times downward. Every method below uses a censored unit for exactly what it tells you (it was still at risk up to that time) and nothing more. When you write up, give the number of events per group, not just units: events drive precision.
The assumption that carries everything
All of these methods assume non-informative censoring: conditional on the covariates in the model, a unit that is censored at time t has the same future failure risk as one that stays under observation. A pump pulled because a technician heard it grinding violates this, because the censoring is a sign of imminent failure. So does a seedling that dies before it can flower, which is a competing event rather than censoring and needs competing-risks methods. No test in your output checks this assumption. It is a question about how the data were collected, and it belongs in your methods section.
A worked example on synthetic data
The data are simulated: 120 pumps, half with standard bearings and half with ceramic bearings, each with an operating temperature. The test stops at 2,000 hours and some pumps are pulled early at random.
Of the 120 pumps, 68 failed during observation (41 standard, 27 ceramic) and 52 were censored.
Kaplan-Meier curves and median survival
The Kaplan-Meier estimator steps down at each observed failure, multiplying the running estimate by the fraction of units still at risk that did not fail at that moment. Censored units leave the risk set without causing a step. The result is an estimate of the survival function, the probability of lasting beyond each time.
The most useful single summary read off the curve is the median survival time, the time by which half the units are estimated to have failed:
| Bearing | Pumps | Failures | Median (h) | 95% CI (h) |
|---|---|---|---|---|
| Standard | 60 | 41 | 956 | 747 to 1308 |
| Ceramic | 60 | 27 | 1732 | 1403 to not reached |
“Not reached” is common and is not an error. The upper confidence limit for the ceramic group cannot be estimated because the curve’s upper band never falls to 0.5 within the 2,000 hours of follow-up. Resist reporting a mean survival time from these data; with heavy censoring the mean depends almost entirely on how the tail is extrapolated.
The log-rank test
The log-rank test compares observed failures in each group with the number expected if the groups shared one survival curve. Here it gives chi-squared = 10.3 on 1 degree of freedom, p = 0.001: 41 failures were observed in the standard group against 28.1 expected.
Two things to keep in mind. The log-rank test is most powerful when one group’s hazard is a constant multiple of the other’s, and it can miss curves that cross. And it gives you a p-value, not a size of effect. For that, you need a model.
Cox regression and the hazard ratio
The Cox proportional hazards model says each unit’s hazard (its instantaneous failure rate, given it has survived so far) is a baseline hazard multiplied by exp(b1x1 + b2x2 + …). The baseline hazard is left completely unspecified, which is why the model is called semi-parametric. The exponentiated coefficients are hazard ratios.
From the pump model:
| Term | Hazard ratio | 95% CI | p |
|---|---|---|---|
| Ceramic vs standard | 0.42 | 0.26 to 0.69 | < 0.001 |
| Temperature (per 1 °C) | 1.04 | 1.01 to 1.07 | 0.011 |
Read the first row as: at any given time, among pumps still running, a ceramic-bearing pump fails at an estimated 42% of the rate of a standard one at the same operating temperature. The second row says each additional degree raises the failure rate by about 4%, so a pump running 10 °C hotter has an estimated hazard 1.04^10, roughly 1.48 times higher.
What a hazard ratio is not: it is not a ratio of median times, and it is not the ratio of the proportions failed by some fixed date. Researchers often write “ceramic pumps were 58% less likely to fail”, which drops the time dimension the model is built on. “A 58% lower hazard of failure” is the accurate phrasing.
Checking proportional hazards
A single hazard ratio only summarises the data if the ratio is roughly constant over time. If ceramic bearings help early but not late, one number averages over that and misleads.
The standard check uses scaled Schoenfeld residuals. If proportional hazards holds, those residuals for a covariate show no trend against time. The test of Grambsch and Therneau (1994) formalises this, per covariate and globally:
| Term | Chi-squared | df | p |
|---|---|---|---|
| Bearing | 1.16 | 1 | 0.28 |
| Temperature | 0.06 | 1 | 0.80 |
| Global | 1.16 | 2 | 0.56 |
No evidence against proportional hazards here, which is expected because the data were simulated from a Weibull model, where hazards are proportional by construction. Two cautions for real data. A non-significant test is absence of evidence, not proof, especially with few events. And with thousands of events a trivial departure becomes significant. Plot the scaled Schoenfeld residuals against time and judge whether the smoothed line drifts meaningfully, not just whether p crosses 0.05.
When the assumption fails, the usual remedies are stratifying on the offending variable, letting its effect change with time, reporting results separately for time windows, or switching to a model that does not rely on proportional hazards at all.
Accelerated failure time models
An accelerated failure time (AFT) model regresses log time directly. Its exponentiated coefficients are time ratios: how much a covariate stretches or shrinks the time to failure. A time ratio above 1 means longer survival, the opposite direction from a hazard ratio. Wei (1992) makes the case that this is often the more natural quantity to report, because it is on the scale people actually think in.
The Weibull AFT fit to the same pumps:
| Term | Time ratio | 95% CI |
|---|---|---|
| Ceramic vs standard | 1.66 | 1.24 to 2.22 |
| Temperature (per 1 °C) | 0.977 | 0.961 to 0.994 |
Ceramic-bearing pumps are estimated to last about 1.66 times as long. The fitted Weibull scale is 0.577, which corresponds to a shape of about 1.73: a hazard that rises with age, as wear failures should.
The Weibull is the one distribution that is both a proportional hazards and an AFT model, so the two fits should agree. They do: converting the time ratio gives a hazard ratio of 1.66^(-1/0.577), about 0.42, the same as the Cox estimate.
When to reach for an AFT model instead of Cox:
- The proportional hazards check fails, and a parametric shape fits the data.
- Your audience needs “how much longer” rather than “how much riskier”.
- You have reason to believe the failure mechanism has a known parametric form (Weibull is standard in reliability engineering).
The price is an extra assumption: the Weibull shape has to be roughly right. Compare the fitted curves with the Kaplan-Meier curves before trusting it.
How Celeus handles this
Given a time column and an event indicator, Celeus gives you Kaplan-Meier median survival and the log-rank test, Cox hazard ratios with 95% confidence intervals and a check of proportional hazards, and a Weibull AFT model as the alternative. Every result comes with a reproducible package, and our validation report compares results with independent computations.
Sources
- Kaplan EL, Meier P (1958). Nonparametric estimation from incomplete observations. Journal of the American Statistical Association 53(282):457-481. https://doi.org/10.1080/01621459.1958.10501452
- Cox DR (1972). Regression models and life-tables. Journal of the Royal Statistical Society: Series B (Methodological) 34(2):187-202. https://doi.org/10.1111/j.2517-6161.1972.tb00899.x
- Grambsch PM, Therneau TM (1994). Proportional hazards tests and diagnostics based on weighted residuals. Biometrika 81(3):515-526. https://doi.org/10.1093/biomet/81.3.515
- Wei LJ (1992). The accelerated failure time model: a useful alternative to the Cox regression model in survival analysis. Statistics in Medicine 11(14-15):1871-1879. https://doi.org/10.1002/sim.4780111409
Try it on your own data
Upload a dataset and Celeus suggests a method, checks its assumptions, runs it in R, and gives you the script to rerun it.
Try it in Celeus