Logistic regression: reading odds ratios and avoiding separation
How to fit and read a logistic regression for a yes/no outcome: odds ratios with confidence intervals, events per variable, and what to do about separation.
Updated .
Logistic regression is the default model when the outcome is binary (passed or failed, present or absent, responded or not) and you want to relate it to one or more predictors at once. A chi-square or Fisher’s exact test answers “is there an association” for a single categorical predictor; logistic regression answers “how much do the odds change with this predictor, holding the others fixed”, and it accepts continuous predictors without binning them.
This guide starts from a fitted model, because most confusion about logistic regression is about reading its output, and then covers the two ways small datasets break it.
An education example
Synthetic data: 200 students, weekly study hours (0 to 20) and whether they attended tutoring, with pass or fail on a final exam as the outcome.
92 students passed and 108 failed. The exponentiated coefficients:
| Odds ratio | 95% CI | |
|---|---|---|
| hours (per 1 hour) | 1.18 | 1.12 to 1.26 |
| tutoring (yes vs no) | 1.84 | 0.99 to 3.47 |
Reading an odds ratio
The model is linear on the log-odds scale: log(p / (1 - p)) = b0 + b1 x hours + b2 x tutoring. Exponentiating a coefficient gives an odds ratio, the multiplicative change in the odds of passing for a one-unit increase in that predictor with the other predictors held constant.
- Hours: 1.18 per hour. Each extra weekly hour multiplies the odds of passing by about 1.18. Effects compound on the odds scale, so five extra hours multiply the odds by the fifth power of the per-hour odds ratio, about 2.31 from the unrounded estimate (95% CI 1.73 to 3.16). Choose a unit that means something in your field and report the odds ratio per that unit; “per hour” and “per 5 hours” describe the same fitted model.
- Tutoring: 1.84. Tutored students had about 1.84 times the odds of passing, adjusted for hours. The interval runs from 0.99 to 3.47 and the Wald p-value is 0.056. That does not show that tutoring has no effect. The data are compatible with anything from no effect to roughly three and a half times the odds. (The simulation used a true odds ratio of e^0.8 = 2.23, which sits comfortably inside the interval. This is what an underpowered comparison looks like.)
Odds are not probabilities
An odds ratio of 2.31 does not mean students are 2.31 times as likely to pass. Odds ratios approximate risk ratios only when the outcome is rare, and passing here is common. To communicate on the probability scale, compute predicted probabilities at meaningful values.
For example, an untutored student who studies 5 hours has a modelled pass probability of 0.21; at 15 hours it is 0.59. That is a ratio of about 2.8 in probability and a difference of 38 percentage points, both of which describe the same effect differently from the odds ratio. Readers generally understand probabilities better, so report both where you can.
Wald or profile intervals
The intervals in the table above are profile-likelihood intervals, which respect the asymmetry of the likelihood. The quick Wald interval (estimate ± 1.96 SE, then exponentiated) is close here because the sample is decent, but it degrades in small samples and fails completely under separation (below). Say which one you report.
Assumptions worth checking
- Independent observations. Students nested in classrooms, or repeated measurements on the same person, need a mixed or clustered model.
- Linearity in the logit for continuous predictors. Hours is assumed to act linearly on the log-odds. Check with a spline or by adding a squared term if you have reason to doubt it.
- No perfect prediction (separation), covered below.
- Enough events for the number of parameters.
There is no normality or equal-variance assumption on the outcome.
Events per variable
A logistic model’s information comes mainly from the rarer outcome class, not from the total sample size. The usual rule of thumb is at least 10 events per estimated parameter (EPV), from a simulation study by Peduzzi et al. (1996). They found no major problems at EPV of 10 or more, whereas below 10 the coefficients were biased in both directions, variance estimates were unreliable and confidence intervals lost coverage.
Two refinements matter in practice:
- Count the minority class, and count parameters, not variables. With 92 passes and 108 failures, the limiting count is 92. A four-level categorical predictor costs three parameters, not one. Our model has two parameters beyond the intercept, so EPV = 92 / 2 = 46, which is comfortable.
- Ten is a heuristic, not a law. Vittinghoff and McCulloch (2007) found a range of situations where coverage and bias were acceptable below 10 EPV, and other factors that mattered as much. A low EPV is a reason for caution, a smaller model, or penalised estimation, not an automatic stop.
Separation, and why a huge effect can come out “non-significant”
Separation happens when a predictor, or a combination of predictors, perfectly predicts the outcome in your sample. The maximum-likelihood estimate then does not exist: the likelihood keeps improving as the coefficient heads toward infinity.
A synthetic pilot with 24 students, where everyone who passed a placement pre-test also passed the exam.
Among the 12 pre-test non-passers, 2 passed the exam; among the 12 pre-test passers, all 12 did. The ordinary fit reports a coefficient of 22.2 with a standard error of 5118 and a Wald p-value of 0.997. On this run R printed no warning at all, so a reader skimming for warnings would miss it. The signs are the absurd standard error and fitted probabilities that equal 1 to eight decimal places. The Wald test is useless here, which is why the p-value says “nothing going on” about what is obviously a strong association. The likelihood-ratio test, which compares deviances instead of dividing by the broken standard error, gives p = 3 x 10⁻⁶.
The odds ratio itself is still infinite, so there is nothing sensible to report from this fit. Dropping the predictor or collapsing categories just to make the problem go away throws away the strongest signal in the data.
Firth’s penalised likelihood
Firth (1993) proposed a modification of the score equations that removes the leading term of the small-sample bias of maximum-likelihood estimates. For logistic regression, it is equivalent to penalising the likelihood by the Jeffreys prior. Heinze and Schemper (2002) showed that this penalty always yields finite estimates under separation, and recommended penalised likelihood-ratio tests and profile penalised-likelihood confidence intervals over Wald ones.
Applied to the pilot data, the Firth odds ratio is 105, with a profile penalised-likelihood 95% CI of 8.8 to about 15,300, and a penalised likelihood-ratio p = 0.00002. For a single binary predictor, this estimate equals the odds ratio you get by adding 0.5 to every cell: (12.5 x 10.5) / (0.5 x 2.5) = 105. The width of the interval is the honest part. The data show a strong association and say almost nothing about how strong beyond that.
With plenty of events and no separation, Firth and ordinary estimates are close. Firth is still worth considering when EPV is low, because it reduces the small-sample bias that makes odds ratios too extreme in exactly those datasets.
What to report
State the outcome and which level counts as the event, the number of events and non-events, every predictor with its unit or reference level, odds ratios with 95% confidence intervals and the interval method, and, for at least the main predictor, a predicted-probability contrast. If the model separated or EPV was low, say so and name the remedy you used. If you fit many candidate models or test many predictors, read the multiple comparisons guide before interpreting the p-values.
How Celeus handles this
Celeus reports each predictor’s odds ratio with a 95% confidence interval, names which outcome level it treated as the event, and checks events per variable for you. It flags separation even when R prints no warning, and offers Firth penalised logistic regression as a re-run.
Sources
- Peduzzi P, Concato J, Kemper E, Holford TR, Feinstein AR (1996). A simulation study of the number of events per variable in logistic regression analysis. Journal of Clinical Epidemiology 49(12):1373-1379. https://doi.org/10.1016/S0895-4356(96)00236-3
- Vittinghoff E, McCulloch CE (2007). Relaxing the rule of ten events per variable in logistic and Cox regression. American Journal of Epidemiology 165(6):710-718. https://doi.org/10.1093/aje/kwk052
- Firth D (1993). Bias reduction of maximum likelihood estimates. Biometrika 80(1):27-38. https://doi.org/10.1093/biomet/80.1.27
- Heinze G, Schemper M (2002). A solution to the problem of separation in logistic regression. Statistics in Medicine 21(16):2409-2419. https://doi.org/10.1002/sim.1047
Try it on your own data
Upload a dataset and Celeus suggests a method, checks its assumptions, runs it in R, and gives you the script to rerun it.
Try it in Celeus