What is the Cox model?
What the Cox proportional hazards model estimates, how to read a hazard ratio, how it is fitted, its assumptions, and when a parametric model is better.
The Cox model, also known as the Cox proportional hazards model, is a regression model for survival data. It estimates how covariates multiply the hazard of an event, reported as hazard ratios, while leaving the shape of the hazard over time unspecified. It was developed by the British statistician David Cox (later Sir David) and published in 1972 (Cox, 1972).
Not having to choose the shape of the baseline hazard is largely why the model became so popular.
The Cox model estimates the relationship between a set of covariates and the hazard function: the instantaneous rate at which the event occurs at a given time, among those who have not yet had it. It is a rate per unit of time, not a probability, so its value depends on the time unit and can be larger than one.
The Cox model is a semi-parametric model, meaning that it makes some assumptions about the underlying distribution of the data but does not require complete specification of the distribution.
Instead, it assumes that the hazard is the product of two parts: a baseline hazard, which is the hazard over time for someone whose covariates are all zero, and a factor that multiplies it according to each individual’s covariates.
The model for the hazard of an individual with covariates is:
where is the baseline hazard, and the linear predictor, which has no intercept because absorbs it.
Reading a hazard ratio
Each coefficient is a log hazard ratio, and is the hazard ratio: the factor by which the hazard is multiplied for a one-unit increase in the covariate, with the others held fixed. A hazard ratio of 0.7 for a treatment means that the treated group’s event rate is 30% lower than the comparison group’s, among people with the same values of the other covariates, at every point in follow-up if hazards are proportional.
A hazard ratio is not a ratio of probabilities, and it does not say how much longer anyone lives. Those answers need the baseline hazard as well, or a different summary, such as survival at a chosen time or restricted mean survival time.
How it is estimated
Cox’s insight was the partial likelihood. At each event time it compares the person who had the event with everyone still at risk, and the baseline hazard cancels out, so the coefficients can be estimated without ever specifying .
In plain terms, suppose that when one death occurs, three people are still at risk: two untreated, and one treated person whose hazard is half of theirs, a hazard ratio of 0.5. Whatever the baseline hazard is at that moment, the probability that the death was the treated person’s is . The partial likelihood multiplies together one such probability for each event, for the person who actually had it, and the estimates are the coefficients that make this product as large as possible. With covariates fixed at baseline, only the order of events and censorings enters, so the estimated hazard ratios would not change if all the times were stretched or squeezed in a way that kept that order.
When several people have the event at the same recorded time, the partial likelihood needs a way to handle these ties. Exact methods exist but can be slow, so Breslow’s and Efron’s approximations are the usual ones. Efron’s is generally the more accurate when ties are common; it is the default in R’s coxph, while Stata’s stcox uses Breslow’s unless you add the efron option.
The baseline cumulative hazard can be estimated after fitting, as a step function that jumps at each event time (the Breslow estimator). Predictions of survival from a Cox model rely on it, and are only available up to the end of follow-up.
Assumptions
- Proportional hazards: each covariate’s effect on the hazard is constant over time. What it means, and how to check it.
- Functional form: a continuous covariate enters as a straight line on the log-hazard scale unless you say otherwise. That is an assumption to check, for example with splines.
- Independent censoring: given the covariates, people who are censored have the same risk afterwards as those still followed. What censoring is.
- Independent individuals: one person’s events tell you nothing about another’s. Clustered or repeated events need robust standard errors or a frailty term.
Stratified Cox models
When a categorical covariate does not have proportional effects, or when groups such as centres differ in ways that are hard to model, the model can be stratified on it. Each stratum then has its own baseline hazard, and the coefficients of the other covariates are assumed to be the same in every stratum unless you add interactions. The stratifying variable gets no hazard ratio of its own, so stratify only on a variable whose effect you do not need to estimate. In Stata this is the strata() option of stcox, and in R a strata() term in coxph.
When a parametric model is the better choice
Because the Cox model leaves the baseline hazard unspecified, it says nothing about survival beyond the end of follow-up, and its estimates near the end rest on few people. A parametric model is more natural when you need to predict or extrapolate, to report absolute measures such as survival at a time or life expectancy, or to model a hazard that changes smoothly over time.
Any extrapolation still rests on assumptions the data cannot check, so the choice of model needs to be justified with clinical and external evidence, as NICE DSU TSD 14 and TSD 21 set out.
Flexible parametric (Royston–Parmar) models keep much of the Cox model’s flexibility by modelling the baseline log cumulative hazard as a restricted cubic spline of log time, so that it is smooth and can be predicted from. Beyond the last knot the spline is a straight line in log time, so an extrapolation from it follows a Weibull shape. When hazards are proportional they give hazard ratios very close to the Cox model’s. Our course on flexible parametric survival analysis covers them.
Want to learn more?
- Our survival analysis course
- What is non-collapsibility?
- Which model do I need? An interactive guide
- Check out our Resources page
MethodSurvival analysis