What is survival analysis?
What survival analysis is, why censoring needs its own methods, the survival and hazard functions, a worked Kaplan–Meier example, main methods and pitfalls.
Survival analysis is the set of statistical methods for analysing the time until an event happens: a death, a relapse, a hospital admission, a machine failing. The outcome has two parts:
- the time each person was followed, and
- whether the event happened in that time (yes or no).
Why it needs its own methods
By the end of follow-up, some people will not have had the event, because the study ended or because they were lost to follow-up or withdrew. We know they were event-free until their follow-up ended, but not how much longer they would have stayed so. This is called censoring, and it is what makes survival data different. Leaving censored people out biases the answer, and so does treating the end of their follow-up as the event. A simple yes-or-no analysis of whether the event happened by a fixed time has no outcome for people whose follow-up ended event-free before that time, so it must drop them or guess. Survival methods use everything each person contributes: how long they were at risk, and whether their follow-up ended in the event.
Every analysis needs clear definitions of the time origin (such as diagnosis or randomisation, or birth when the timescale is attained age), of when each person's follow-up starts, which can be later, and of what counts as the event.
The key quantities
- The survival function : the probability of still being event-free at time . It starts at 1 and never rises.
- The hazard function : the rate at which the event happens at time among those still at risk. It can rise, fall, or do both over time.
- The cumulative hazard : the hazard added up from the start. For an event time measured continuously, it ties the other two together: .
- Summaries: the median survival time, survival at a fixed time such as five years, and the restricted mean survival time, the average time event-free up to a chosen horizon.
The main methods
- Kaplan–Meier estimates the survival function without assuming a shape for it, and the log-rank test compares survival between groups. The test is most powerful when hazards are proportional, and can miss curves that cross.
- The Cox model relates the hazard to covariates as hazard ratios, assumed constant over time, without specifying the baseline hazard. It is the most widely used regression model for survival data; see what is the Cox model? and the proportional hazards assumption.
- Parametric models, such as the exponential, Weibull and Gompertz, specify the shape of the hazard. Flexible parametric models use a restricted cubic spline, usually for the log cumulative hazard, so the shape can follow the data closely. Both give smooth estimates of survival, the hazard and other quantities for any covariate pattern, and a hazard that can be carried beyond follow-up, which health economic models usually need. Data from within follow-up cannot confirm an extrapolation, so it needs external evidence and sensitivity analyses to the choice of model.
A worked Kaplan–Meier calculation
Kaplan–Meier multiplies together, at each event time, the proportion of those still at risk who did not have the event. Take ten hypothetical patients followed for up to 12 months. Three have the event, at 2, 4 and 7 months. Three are censored, at 3, 5 and 9 months, and four are still event-free at 12 months.
- At 2 months, 10 are at risk and 1 has the event, so .
- The patient censored at 3 months leaves 8 at risk at 4 months, so .
- The patient censored at 5 months leaves 6 at risk at 7 months, so .
- The censoring at 9 months and the end of follow-up at 12 cause no step, so the estimate of survival to 12 months stays at 0.66.
Two shortcuts give other answers. Dropping the three patients censored before 12 months leaves 4 event-free out of 7, or 0.57, and understates survival, because those dropped had not had the event when last seen. Counting them as event-free to 12 months gives 7 out of 10, or 0.70, which assumes none of them had the event after leaving. Kaplan–Meier uses each censored patient for as long as they were followed, and is valid when those censored have the same risk afterwards as those still followed.
When there is more than one kind of event
The standard methods assume a single event. When several can happen, the analysis has to say how they relate:
- Competing risks: one of several events happens first and prevents the others, as with deaths from different causes.
- Multi-state models: people move between states, such as well, relapsed and dead, with a hazard for each move.
- Recurrent events: the same event can happen more than once, such as repeated hospital admissions.
- Joint models: a repeatedly measured biomarker and the time to an event are modelled together.
Common pitfalls
- Grouping people by something that only happens during follow-up, which causes immortal time bias.
- Treating a competing event as censoring when the question is about risk; see censoring a competing event.
- Censoring that is related to the event, for example when the sickest patients are the ones lost to follow-up.
- Reading a single hazard ratio as constant when hazards are not proportional.
Where it is used
Survival analysis is used wherever the timing of an event matters: overall and progression-free survival in clinical trials, time to readmission in health services research, the reliability of components in engineering, and the length of unemployment in economics. It is also called time-to-event analysis, reliability analysis in engineering, duration analysis in economics and event history analysis in sociology.
In R, the survival package provides Kaplan–Meier estimates, the log-rank test, the Cox model and parametric models; in Stata, the st commands do the same. Our merlin package fits parametric and flexible parametric survival models, and extends them to competing risks, multi-state and joint models.
Want to learn more?
MethodSurvival analysis