What is censoring?
What censoring means in survival analysis: right, left and interval censoring, independent censoring, and how Kaplan–Meier and other methods handle it.
Censoring means that a person’s event time is only partly known. We know that the event had not happened by a certain time, or that it happened before a certain time, or between two times, but not exactly when. It is the defining feature of survival data, and the reason survival analysis needs its own methods.
Why it matters
Take a study that follows people for five years. Someone who leaves after two years without having had the event has contributed two years in which we know they were event-free. Dropping them throws that information away and usually overstates the risk, because the people dropped are ones who had not had the event. Treating them as event-free for all five years invents information we don’t have.
Methods for survival data use exactly what is known: how long each person was at risk, and whether their follow-up ended with the event or with censoring.
Right censoring
The most common kind. The event has not happened by the time follow-up ends, so we know only that the event time is later than the censoring time. It happens when the study ends (administrative censoring), when someone is lost to follow-up or withdraws, or when follow-up stops for another reason.
Left censoring
The event happened before a known time, but we don’t know when. For example, the first test in a study shows that an infection is already present, but not when it began. We know only that the event time is earlier than the first observation.
Left censoring is not the same as left truncation, or delayed entry, where people join a study some time after time zero and are included only because they had not yet had the event. Truncation is handled by starting each person’s time at risk at their entry time.
Interval censoring
The event is known to have happened between two times, but not exactly when. It is typical when status is checked only at visits: someone shows no progression at one scan and has progressed at the next, so progression happened somewhere in between. For example, if a person develops glaucoma between visits to the optician, the onset is interval-censored.
Right and left censoring are special cases: for right censoring the interval has no upper end, and for left censoring it starts at time zero. Interval-censored data need methods built for them. Recording the event at the later visit makes event times too long, and survival looks better than it is. Our tutorial on survival analysis with interval censoring works through an example.
Independent (non-informative) censoring
Standard methods assume that censoring is independent of the event: people censored at a given time have the same risk of the event afterwards as people still under follow-up at that time. Kaplan–Meier assumes this within each group it is calculated for, and the Cox model given the covariates in the model. This is called independent censoring. It is often called non-informative censoring, although strictly that is a separate, technical condition: that the censoring process carries no information about the parameters of the event-time distribution, so it can be left out of the likelihood.
Administrative censoring at the end of a study usually satisfies it, unless people who joined later, and so were followed for less time, differ in risk from those who joined early. Dropout often does not. If people leave a trial because they are becoming more unwell, those who remain are healthier than those who left, and survival will be overestimated.
The assumption cannot be tested from the observed data alone, because what happens after censoring is never seen. It has to be argued from how the data were collected, and probed with sensitivity analyses. What can be checked is whether censoring is related to measured factors. If it is, including them in the model, or weighting by the inverse probability of remaining uncensored, deals with that part of the problem.
Censoring and competing events
A death from another cause is not the same as censoring. After it, the event of interest can no longer happen, whereas a censored person could still have it later. Treating competing events as censored is correct for estimating the cause-specific hazard, but not for turning it into a probability: one minus the Kaplan–Meier estimate then overstates the real risk of the event of interest. At best it estimates the risk in a hypothetical world where the competing event could not happen, and even that relies on the two events being independent, which the data cannot confirm. See what competing risks are and when censoring a competing event is right, and when it is wrong.
How methods handle it
- Kaplan–Meier counts each person in the risk set while they are under follow-up, and removes them when they are censored.
- The Cox model uses each person’s time at risk: they take part in the comparison at every event time until they are censored.
- Parametric models use the likelihood: an event contributes the density at its time, and a right-censored person contributes the probability of surviving beyond their censoring time.
- Left and interval censoring contribute the probability that the event fell in the known interval, which needs software written for it.
How Kaplan–Meier uses a censored person
Ten people are followed for up to six months. One dies at month 2, two are lost to follow-up at month 3, one dies at month 5, and the other six are alive at month 6.
- Month 2. Ten people are at risk and one dies, so the estimated survival is .
- Month 3. The two censored people leave the risk set. Nobody has died, so the estimate stays at 0.900, and seven people remain at risk.
- Month 5. Seven are at risk and one dies, so the estimated survival is .
The estimated risk of death by month 6 is therefore , about 23%. Dropping the two censored people gives 2 deaths among 8, or 25%, which overstates the risk. Counting them as alive at month 6 gives 2 deaths among 10, or 20%, which assumes they survived three months that nobody saw. Kaplan–Meier uses the three months in which they were seen, and assumes that afterwards they had the same risk as the seven still followed, which is the independent censoring assumption described above. Had the two been lost at month 2 instead, they would still count among the ten at risk for that death, because when a death and a censoring are recorded at the same time the death is taken to come first (Kaplan and Meier, 1958).
Want to learn more?
- Our survival analysis course
- What are competing risks?
- Reconstruct patient data from a published Kaplan–Meier curve
- Check out our Resources page
MethodSurvival analysis