stexcess Excess hazard
Flexible parametric excess hazard models in Stata, with the expected rate modelled from a reference population in the same data rather than read from a life table. Built on merlin.
What stexcess does
stexcess fits an excess hazard model in which the expected (reference) hazard is not read from a life table but modelled, jointly with the excess hazard, from a reference population in the same data. It is built on merlin, and each model it takes accepts merlin’s extended linear predictor syntax.
After stset, an indicator variable coded 0 for the reference population and 1 for the population with an excess hazard divides the data. You give stexcess two models, (reference_model)(excess_model). Everyone has the reference hazard; those with the indicator equal to 1 have an excess hazard as well.
After estimation, predict gives survival, hazard, cumulative hazard and cumulative incidence functions, restricted mean survival time and time lost, and differences and ratios of these between covariate patterns, with confidence intervals.
What the command fits
Everything below is what the package’s help file documents.
A reference model and an excess model
Each has its own covariates and its own baseline. Each linear predictor accepts merlin’s extended syntax, so a spline or a fractional polynomial of a continuous covariate can be written directly in the model.
Restricted cubic splines on the log hazard scale
The baseline is a restricted cubic spline of log time, or of time, with the degrees of freedom or the knots you choose. By default internal knots sit at centiles of the event times and boundary knots at the earliest and latest event times.
Time-dependent effects in either model
Through restricted cubic splines of log time or of time, with tvc() and dftvc(), or written directly in the linear predictor.
Further timescales
Further timescales in either model, time2() to time5(): attained age as a second timescale, for example, when time since diagnosis is the main one.
Predictions after estimation
Survival, hazard, cumulative hazard and cumulative incidence functions; restricted mean survival time and time lost; differences and ratios of these between covariate patterns; confidence intervals. See help stexcess postestimation.
Get started
stexcess 1.1.2 is available now. It needs Stata 17 or later, merlin 2.5.0 and stmerlin 1.1.2.
* Stata 17 or later · requires merlin 2.5.0 and stmerlin 1.1.2
net install merlin, from("https://reddooranalytics.se/install/stata/merlin/2.5.0/")
net install stmerlin, from("https://reddooranalytics.se/install/stata/stmerlin/1.1.2/")
net install stexcess, from("https://reddooranalytics.se/install/stata/stexcess/latest/")A first model
// a reference population and a cancer population in one dataset: everyone has the
// reference hazard, and those with cancer==1 an excess hazard as well
clear
set seed 725
set obs 3000
generate byte cancer = runiform()>0.5
generate double age = rnormal(0,10)
generate byte sex = runiform()>0.5
generate double t1 = (-ln(runiform())/(0.01*exp(0.01*age + 0.2*sex)))^(1/1.5)
generate double t2 = (-ln(runiform())/(0.02*exp(0.02*age + 0.3*sex)))^(1/1.2) if cancer
generate double stime = min(t1, t2, 10)
generate byte died = stime<10
stset stime, failure(died)
// a flexible parametric excess hazard model
stexcess (age sex, df(3))(age sex, df(2)), indicator(cancer)
// a time-dependent effect of age in the reference model
stexcess (age sex, df(3) tvc(age) dftvc(1))(age sex, df(2)), indicator(cancer)Checked with merlin 2.5.0 and stmerlin 1.1.2 on Stata 19.5: the example below, predictions after it, a spline and a fractional polynomial of a covariate, a model with a further timescale, and a model with delayed entry.