What is real-world evidence?
What real-world data and evidence are, where the data come from, what RWE is used for, the main biases, and what makes a study credible.
Real-world evidence (RWE) is evidence about how a medical treatment is used, and about its benefits and risks, obtained by analysing real-world data (RWD). These are data on patients’ health and care collected outside a tightly controlled clinical trial, usually during routine care, in electronic health records, insurance claims and registers. How far the evidence can be trusted depends on the study design as much as on the data.
Real-world data and real-world evidence
The US Food and Drug Administration (FDA) set out the distinction in its Framework for FDA’s Real-World Evidence Program (December 2018), which the 21st Century Cures Act of 2016 required, for drugs and biologics. Real-world data are data on patients’ health status or the delivery of health care, collected routinely from a variety of sources. Real-world evidence is the clinical evidence about the use and potential benefits or risks of a medical product that comes from analysing them.
NICE’s real-world evidence framework (first published in June 2022, and updated since) uses similar definitions. It covers disease epidemiology and health services research as well as treatment effects, and counts a single-arm trial with an external control built from real-world data as an RWE study. Both frameworks also count randomised trials that use real-world data, such as pragmatic or registry-based trials, as sources of real-world evidence.
Where the data come from
- Electronic health records. Diagnoses, test results and prescriptions, recorded when the patient’s care needed them.
- Claims and billing data. Diagnoses, procedures and dispensed prescriptions that were paid for, with little clinical detail such as disease stage or laboratory values.
- Disease and product registers. Cancer registers, clinical quality registers and drug or device registries record defined clinical variables for patients with a condition or treatment.
- National health registers. In Denmark, Finland, Iceland, Norway and Sweden, each resident has a personal identity number used across the national health registers, including those of hospital contacts, prescriptions, cancers and causes of death. Linking them gives population-based data with virtually complete follow-up (Laugesen et al., 2021).
- Patient-generated data. Questionnaires, home-use devices, wearables and apps.
What it is used for
- Effectiveness in routine practice. Trials often exclude older and sicker patients and follow them briefly. Real-world data cover the people actually treated, for longer, and can compare treatments never tested head to head.
- Safety. Rare or delayed adverse events need more patients and longer follow-up than most trials have.
- Natural history and external controls. How a disease progresses under current care and, when randomisation is not feasible, a comparison group for a single-arm trial. The FDA notes that differences in practice, diagnostic criteria, outcome measures and follow-up make a comparable external group hard to select.
- Inputs to health economic models. Baseline risks, treatment patterns, resource use and long-term survival in the population a decision concerns.
- Regulatory and HTA decisions. The European Medicines Agency set up DARWIN EU in 2022, a federated network in which data partners keep their data, run the analyses locally and return only aggregated results, giving regulators studies on the use, safety and effectiveness of medicines. ICH’s M14 guideline sets out harmonised principles for safety studies of medicines that use real-world data. NICE’s framework sets out how RWE can inform its guidance.
What can go wrong
In a trial, randomisation makes the groups comparable, and follow-up starts at randomisation. An observational study has neither, so its design and analysis must supply them.
- Confounding by indication. Treatments are given for reasons that often predict the outcome. If sicker patients get the more intensive treatment, it can look harmful. Adjustment and weighting remove confounding only by measured factors, and frailty or disease severity are often poorly recorded. NICE prefers an active comparator, another treatment for the same indication, over untreated patients, because the patients are more alike.
- Selection. Who enters the data, and who stays, can depend on health. Comparing people already taking a drug (prevalent users) with non-users compares people who have tolerated it and survived its early risks. Restricting to new users avoids this (Ray, 2003).
- Measurement and missing data. Codes recorded for care or billing need validating. Values are missing when nobody measured them, which often depends on the patient’s health, so analysing only complete records can mislead.
- Immortal time bias. Grouping people by a treatment they start after follow-up begins credits the treatment with the time they had to survive to start it. The primers on immortal time bias and landmark analysis show how to avoid it.
Designing the study as a trial
Hernán and Robins proposed designing an observational study of a treatment’s effect around the randomised trial one would run if one could, the target trial (Am J Epidemiol 2016;183:758–764). Its protocol is specified first, sometimes revised after checking what the data can support, covering eligibility criteria, treatment strategies, assignment procedures, follow-up period, outcome, causal contrasts and analysis plan, and the study emulates each one. Ideally, time zero is when an eligible person starts a treatment strategy, so eligibility, assignment and the start of follow-up coincide, as at randomisation. This prevents immortal time bias and excludes prevalent users. When a person’s strategy is not yet known at time zero, for example with a grace period to start treatment, cloning, censoring and weighting keep time zero aligned. Confounding, measurement error and missing data remain, and are handled with the methods of causal inference.
NICE asks for comparative studies to be designed this way. The RCT-DUPLICATE initiative emulated 32 medication trials, chosen as feasible to emulate, in three US claims databases. Across all 32, trial and database effect estimates had a correlation of 0.82, and the database estimate fell within the trial’s 95% confidence interval for 66%. In a post hoc split, the correlation was 0.93 for the 16 trials whose design and measurements could be emulated closely and 0.53 for the other 16. The authors conclude that such studies do not replace trials (Wang et al., 2023).
What makes a study credible
- A prespecified protocol. Objectives, data, design and planned analyses, including sensitivity analyses, are fixed before the final analysis, and changes are recorded and justified. NICE points to the HARPER protocol template and encourages publishing the protocol on a public platform, such as ClinicalTrials.gov or the HMA–EMA Catalogues of real-world data sources and studies.
- A clear estimand. The population, treatments, outcome, how events such as switching treatment are handled, and the summary measure, for example a difference in five-year risk. These are the attributes of an estimand in the ICH E9(R1) addendum.
- Fit-for-purpose data. The FDA asks whether the data are reliable, meaning accurate, complete and traceable, and relevant, meaning they record the exposure, outcome and covariates the question needs in enough representative patients (FDA guidance, 2024). NICE provides a Data Suitability Assessment Tool (DataSAT) for reporting this.
- Sensitivity analyses. The analysis is repeated under other plausible definitions and assumptions, with quantitative bias analysis for the main threats. The E-value, for example, is the minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need with both treatment and outcome, beyond the measured covariates, to explain away an estimate (VanderWeele and Ding, 2017). NICE advises deciding in advance what strength is plausible.
- Reproducible reporting. Enough detail, such as code lists and the flow of patients, for an independent team to repeat the study, following a reporting guideline such as RECORD-PE or, for target trial emulations, TARGET (Cashin et al., 2025).