What Evidence-Based Medicine Got Wrong (And What It Got Right)
Evidence-based medicine transformed clinical practice. It also produced a hierarchy of evidence that has been applied with a rigidity its founders would not have recognised or endorsed.
When Archie Cochrane published Effectiveness and Efficiency in 1972, the central argument was that medical practice was full of interventions that had never been properly evaluated and that patients were receiving treatments of unknown benefit at substantial cost and some risk. The randomised controlled trial, he argued, was the only reliable way to know whether a treatment worked.
That argument was correct, and important, and it drove one of the genuine transformations of twentieth-century medicine. The infrastructure it produced, clinical trial methodology, systematic reviews, meta-analyses, clinical guidelines, evidence grading systems, changed how decisions were made across every clinical speciality. The idea that you should know whether what you are doing works before you do it to a patient seems obvious now. It was not obvious then, and making it standard required decades of difficult institutional change.
But somewhere along the way, a useful heuristic, RCTs are more reliable than observational studies, became a rigid hierarchy, and the hierarchy was applied in ways its founders would not have recognised or endorsed. The evidence-based medicine movement was right that evidence should inform clinical decisions. It was wrong, or at least incomplete, about which evidence counts as evidence.
The hierarchy problem
The standard hierarchy, systematic reviews and meta-analyses at the top, followed by RCTs, followed by observational studies, followed by expert opinion, made sense as a guide to scepticism. Evidence from well-conducted RCTs is, on average, more reliable than evidence from uncontrolled observational studies, and treating them as equivalent would have produced predictable errors.
The problem is that the hierarchy became prescriptive in ways that did not account for context. A Cochrane review of RCTs in a specific patient population was treated as more authoritative than extensive, well-characterised real-world evidence in a different population, even when the real-world population was more representative of clinical practice. A single well-conducted observational study with millions of patients and years of follow-up was discounted relative to a meta-analysis of several small, short-duration RCTs with high dropout rates.
The hierarchy tells you which study designs are most protected against specific types of bias. It does not tell you which studies are most relevant to your specific clinical question. Those are different questions, and conflating them has produced clinical guidance that is methodologically impeccable and sometimes clinically unhelpful.
The population mismatch problem
The populations enrolled in clinical trials are systematically different from the populations that receive treatments in clinical practice. Trial populations tend to be younger, with fewer comorbidities, higher adherence to treatment protocols, and more frequent monitoring than would be typical in routine care. This is not a failure of trial design, it is often deliberate, intended to reduce noise and improve the probability of detecting a treatment effect.
But it creates a generalisation problem. A treatment demonstrated to be effective in a trial population of 18- to 65-year-olds without significant comorbidities may behave very differently when given to the 78-year-olds with four comorbidities and six concurrent medications who make up a substantial portion of the patients who will actually receive it.
Observational evidence from the real-world population can address this directly in ways that even the most perfectly designed RCT cannot. A randomised trial tells you about efficacy. Real-world data tells you about effectiveness in practice. Both questions matter, and conflating them produces systematic errors in clinical guidance.
What guidelines get wrong when they grade evidence
Most clinical guideline development bodies use evidence grading systems, GRADE being the most widely adopted, that rate the quality of evidence and the strength of resulting recommendations. The system is used by the WHO, NICE, the CDC's Advisory Committee on Immunization Practices, and dozens of other bodies worldwide. These systems are designed to protect against overconfident recommendations from weak evidence. They are valuable, and the guidance they produce is generally more careful and more transparent than the pre-EBM era of expert opinion unconstrained by explicit evidence assessment.
The problem is that evidence grade and recommendation strength are sometimes interpreted as inversely related to clinical judgment. A "strong recommendation, moderate quality evidence" becomes the threshold for changing practice, and evidence that does not fit neatly into randomised trial categories struggles to achieve that threshold regardless of how compelling it is.
This creates perverse incentives. Clinical questions that are amenable to RCTs get answered. Clinical questions where randomisation is impractical, unethical, or produces populations irrelevant to clinical practice get answered inadequately or not at all. The evidence base ends up shaped by what is easy to study rather than by what is most clinically important.
What the founders actually said
David Sackett, one of EBM's founders, was explicit that evidence-based medicine was never meant to replace clinical judgment. His 1996 definition is worth quoting: "Evidence based medicine is the conscientious, explicit and judicious use of current best evidence in making decisions about the care of individual patients. The practice of evidence based medicine means integrating individual clinical expertise with the best available external clinical evidence from systematic research."
The phrase "best available external clinical evidence" is doing important work there. It does not say "evidence from RCTs only." It says best available, which means using the highest-quality evidence that exists for the question at hand, not requiring a type of evidence that may not exist for that question. The original Sackett paper is short and worth reading if you have not done so recently.
A more useful relationship with evidence
The researchers and clinicians who use evidence most effectively do not apply the hierarchy mechanically. They ask, for each clinical question, what kind of evidence would most reliably answer it, and then they assess what evidence of that type exists and how well it was conducted.
Sometimes the answer is that a well-conducted RCT provides the most reliable answer. Sometimes it is that real-world evidence from a large, representative population is more informative than the available trials. Sometimes it is that the evidence on the specific question is genuinely sparse and that honest uncertainty is the right response, not a forced recommendation derived from imperfect analogy.
That judgment requires a comprehensive and honest picture of the available evidence in all its forms. Not just the RCT literature, not just the indexed journal articles, but the full landscape of what has been studied, in what populations, using what methods, and with what results. Building that picture is difficult manually and increasingly feasible with the right tools. It is one of the core problems NousLab was built to address. See how our evidence mapping approach works across different evidence types.