How We Know Any of This

TL;DR. Medicine's hardest problem is not finding treatments, it is telling whether one worked. People get better on their own, get worse on their own, and differ from each other in a thousand ways that also predict their outcome. The randomised controlled trial exists to solve exactly one problem: making the treated group and the untreated group identical in every respect except the treatment. Every weaker form of evidence (observational studies, expert opinion, "I tried it and felt better") is vulnerable to confounding, and the history of medicine is a long list of treatments that looked obviously effective and were not. Some of them killed people at scale.

Key takeaways

  • Correlation is not causation is a slogan; confounding is the mechanism. If the people who take a treatment differ from those who do not, the comparison is broken before it starts.
  • Randomisation is the only method that balances the confounders you never thought of. That is the whole reason it is the gold standard.
  • Relative risk reduction ("cuts risk 50 percent") is nearly meaningless without the baseline. Ask for the absolute reduction and the number needed to treat.
  • A surrogate endpoint (a lab value) is not an outcome (a death, a stroke, a fracture). Drugs that improved surrogates and killed patients are the reason this distinction is written in blood.
  • Screening statistics are systematically distorted by lead-time bias, length bias, and overdiagnosis. Five-year survival is the wrong measure for screening and is quoted for it constantly.
  • A test's usefulness depends on how common the disease is. The same 99 percent accurate test is excellent for a common condition and misleading for a rare one.

The hierarchy of evidence, and what each level can and cannot do

In short: Case reports raise alarms, cohort studies show associations, and only randomisation can demonstrate cause.

LevelWhat it isGood forFails at
Case reportOne patient describedSpotting something new, raising an alarmProving anything
Case seriesA handful of similar patientsThe same, with slightly more weightThe same
Case-controlCompare people with the disease to people without, look backward for exposuresRare diseases, fast and cheapRecall bias, choosing controls badly
Cohort studyFollow a large group forward, see who gets sickReal-world exposures, long horizons, questions you cannot randomiseConfounding, always
Randomised controlled trial (RCT)Assign treatment by chance, then compareCausationCost, duration, artificial populations, ethics
Systematic review and meta-analysisPool all studies to a rigorous protocolPrecision, resolving disagreementGarbage in, garbage out; publication bias

The reason case reports still matter is that they are how the system notices. The AIDS epidemic was first visible as a June 1981 CDC report of five cases of an unusual pneumonia in previously healthy young men in Los Angeles. It proved nothing and it started everything.

Confounding, in one example

In short: Moderate drinkers looked healthier largely because the comparison group included people who had stopped drinking because they were already ill.

Observational studies found for years that people who drank moderately had lower mortality than people who did not drink at all. The obvious conclusion is that a glass of wine is protective. The problem is the comparison group. "Non-drinkers" includes former heavy drinkers who quit because they were ill, people too sick to drink, and people avoiding alcohol for religious or health reasons that correlate with everything else. Once studies separated lifelong abstainers and adjusted properly, and once genetic (Mendelian randomisation) analyses used alcohol-metabolism variants as a natural experiment, most of the protective effect disappeared.

That is confounding: a third factor that causes both the exposure and the outcome. It is not an occasional nuisance in observational research. It is the default condition.

Statistical adjustment helps and cannot finish the job, because you can only adjust for what you measured. Randomisation balances everything, measured and unmeasured, by construction. That is its entire and sufficient justification.

The most expensive lesson came from hormone replacement therapy. Observational studies through the 1990s consistently showed lower heart disease in women taking HRT, and it was prescribed widely on that basis. The Women's Health Initiative randomised trial, reported in 2002, found the combined therapy did not prevent heart disease and increased breast cancer and stroke. The observational studies had been comparing healthier, wealthier, more health-engaged women to everyone else. (The full picture is more nuanced than the 2002 headlines, with age at initiation mattering a great deal, which is its own lesson about how trial results get flattened in transmission.)

What makes a trial trustworthy

  • Randomised, ideally with concealed allocation so no one can steer patients into arms.
  • Blinded: single (patient), double (patient and clinician), or triple (plus analysts). Blinding matters most for subjective outcomes, and is impossible for some interventions, which is why surgical and behavioural trials are hard.
  • Controlled against placebo, or against current best treatment where withholding it would be unethical.
  • Analysed by intention to treat: everyone counted in the group they were assigned to, even if they stopped the drug. Analysing only completers reintroduces the confounding randomisation removed, because the people who quit are different.
  • Preregistered: outcomes declared before the data are seen. Without this, a negative trial can be rewritten around whatever subgroup happened to look good, a practice common enough to have a name (HARKing).
  • Powered and long enough to measure the outcome that matters.

Absolute vs relative: the number that gets hidden

In short: A 50 percent reduction can mean one percentage point, so ask for the absolute change and the number needed to treat.

A drug reduces heart attacks from 2 percent to 1 percent over five years.

  • Relative risk reduction: 50 percent. True, and the number in the press release.
  • Absolute risk reduction: 1 percentage point.
  • Number needed to treat (NNT): 100. One hundred people take it for five years for one to avoid a heart attack. The other 99 get no benefit and all of the side effects.

Whether 100 is a good deal depends on the drug's cost, its harms, and the seriousness of the event prevented. For a cheap, well-tolerated pill preventing a heart attack, it is a good deal. For an expensive drug with serious toxicity preventing a mild symptom, it is not. None of that is decidable from the 50 percent figure alone.

Number needed to harm (NNH) is the same arithmetic for adverse effects, and a sensible treatment discussion is a comparison of the two.

Illustrative exampleARR over 5 yearsNNTReading
Statin, someone at high cardiovascular riskAbout 2 to 4 points25 to 50Strong case
Statin, someone at low riskWell under 1 pointSeveral hundredWeak case; preference and cost decide
Blood pressure treatment, stage 2 hypertensionLargeLow tensClear case

This is the single most useful piece of numeracy a patient can carry into a consultation: "What is my absolute risk now, and what is it with treatment?"

Surrogates and hard endpoints

In short: Drugs that improved a lab value while killing patients are why a surrogate endpoint is a hypothesis until an outcome trial confirms it.

A surrogate endpoint is a measurable stand-in assumed to track the outcome you care about: blood pressure for stroke, LDL cholesterol for heart attack, HbA1c for diabetic complications, bone density for fractures, tumour shrinkage for survival. They make trials faster and cheaper, and they are sometimes wrong in ways that kill people.

The CAST trial (1989 to 1991) is the case every medical student learns. Extra heartbeats after a heart attack predict sudden death, and antiarrhythmic drugs suppressed them beautifully. The surrogate improved. When the drugs were finally tested against actual mortality, the treated patients died at roughly three times the rate of placebo. The trial was stopped early. Estimates of the excess deaths from the years of use beforehand run into the tens of thousands in the United States alone.

Other examples of the same failure mode: rosiglitazone lowered blood sugar and raised cardiovascular risk; torcetrapib raised "good" HDL cholesterol and increased mortality; high-dose oxygen, tight glucose control in critical illness, and several cancer drugs approved on tumour response later showed no survival benefit.

The rule that follows: a surrogate is a hypothesis until an outcome trial confirms it. When a chapter in this book says a treatment reduces deaths or strokes, that is a stronger claim than saying it improves a number.

Screening: why it needs its own statistics

In short: Lead-time bias, length bias, and overdiagnosis all inflate survival without saving anyone, so five-year survival is the wrong measure for screening.

Screening tests healthy people for early disease. It is intuitively obvious that finding cancer early must be good. It is often true and it is not automatically true, and three biases make screening look better than it is.

Lead-time bias. Diagnose a cancer three years earlier without changing the date of death, and survival time from diagnosis rises by three years. The patient gained nothing but three more years of knowing.

Length bias. Screening at intervals preferentially catches slow-growing tumours, because fast ones appear and cause symptoms between screens. So screen-detected cancers look more survivable, partly because they are a gentler selection of cancers.

Overdiagnosis. The extreme of length bias: finding disease that would never have caused symptoms in the person's lifetime. Every such case is counted as a life saved and is actually a patient given a diagnosis, treatment, and its harms for no benefit. Neuroblastoma screening in infants in Japan and Germany found many more tumours without reducing deaths and was stopped. Thyroid cancer diagnoses in South Korea rose roughly fifteen-fold after widespread ultrasound screening while mortality stayed flat.

The consequence: five-year survival is the wrong measure for screening, because all three biases inflate it. The right measure is disease-specific mortality in a randomised comparison, and by that standard some screening programmes are clearly worthwhile (cervical, colorectal), some are worthwhile with real trade-offs (mammography, lung CT in heavy smokers), and some have been abandoned.

Test accuracy and the base rate

In short: The same 99 percent accurate test is useful for a common disease and misleading for a rare one, because the population changed rather than the test.

Two properties describe a test. Sensitivity is the proportion of people with the disease it catches. Specificity is the proportion without it that it correctly clears. Neither answers the question a patient actually has, which is "I tested positive, do I have it?" That is the positive predictive value, and it depends on how common the disease is.

Take a test with 99 percent sensitivity and 99 percent specificity, applied to 10,000 people where 1 percent have the disease:

  • 100 have it; 99 test positive.
  • 9,900 do not; 99 test positive anyway (1 percent of 9,900).
  • Total positives: 198. Half are wrong. PPV = 50 percent.

Now apply the same test where only 1 in 10,000 has the disease: 1 true positive and about 100 false positives. PPV under 1 percent. The test did not change. The population did.

This is why screening rare conditions in the general population produces mostly false alarms, why confirmatory testing exists, and why "the test is 99 percent accurate" is not an answer to anything on its own.

When randomising is impossible

Nobody will randomise people to smoke for thirty years. So how was smoking established as a cause of lung cancer? By meeting, across many observational studies, criteria articulated by Austin Bradford Hill in 1965:

Strength (smokers had many times the risk, not a few percent more), consistency (the same result in different countries, designs, and decades), temporality (smoking came first), biological gradient (more cigarettes, more cancer), plausibility and coherence (carcinogens in smoke, changes visible in airway tissue), experiment (quitting lowers risk; animal exposure causes tumours), and specificity where applicable.

No single criterion proves cause. Together, and with a dose-response relationship this steep, they are conclusive. The same framework now applies to air pollution, asbestos, alcohol and cancer, and lead and cognitive development.

Reading a health headline without being fooled

In short: Eight questions dispose of most health headlines, starting with whether the study was in humans and whether it was randomised.

A short checklist that resolves most of them:

  1. Humans? Mice and cell cultures are not people, and most mouse results do not replicate in humans.
  2. Randomised or observational? If observational, assume confounding until shown otherwise.
  3. How many people, for how long? A 40-person 6-week study is a hypothesis.
  4. Absolute numbers? If only percentages are given, the absolute effect is probably small.
  5. Real outcome or surrogate?
  6. Compared to what? Placebo, nothing, or the current best treatment (which is the only comparison that matters clinically).
  7. Who funded it, and was it preregistered? Industry funding does not invalidate a trial, and it does predict which questions get asked and which results get published.
  8. Has anyone replicated it? A single positive study is where science starts.

Don't be confused: "no evidence of benefit" is not "evidence of no benefit." A small trial that fails to find an effect may simply have been too small to see one. The two statements are routinely swapped in both directions, by enthusiasts and by sceptics.

Sources and notes

Evidence hierarchy and trial design: standard clinical epidemiology (Sackett et al., Evidence-Based Medicine; Guyatt et al., Users' Guides to the Medical Literature). CAST: the Cardiac Arrhythmia Suppression Trial, New England Journal of Medicine, 1989 and 1991. Women's Health Initiative: JAMA, 2002. The first AIDS report: MMWR, 5 June 1981. Bradford Hill criteria: Proceedings of the Royal Society of Medicine, 1965. South Korean thyroid cancer overdiagnosis: Ahn, Kim, and Welch, NEJM, 2014. Neuroblastoma screening: Schilling et al. and Woods et al., 2002. Alcohol and Mendelian randomisation: Millwood et al., The Lancet, 2019, among others. NNT figures here are illustrative round numbers, not guideline values.

Open questions. How much observational "real world evidence" can substitute for trials is a live methodological argument, sharpened by the availability of very large health datasets. The balance of benefit and harm in several screening programmes, prostate and breast especially, remains genuinely contested among reasonable experts.

That is the toolkit. Now the diseases, starting with the one that best rewards understanding it. 👉