Correlation is not causation because two things can move together without one producing the other: a third factor may drive both, the timing may be coincidence, or the causation may run backward. A famous 2012 analysis in The New England Journal of Medicine found that a country's chocolate consumption correlated with its Nobel Prizes per capita — a nearly perfect statistical association that no one seriously believes is causal.
Engevity News publishes information, not medical advice, and this piece is about reading claims, not acting on them. The chocolate example, reported by cardiologist Franz Messerli, is a staple of statistical education precisely because it is absurd enough to make the mechanism visible.
What does a correlation actually measure?
A correlation measures how closely two measured quantities move together across a set of observations — countries, people, years — and nothing more. It has no sense of time order, no mechanism, and no ability to distinguish which variable is the horse and which is the cart. A strong correlation between exercise and longevity, for example, is consistent with exercise extending life, with healthy people exercising more, and with a third factor such as income supporting both. Each of those readings fits the same number. That is why statisticians describe correlation as an association: a summary of co-movement that narrows the possibilities without choosing among them.
Why do headlines convert correlations into causes?
Headlines compress, and causation compresses better than association. "Coffee may protect the heart" fits a headline; "coffee drinkers in one observational cohort had lower rates of a heart outcome, for reasons this design cannot isolate" does not. The incentive is structural: studies that measure thousands of people over years produce many correlations, and only some deserve the causal verb. When a headline says a food "boosts", "protects", or "raises the risk of" something, and the underlying study observed rather than assigned, the causal verb is doing work the data never did. The reader's corrective habit is simple: on seeing a causal verb, ask what kind of study produced the finding.
What is a confounder?
A confounder is a third variable that influences both the suspected cause and the outcome, manufacturing an association between them. The classic teaching example pairs ice cream sales with drownings: both rise in the same months, and the shared driver is summer, not dessert. In health research confounders are subtler — income, geography, baseline health, access to care — and researchers adjust for them statistically. Adjustment helps, but it can only account for confounders someone measured. The ice cream lesson survives the sophistication: the question is never only "how strong is the association" but "what else moves with it".
Can the causation run backward?
It can, and this pattern is called reverse causation: the outcome produces the exposure rather than the reverse. People in the early stages of an illness often change their behavior — cutting back on activity, seeking medication, losing weight — so a snapshot taken afterward can make the remedy look like the trigger. Studies of coffee and health wrestle with this constantly: sick people may avoid coffee, making moderate drinkers look artificially healthy. Longitudinal designs that follow people forward, and measure exposure before the outcome appears, exist largely to break this loop.
How do we know when causation is justified?
Causation earns support from converging designs, not from a single large association. The strongest single design is the randomized controlled trial, where participants are assigned to conditions by chance, distributing confounders evenly across groups in expectation. Where trials are impossible, researchers look for the checklist Bradford Hill articulated in 1965: consistency across studies, a plausible mechanism, a time order with cause before effect, a dose-response pattern, and strength of association. The link between smoking and lung cancer was established this way — through decades of converging evidence, not one correlation — because no single observational study could carry a causal conclusion alone. That standard, not the size of any one dataset, is what the word "cause" has to meet.
Which phrases should trigger suspicion?
Certain headline formulas reliably signal an association dressed as a cause.
- "X is linked to Y" — accurate phrasing, but the story often escalates it to cause.
- "Doing X raises your risk of Y" from an observational study.
- "New study proves" — studies rarely prove; they support or fail to support.
- "Scientists say" without a named study, journal, or institution.
How do researchers isolate cause without experiments?
Some questions can never be answered by assignment — no ethics board approves assigning smoking, poverty, or decades of stress — so researchers have developed designs that borrow the logic of experiments from accidents of circumstance. Natural experiments exploit events that mimic random assignment: a policy change, a boundary, a lottery. Mendelian randomization, a method from genetic epidemiology, uses the random shuffle of genes at conception as if it were assignment, since variants associated with, say, higher cholesterol are distributed independently of the lifestyle factors that usually confound the comparison. Each approach has assumptions that can fail, and researchers debate individual applications. What the designs share is the goal: to break the link between the exposure and the confounders that observational data cannot separate on their own. When a headline about diet or alcohol cites such a design, the honest reading is not "causation proven" but "a clever attempt to approximate an experiment" — closer to cause than a plain correlation, and short of a trial.
How can a reader practice the distinction?
A short drill builds the reflex. On the next health headline, find the named study and year, identify the design — observational or randomized — and restate the claim in the language the design supports: "people who did X had more Y", not "X causes Y". The restated version is usually smaller and more conditional. It is also, almost always, what the data actually showed.
For more context, read How to check a health claim in five minutes.
For more context, read confidence interval.
For more context, read What absolute versus relative risk really means.
