A cohort study follows a defined group of people forward in time, recording exposures and outcomes as they occur, and it can show that two things travel together with notable strength. It cannot, on its own, show that one causes the other. The Framingham Heart Study, begun in 1948 with 5,209 residents of a Massachusetts town and still running under federal support, has traced how smoking, blood pressure and cholesterol track with heart disease across generations, and it says so in the language of association.
This piece is about research design, not personal health advice; individual decisions belong with a clinician who knows the person making them.
How does a cohort study work?
Researchers enroll participants free of the outcome under study, measure their characteristics, and wait. Visits, questionnaires and records accumulate over years or decades, and analysts compare outcome rates among people with different exposures. Because observation happens without assignment, a cohort captures real behavior: people smoke or quit, take hormones or do not, live where they live. No randomized trial can reproduce that breadth, and none should try for exposures such as smoking.
The classics define the form. Framingham, funded by the United States National Heart, Lung, and Blood Institute, has followed original participants and, later, their children and grandchildren. The Nurses' Health Study, launched by Harvard researchers in 1976 with more than 100,000 nurses, showed how a willing, well-documented cohort can support thousands of analyses on diet, hormones, work and disease. Doll and Hill's study of more than 34,000 British doctors, begun in 1951, tracked smoking and lung cancer across 50 years and remains the template for smoking evidence.
What can a cohort show that a trial cannot?
Cohorts answer questions randomized trials are unfit to ask. Randomly assigning people to smoke, to hold a stressful job, or to live near a highway for decades would be unethical, so for harmful or self-chosen exposures, observation is the only ethical instrument. Cohorts also excel at rare outcomes that need huge populations and long horizons, at natural experiments such as quitting smoking at different ages, and at describing how risk unfolds over a lifetime rather than over a trial's few years.
Their second strength is realism. Trial volunteers are screened, adherent and watched; cohort members skip medications, move house and age in ordinary ways. When the question is how an exposure behaves in the population, that messiness is a feature.
Why can't a cohort prove cause?
Because people are not assigned to their exposures, so exposed and unexposed groups differ in ways beyond the exposure itself. Smokers differ from non-smokers in more than smoking; coffee drinkers differ from abstainers in sleep, work and health habits. These lurking differences, called confounders, can manufacture an association that looks like cause. Researchers adjust statistically for what they can measure, but no adjustment covers what nobody measured.
Another trap is reverse causation: early disease may change behavior rather than behavior causing disease. People who feel unwell may give up alcohol or exercise less, so abstinence can appear to precede illness. And measurement that happens once, decades before outcomes, can miss changes in between. Each cohort paper lists these limits; the honest ones lead with them.
| Feature | Cohort study | Randomized trial |
|---|---|---|
| Assignment | None; exposures observed | Random allocation |
| Best at | Real-life exposures, long horizons, rare outcomes | Causal tests of assignable interventions |
| Main weakness | Confounding and reverse causation | Artificial conditions, selected volunteers |
| Typical scale | Thousands to hundreds of thousands | Dozens to tens of thousands |
| Language of results | Associated with, linked to | Reduced, increased (within trial) |
How do we know when a cohort finding is strong?
Strength is checkable, item by item. A strong finding shows a large association that is hard to explain by measured confounders, appears consistently across subgroups and study designs, and follows a plausible time order, exposure first, outcome later. Doll and Hill's smoking results met every bar: enormous effect sizes, dose-response gradients, consistency with earlier case-control work, and later mechanistic corroboration. Weak findings tend to be modest in size, sensitive to adjustment choices, and inconsistent between cohorts.
Triangulation is the modern habit: when a cohort, a natural experiment and a trial all point the same way, confidence rises. When they diverge, as hormone therapy did in the 1990s, before randomized results contradicted optimistic observational findings, the divergence itself is the lesson. The Women's Health Initiative trials, begun in the 1990s, showed why: the women who chose hormone therapy were healthier to start with than those who did not.
What should a reader ask of a cohort headline?
Three questions do most of the work. Compared with whom, since a cohort's comparison group defines the association. Adjusted for what, since statistical adjustment handles only measured variables. And what outcome rate, since a doubling of a tiny risk remains tiny. Headlines that convert association into cause, coffee prevents disease, sitting shortens life, should be read as hypotheses the cohort supported, not verdicts.
The record justifies both respect and caution. Cohort studies built the modern understanding of cardiovascular risk and of smoking, and they also produced confident findings that trials later overturned. The design is not the problem; forgetting what it measures is.
Prospective cohorts, which recruit participants before outcomes occur, hold a timing advantage over retrospective designs that reconstruct the past from records and memory; recall degrades, and the records were never kept for research purposes. Loss to follow-up is the second structural risk: participants who drop out rarely do so at random, and if attrition differs by exposure, the surviving comparison drifts. The best cohorts report retention by group and test whether leavers resemble stayers, which lets readers judge how much of an association might be attrition in disguise.
For more context, read What makes an observational study strong.
For more context, read animal studies.
For more context, read How does peer review actually work before publication?.
