← Disease Detectives
0/8 sections complete done☰ All chapters
Chasing the Curve/Chapter 6
06

Crunching the Numbers

Attack rate, relative risk & odds ratio: the math that turns hypotheses into evidence

This is where Step 7, evaluating a hypothesis epidemiologically, actually happens. Everything starts with sorting people into a 2×2 table: exposed or not, ill or not.

🥔 Worked Example: The Cedarwood Picnic (Cohort Study)

Investigators suspect the potato salad at a class picnic made people sick. Because everyone who attended is known (a defined cohort), they can survey all 120 attendees about what they ate and whether they got sick:

IllNot IllTotal
Ate potato salad602080
Didn't eat potato salad83240
Total6852120
Attack Rate
Cases in groupTotal in that group

Exposed: 60 ÷ 80 = 75%
Unexposed: 8 ÷ 40 = 20%

A special "incidence" just for outbreaks
Relative Risk
Attack rate, exposedAttack rate, unexposed

75% ÷ 20% = 3.75

The exposed group got sick 3.75× as often
Odds Ratio
(a × d)(b × c)

Used for case-control studies, where relative risk can't be calculated directly

See worked example below ↓
Reading the results: A relative risk of 3.75 is strong evidence. It means people who ate the potato salad were nearly 4 times as likely to get sick as people who didn't. (Say "times as likely": some graders take points off for "times more likely.") An RR (or OR) of exactly 1 would mean no difference in risk at all between the two groups. Above 1, the exposure goes with more illness; below 1, with less (it's protective, like a vaccine).
The "everybody ate it" trap: a food that 95% of guests ate will show a high attack rate among eaters, simply because nearly every sick person ate it too. Don't stop at the exposed attack rate: compare it to the unexposed attack rate. If the few people who skipped the dish got sick just as often, the RR is near 1 and that dish isn't the source.
Risk difference (attributable risk): relative risk says how many times as likely; the risk difference says how much extra risk. Risk difference = attack rate exposed − attack rate unexposed. At the Cedarwood Picnic: 0.75 − 0.20 = 0.55, so 55 more cases per 100 people who ate the potato salad.
Turning an RR into a percent: the exposed group's risk is (RR − 1) × 100% higher. RR = 1.8 means 80% higher risk (not 1.8% or 180%); RR = 3.75 means 275% higher. Below 1, the risk is (1 − RR) × 100% lower: RR = 0.6 means 40% lower. For an odds ratio, say "odds" instead of "risk."

How Much Illness Is Due to the Exposure? State/Nats

🥗
Attributable risk percent

Among the exposed: (AR exposed − AR unexposed) ÷ AR exposed = (RR − 1) ÷ RR. At the Cedarwood Picnic: (3.75 − 1) ÷ 3.75 = 73%, so about three-quarters of the illness among salad eaters came from the salad.

👥
Population attributable fraction

The same question for everyone: (attack rate in the whole group − attack rate unexposed) ÷ attack rate in the whole group. It depends on how common the exposure is, so a strong but rare exposure causes only a small share of cases.

💊
Number needed to treat

NNT = 1 ÷ the risk difference (the absolute risk reduction). If a drug lowers heart-attack risk from 2.0% to 1.0%, NNT = 1 ÷ 0.010 = 100: treat 100 people to prevent one heart attack.

Confidence interval: a study only samples some people, so a result like RR = 3.75 is an estimate. A 95% confidence interval gives a range likely to contain the true value (for example 2.1 to 6.7). If the interval for an RR or OR includes 1, the study can't rule out "no difference." A bigger study gives a narrower (more precise) interval. State/Nats For an odds ratio the interval is worked out on the log scale: 95% CI = eln(OR) ± 1.96 × SE.

🧫 Worked Example: A Case-Control Study

Now imagine the outbreak was community-wide, so investigators can't identify every exposed person. Instead, they compare 50 cases (confirmed ill) to 50 controls (healthy people, similar in age and location) and ask both groups about past exposures:

CasesControls
Exposed (a, b)4015
Not exposed (c, d)1035
Total5050

Odds ratio = (40 × 35) ÷ (10 × 15) = 1,400 ÷ 150 = ≈ 9.3. Because a case-control study starts with the outcome (already-diagnosed cases) rather than following a group forward in time, investigators can't calculate a true attack rate or relative risk, but the odds ratio is close to the relative risk when the disease is rare. When lots of people got sick, the OR comes out bigger than the RR.

More than one exposure in the table? Work out a separate 2×2 for each exposure. For exposure A, "exposed" is everyone who had A and "not exposed" is everyone else, including people who had exposure B. Then compare the odds ratios: the largest one (with a sensible story) points to the likely source. Say it the right way round: "cases had 12 times the odds of having been exposed," not "exposed people had 12 times the risk."
🤔 Which Formula Do I Use?
Did the study start with an exposed group, followed forward to see who got sick? (like the picnic: a cohort)
→ Use Relative Risk
Did the study start by picking already-sick vs. healthy people, then look backward at exposure? (a case-control)
→ Use Odds Ratio

How Good Is the Test? Sensitivity & Specificity

Lab tests confirm cases, but no test is perfect. Two numbers describe how good a test is. Learn them by meaning, not by letter, because tests put the table together in different ways:

Example: a quick test is tried on 40 people who have the disease and 160 who don't. It comes back positive for 36 of the sick people and 32 of the healthy ones.

Layout 1: disease in the rows
Test +Test −
Has disease36 (TP)4 (FN)
No disease32 (FP)128 (TN)
Layout 2: disease in the columns
Has diseaseNo disease
Test +36 (TP)32 (FP)
Test −4 (FN)128 (TN)
Sensitivity
TPTP + FN

36 ÷ (36 + 4) = 90%

Misses 10% of real cases
Specificity
TNTN + FP

128 ÷ (128 + 32) = 80%

Wrongly flags 20% of healthy people
Positive predictive value
TPTP + FP

36 ÷ (36 + 32) = 53%

Only half of positives are really sick
Same numbers, same answers, either layout. Before plugging in, circle the two cells where the test was right (true positive, true negative). Then ask: sick people go on the bottom of sensitivity, healthy people on the bottom of specificity.

What low values cost: low sensitivity means false negatives: sick people are told they're fine, go untreated and keep spreading the disease. Low specificity means false positives: healthy people get unneeded treatment, worry and cost, and case counts get inflated.

Predictive Values: What a Result Means

➕
PPV

Positive predictive value = true positives ÷ everyone who tested positive. "If I test positive, how likely am I to really have it?"

➖
NPV

Negative predictive value = true negatives ÷ everyone who tested negative. "If I test negative, how likely am I to really be healthy?"

📉
Rarer disease, lower PPV

Sensitivity and specificity belong to the test, but PPV falls when a disease becomes rarer: the false positives from many healthy people make up a bigger share of the positives. That's why screening a low-risk group brings lots of false alarms.

Screening vs. confirming: a screening test should be highly sensitive, so it misses as few sick people as possible (a negative then really means "probably not sick"). Anyone who screens positive gets a second, highly specific test to weed out the false positives. Why not always use the best "gold standard" test? It's usually slower, costlier or more invasive.
To raise a test's sensitivity: lower the cutoff for a positive, run two tests and count a positive on either one, or collect the sample at a better time. Requiring both tests to be positive does the opposite: it raises specificity.

Vaccine Effectiveness State/Nats

Vaccine effectiveness (VE, also called vaccine efficacy in a trial) uses the same cohort table as relative risk, with vaccinated as the "exposed" row. It answers: by what percent did the vaccine cut the risk of getting sick?

Step 1: Attack rates

Vaccinated: 10 of 200 ill = 5%
Unvaccinated: 50 of 200 ill = 25%

Step 2: Relative risk
AR vaccinatedAR unvaccinated

5% ÷ 25% = 0.20

Step 3: VE = 1 − RR

1 − 0.20 = 0.80 = 80%

Vaccinated people had 80% lower risk
Say it in context: "Vaccinated people had an 80% lower risk of getting the disease than unvaccinated people." An RR below 1 means the "exposure" (the vaccine) is protective. The best study for testing a new vaccine is a randomized controlled trial: volunteers are randomly assigned to the vaccine or a placebo, then followed to see who gets sick. The main concerns are ethics (withholding a working vaccine from the placebo group) and cost.

Try Your Own Numbers

Build your own 2×2 table and watch the math update live: a good way to build a feel for how each number moves the result.

IllNot Ill
Exposed
Unexposed
Attack Rate, Exposed
N/A
Attack Rate, Unexposed
N/A
Relative Risk
N/A

Try setting exposed and unexposed attack rates equal: watch the relative risk head toward 1, meaning no difference in risk at all.

✓ Check Yourself

Q1At the Cedarwood Picnic, the attack rate was 75% for people who ate the potato salad and 20% for people who didn't: RR = 3.75. What does that mean?

Q2Why can't investigators calculate a true relative risk from a case-control study?
A case-control study starts by selecting people based on their outcome (already sick or not), not by following a defined population forward from exposure, so there's no true "population at risk" to calculate a rate from. An odds ratio is used instead.
Q3A test is given to 23 people with a disease (18 test positive) and 88 people without it (52 test negative). What are its sensitivity and specificity?
Sensitivity = 18 ÷ 23 = 78%. Specificity = 52 ÷ 88 = 59%. Divide by everyone who truly has (or doesn't have) the disease, whichever way the table is drawn.

Q4Of 30 unvaccinated birds, 18 got sick; of 30 vaccinated birds, 4 got sick. What is the vaccine effectiveness, in percent?

Q5Of 40 people who ate the chicken, 30 got sick. What is the attack rate among those who ate it, in percent?

Q6A relative risk of exactly 1 means...