NovaMarket has 8,000 employees and, for the past two years, a wellness programme (preventive physiotherapy, active breaks in the warehouse, psychological support) run jointly with an insurer. The insurer has proposed evaluating the results together, and its reports speak a language of their own: prevalences, relative risks, odds ratios, screening sensitivities, survival curves. It is the language of biostatistics, and this lesson teaches you to read it with the tools you already have — because almost all of it is Bayes, 2×2 tables, intervals and experimental design under other names. The goal is not for you to play healthcare professional: it is for you, as an analyst, to be able to collaborate with those who are, and to read a health report or headline critically.
Important warning: health data are a special category under the GDPR, and any analysis involving employees requires a legal basis, anonymization, and prior review by the DPO and the occupational health/medical service. And healthcare decisions — diagnoses, treatments, screenings — always belong to healthcare professionals: this lesson teaches you to read the evidence, never to replace them.
Contents
- Measuring disease: incidence vs prevalence
- The wellness programme's 2×2 table: relative risk and odds ratio
- Screening: sensitivity, specificity and the positive-result illusion
- Study design: the hierarchy of evidence
- Survival analysis: reading a Kaplan-Meier curve
- Critical reading: headlines, relative and absolute risks
Measuring disease: incidence vs prevalence
The first distinction, and an inexhaustible source of confusion in reports:
- Prevalence: the proportion of people who have the condition at a given moment. It is a snapshot. Example: in the occupational health survey, 18% of warehouse staff report current lower back pain.
- Incidence: the rate at which new cases appear among those who did not have it. It is a film, and it is measured over observation time: 6 new episodes of back pain with sick leave per 100 person-years in the warehouse.
They are related but not interchangeable: prevalence ≈ incidence × average duration. A long-lasting chronic condition can have low incidence and high prevalence (diabetes); the flu, sky-high incidence and modest prevalence in any snapshot. The practical consequence: an effective prevention programme lowers incidence almost immediately, but prevalence can take years to show it (the existing cases are still there) — judging prevention with the wrong picture condemns it unfairly. The parallel with Club Nova is exact: the 14% annual churn rate is an incidence; the 64% of active members, a prevalence.
The wellness programme's 2×2 table: relative risk and odds ratio
Of the 4,000 logistics employees, 1,200 take part in the preventive physiotherapy programme. Episodes of back pain with sick leave over the past year:
| Back pain | No back pain | Total | Risk | |
|---|---|---|---|---|
| Participates | 60 | 1,140 | 1,200 | 5.0% |
| Does not participate | 224 | 2,576 | 2,800 | 8.0% |
| Total | 284 | 3,716 | 4,000 | 7.1% |
Relative risk (RR)
The ratio of risks (conditional probabilities, as in Probability Rules):
\[ RR = \frac{60/1,200}{224/2,800} = \frac{0.050}{0.080} = 0.625 \]
Participants carry 62.5% of the risk of non-participants: a relative reduction of 37.5%. And in absolute terms: the absolute risk reduction is \(8.0 - 5.0 = 3\) percentage points, which yields the NNT (number needed to treat): \(1/0.03 \approx 34\) — you need to enrol 34 employees in the programme to prevent one episode per year. The NNT is efficacy translated into resources, the healthcare equivalent of "turning the effect into euros".
Odds ratio (OR)
The odds of an event is \(p/(1-p)\) — the "3 to 1" form used in betting. The OR compares odds:
\[ OR = \frac{60/1,140}{224/2,576} = \frac{0.0526}{0.0870} = 0.605 \]
When is each one used?
| Relative risk | Odds ratio | |
|---|---|---|
| Interpretation | Direct ("half the risk") | Less intuitive (odds, not probabilities) |
| Requires | Being able to estimate incidences → cohort studies and trials | Only the table → also in case-control studies |
| Appears in | Cohort reports, trials | Case-control studies and logistic regression |
| With a rare event | — | OR ≈ RR |
Here the OR (0.605) and RR (0.625) nearly coincide because back pain with sick leave is relatively infrequent; with common events the OR exaggerates relative to the RR, and reading it as if it were an RR is a classic headline mistake. You already knew the OR without quite knowing it: the logistic regression on Club Nova churn returned an OR = 0.30 for active members — the same object, now with a birth certificate.
One caution before raising a glass: participation in the programme is voluntary. This is an observational study, and whoever signs up for preventive physiotherapy may have been taking better care of themselves already (the "healthy user bias") — the confounder from the previous lesson in its healthcare edition. The RR of 0.625 is promising, not probative; the study-design section shows what it would take to prove it.
Screening: sensitivity, specificity and the positive-result illusion
The insurer offers voluntary screening for a metabolic condition whose prevalence in the working population is 0.8%. The test has:
- Sensitivity = 90%: \(P(\text{positive} \mid \text{diseased})\) — out of every 100 people with the condition, it detects 90.
- Specificity = 95%: \(P(\text{negative} \mid \text{healthy})\) — out of every 100 healthy people, it gets 95 right (and falsely alarms 5).
The question that matters to whoever receives the result is the reverse one: "I tested positive — do I have it?" — that is, the positive predictive value \(P(\text{diseased} \mid \text{positive})\). Inverting conditionals is exactly Bayes' theorem, and the natural frequencies table makes it transparent. Imagine 10,000 employees screened:
| Positive | Negative | Total | |
|---|---|---|---|
| Diseased (0.8%) | 72 | 8 | 80 |
| Healthy | 496 | 9,424 | 9,920 |
| Total | 568 | 9,432 | 10,000 |
\[ PPV = \frac{72}{568} \approx 12.7\ % \qquad NPV = \frac{9,424}{9,432} \approx 99.9\ % \]
A counterintuitive and crucial result: with a "90/95%" test, nearly 9 out of every 10 positives are false alarms, because the disease is rare: 496 alarmed healthy people swamp the 72 detected cases. That is why:
- Population screening programmes chain a cheap, sensitive test with a more specific confirmatory test for the positives.
- The same test performs differently depending on who it is applied to: in a high-risk group with 10% prevalence, the PPV rises to \(0.9 \times 0.10 / (0.9 \times 0.10 + 0.05 \times 0.90) = 66.7\ %\). Prevalence — Bayes' prior probability — is in charge.
- Communicating it badly generates anxiety and overdiagnosis; communicating it well ("a positive means moving on to the confirmatory test, and most turn out to be nothing") is part of the screening design.
Study design: the hierarchy of evidence
How do we come to know whether something — the physiotherapy programme, a drug — works? Three major designs, from least to most control:
- Case-control study (retrospective): you start from the outcome — employees with back pain (cases) and without it (controls) — and look backwards at what exposures they had. Fast and cheap, ideal for rare conditions; but it lives under threat from recall bias and selection bias, and it can only estimate an OR (since you decide how many cases and controls to recruit, the real incidences are out of reach).
- Cohort study (prospective): you start from the exposure — participants and non-participants — and follow them over time counting new cases. It allows incidences and RR; it remains observational: confounding (the healthy user) does not disappear, it can only be adjusted for what has been measured.
- Randomized controlled trial (RCT): the exposure is assigned at random. Randomization spreads the confounders — measured and unmeasured — evenly across the groups: it is the A/B test from Statistics in Business in a white coat, and the only design that licenses the word "cause" without a footnote. Two reinforcements specific to this field: the placebo (the control group receives something indistinguishable, because mere expectation improves symptoms) and double-blind (neither participant nor assessor knows the group, so that hope contaminates neither the response nor its measurement).
The hierarchy of evidence ranks confidence: systematic reviews and meta-analyses of RCTs > individual RCT > cohort studies > case-control studies > case series > expert opinion. It is not a dogma — a small, clumsy RCT can be worth less than a huge, clean cohort — but it is the default map. For the wellness programme, the team's proposal to the insurer was straight out of the textbook: offer the next expansion by lottery among those interested (a randomized waiting list), turning the expansion into an RCT without denying the programme to anyone.
Survival analysis: reading a Kaplan-Meier curve
The insurer's reports include "episode-free survival curves". Despite the name, the technique works for the time until any event — death, relapse, first sick leave... or a Club Nova member churning. Its characteristic difficulty is censoring: employees who leave the company mid-follow-up, or who reach the end without an event. You do not know when it would have happened to them, but you do know it did not happen while you were watching — throwing that information away would bias the result, and the Kaplan-Meier estimator is built to use it.
The curve \(S(t)\) starts at 1 (everyone event-free) and steps down at each observed event. Reading the curve from the report (24-month follow-up, "back-episode-free survival"):
| Month | \(S(t)\) participants | \(S(t)\) non-participants |
|---|---|---|
| 6 | 0.98 | 0.96 |
| 12 | 0.95 | 0.92 |
| 18 | 0.93 | 0.87 |
| 24 | 0.91 | 0.83 |
Keys to reading it:
- Height at a given moment: at 24 months, 91% of participants remain episode-free versus 83% — consistent with the RR from the 2×2 table, but now you can see when the curves separate (here, mostly from month 12 onwards: prevention takes time to pay off).
- Median survival: the moment the curve crosses 0.50. Neither of these two does within 24 months ("median not reached" — good news); the 25.5 h median of the deliveries was, deep down, the same idea.
- Formal comparison: the usual test between curves is the log-rank (a relative of chi-square that compares observed and expected events at each moment); keep in mind that its p-value answers "are the curves distinguishable?", and that survival regression models (Cox) return hazard ratios, cousins of the RR.
- Sparse tails: at the end of follow-up few individuals remain at risk and the curve becomes unstable; serious reports print the "patients at risk" table underneath for exactly that reason.
Critical reading: headlines, relative and absolute risks
The lesson closes with the most transferable muscle: reading a health headline without being swept along.
Headline: "A supplement cuts the risk of a certain complication by 50%." The fine print: the complication drops from 2 cases to 1 per 10,000 person-years.
- Relative reduction: 50%. True and spectacular.
- Absolute reduction: 0.01 percentage points. \(NNT = 1/0.0001 = 10,000\) people treated for a year to prevent one case.
- With the supplement's price and its possible adverse effects multiplied by 10,000, the rational decision can flip completely. Rule: always demand the absolute risks; the relative one alone, without a base, is marketing.
A quick checklist for any health claim:
- Absolute risk and NNT, or only relative?
- Confidence interval or only a p-value? An RR of 0.60 with CI \((0.45;; 0.80)\) is a solid effect; with CI \((0.35;; 1.05)\) it is a "maybe" that crosses 1 (the null RR). The CI reports magnitude and precision; the asterisk, almost nothing — the old lesson from Hypothesis Testing.
- What design? An RCT, or a cohort with confounders lying in wait?
- Pre-specified primary outcome, or a finding among twenty subgroups? (With 20 tests at \(\alpha = 0.05\), one comes out on its own — the multiple comparisons problem.)
- Who was it measured in? An effect in 60-year-old men with prior disease does not extrapolate to a supermarket's workforce.
And the final warning, repeated on purpose: this judgment is for reading and for asking professionals better questions — not for self-medicating or for making clinical or organizational health decisions without them.
Common Mistakes and Tips
- Confusing prevalence with incidence. The snapshot does not measure the rate: a prevention programme is evaluated by incidence; judging it by short-term prevalence undervalues it.
- Reading an OR as if it were an RR. They only resemble each other with rare events; with common events the OR exaggerates. Always check which measure it is and how frequent the event is.
- Forgetting prevalence when interpreting a positive. Sensitivity and specificity are properties of the test; the PPV depends on who you apply it to. With low prevalence, most positives are false: a table of 10,000 and Bayes before anyone panics.
- Taking an observational study as causal. Voluntary participation = healthy user = confounding. Always ask about the design before asking about the p-value.
- Buying relative reductions without an absolute base. "Cuts it by 50%" can be 2→1 per 10,000. Demand the absolute risk and the NNT.
- Ignoring censoring. Excluding whoever dropped out of follow-up biases the curves; Kaplan-Meier exists precisely for that.
- Analysing health data as if they were receipts. Special category under the GDPR: anonymization, legal basis, and review by compliance and the medical service before the first query hits the database.
Exercises
Exercise 1
In a cohort of 1,000 office employees, 400 sit for more than 9 hours a day (exposed) and 600 do not. Episodes of neck pain over the year: 60 among the exposed and 54 among the unexposed. Calculate the risk in each group, the RR, the OR and the absolute reduction/increase. Why are the RR and OR similar?
Exercise 2
A screening test has 85% sensitivity and 92% specificity, and the prevalence is 2%. Using the 10,000-person table, calculate the PPV and interpret the result for someone who tests positive.
Exercise 3
Headline: "A new habit cuts the risk of a certain injury by 40%". The study: observational, a risk of 5 per 1,000 in non-practitioners versus 3 per 1,000 in practitioners, RR = 0.60 with 95% CI \((0.33;; 1.08)\). Give the full critique: absolutes, NNT, precision and design.
Solutions
Exercise 1.
Risks: \(60/400 = 15.0\ %\) and \(54/600 = 9.0\ %\). \(RR = 0.15/0.09 = 1.67\): 67% more relative risk. \(OR = (60/340)/(54/546) = 0.1765/0.0989 = 1.78\). Absolute increase: 6 percentage points (equivalent to a "number needed to harm" of \(1/0.06 \approx 17\)). RR and OR are similar (1.67 vs 1.78) because the event runs at around 10%; you can already see the OR inflating somewhat. Reminder: this is an observational cohort — sedentary work and neck pain share confounders (age, role, ergonomics) and this RR does not prove causation.
Exercise 2.
Out of 10,000: 200 diseased → 170 true positives, 30 false negatives; 9,800 healthy → 784 false positives, 9,016 true negatives. Total positives: 954. \(PPV = 170/954 \approx 17.8\ %\). Interpretation: even with a positive, the most likely outcome (82%) is not having the condition; the result means "move on to confirmation", not "diagnosis". Common mistake: answering 85% (the sensitivity) — that is the confusion of conditionals \(P(+\mid D) \neq P(D \mid +)\) that Bayes exists to undo.
Exercise 3.
(1) Absolutes: from 5 to 3 per 1,000 = an absolute reduction of 2 per 1,000; \(NNT = 1/0.002 = 500\) practitioner-years per injury prevented — the relative "40%" was far less glamorous in absolute terms. (2) Precision: the CI \((0.33;; 1.08)\) crosses 1: the data are compatible both with a sizeable reduction and with no effect at all (or a slight increase); in testing terms, not significant at 5%. (3) Design: observational — those who adopt the habit probably differ in everything else (healthy user). Verdict: an interesting hypothesis that will call for a better study, not a headline. Tip: this trio — absolutes, CI, design — dismantles or validates 90% of health headlines.
Conclusion
You can now read the healthcare language with your own tools: incidence and prevalence as film and snapshot; the wellness programme's RR of 0.625 with its NNT of 34, and the OR as the relative that survives in case-control studies and in the logistic regression you already used; the screening where 87% of positives were false — Bayes in its purest form; the hierarchy crowned by the randomized trial (your A/B test with placebo and double-blinding); the Kaplan-Meier curves that respect censoring; and the critical reading that demands absolutes and intervals before believing a headline. And two limits that are not optional: health data go through compliance, and healthcare decisions go through professionals.
That completes the sector tour. One last piece of the course remains, the most pragmatic of all: which tools — from the spreadsheet to Python, by way of jamovi and generative AI — are used to do all of this day to day, and how to choose the right one for each task. It is the end of the journey: Statistical Tools in Practice.
Statistics Course
Module 1: Introduction to Statistics
Module 2: Describing Data
- Measures of Central Tendency
- Measures of Dispersion
- Measures of Position and Outliers
- Graphical Representation of Data
Module 3: Probability
Module 4: Probability Distributions
- The Binomial Distribution
- The Normal Distribution
- Other Important Distributions
- The Central Limit Theorem
Module 5: Statistical Inference
Module 6: Data Analysis
- Correlation Analysis
- Regression Analysis
- Analysis of Variance (ANOVA)
- Categorical Data Analysis: Chi-Square
