In the Errors, Power and Sample Size lesson we left a debt outstanding: when we tried to compare NovaMarket's 42 stores with pairwise t tests, we discovered that multiplying comparisons sends the probability of false alarms through the roof, and we promised a tool that compares many groups at once with a single test. That tool is the Analysis of Variance (ANOVA), and its logic is an elegant paradox: to compare means, it analyzes variances — specifically, it compares how much the groups vary among themselves with how much people vary within each group. In this lesson we will build that logic from scratch, compute a complete ANOVA table by hand using satisfaction data from three store formats, learn to read the F distribution table, review the method's assumptions, and see what to do after rejecting: post-hoc comparisons and effect size.
Contents
- Why not run many t tests: the debt from Module 5
- The logic of ANOVA: between-group vs within-group variability
- Hypotheses and case data: satisfaction across three store formats
- The ANOVA table step by step: SS, df, MS and F
- The F distribution and how to read its table
- ANOVA assumptions (the working version)
- After rejecting: post-hoc comparisons (Tukey and Bonferroni)
- How big is the effect? Eta-squared
- Where this lesson stops: two factors and non-parametric alternatives
Why not run many t tests: the debt from Module 5
Marta wants to know whether customer satisfaction depends on the store format. NovaMarket operates three formats: Express (urban, small), Standard (neighborhood supermarket) and Superstore. Comparing three groups with t tests requires three comparisons (Express–Standard, Express–Superstore, Standard–Superstore), and we already know what that means: if each test carries a 5% false-alarm risk, the probability that at least one of the three throws a false positive is
\[ 1 - 0.95^3 = 1 - 0.857 = 0.143 \]
that is, 14.3% — almost triple the agreed 5%. With more groups the problem explodes (recall the 42-store thought experiment from Module 5, with an 88% probability of at least one false alarm). ANOVA solves the problem at the root: it states one single global hypothesis with one single α:
- \( H_0: \mu_{\text{Express}} = \mu_{\text{Standard}} = \mu_{\text{Superstore}} \) — all means are equal; format has no influence
- \( H_1: \) at least one mean is different
Mind the subtlety in \( H_1 \): rejecting \( H_0 \) says "they are not all equal", not which ones differ. That second step comes later, with the post-hoc comparisons.
The logic of ANOVA: between-group vs within-group variability
How can variances detect whether means differ? Picture two scenarios with the same three sample means (7, 8 and 9):
- Scenario A: within each format, the stores score almost identically (all Express stores between 6.8 and 7.2…). The differences between groups (7 vs 8 vs 9) stand out crisply against that internal silence: they look real.
- Scenario B: within each format there is a bit of everything (Express stores ranging from 4 to 10…). Means of 7, 8 and 9 could perfectly well be the fruit of sampling chance: the signal drowns in the noise.
ANOVA formalizes that intuition with a signal-to-noise ratio:
\[ F = \frac{\text{variability BETWEEN groups (signal: what separates the means)}}{\text{variability WITHIN groups (noise: how much people vary by chance)}} \]
- If \( H_0 \) is true, both variabilities estimate the same thing (the population's natural variance) and F will hover around 1.
- If some group has a different mean, the between variability inflates and F climbs well above 1.
The question "how large must F be before we reject?" will be answered by the F distribution. First, the calculations.
Hypotheses and case data: satisfaction across three store formats
Marta's team takes a sample of 5 stores per format and computes the average customer satisfaction of each store (0–10 scale, rounded to whole numbers for the hand calculation):
| Format | Scores | Sum | Mean \( \bar{x}_j \) |
|---|---|---|---|
| Express | 7, 8, 6, 7, 7 | 35 | 7.0 |
| Standard | 8, 9, 8, 7, 8 | 40 | 8.0 |
| Superstore | 9, 8, 9, 10, 9 | 45 | 9.0 |
In total \( N = 15 \) observations, \( k = 3 \) groups, and the grand mean is \( \bar{x} = (35+40+45)/15 = 120/15 = 8.0 \). The sample means differ (7, 8, 9), but we already know that alone is not enough: we must weigh the signal against the noise.
The ANOVA table step by step: SS, df, MS and F
The heart of ANOVA is the same variability decomposition we saw in regression: the total sum of squares splits into two pieces.
\[ \underbrace{\sum_{j}\sum_{i} (x_{ij} - \bar{x})^2}{\text{Total SS}} ;=; \underbrace{\sum{j} n_j (\bar{x}j - \bar{x})^2}{\text{Between SS}} ;+; \underbrace{\sum_{j}\sum_{i} (x_{ij} - \bar{x}j)^2}{\text{Within SS}} \]
Step 1 — Between SS (each group mean against the grand mean, weighted by its size):
\[ \text{SS}_{\text{Between}} = 5(7-8)^2 + 5(8-8)^2 + 5(9-8)^2 = 5 + 0 + 5 = 10 \]
Step 2 — Within SS (each observation against the mean of its own group):
| Format | Deviations from its mean | Squares | Sum |
|---|---|---|---|
| Express (mean 7) | 0, +1, −1, 0, 0 | 0, 1, 1, 0, 0 | 2 |
| Standard (mean 8) | 0, +1, 0, −1, 0 | 0, 1, 0, 1, 0 | 2 |
| Superstore (mean 9) | 0, −1, 0, +1, 0 | 0, 1, 0, 1, 0 | 2 |
\[ \text{SS}{\text{Within}} = 2 + 2 + 2 = 6 \qquad\qquad \text{SS}{\text{Total}} = 10 + 6 = 16 ; \checkmark \]
Step 3 — Degrees of freedom. Between: \( k - 1 = 2 \). Within: \( N - k = 12 \). (Total: \( N - 1 = 14 = 2 + 12 \), the usual sanity check.)
Step 4 — Mean squares: each SS divided by its df, which turns it into a comparable variance:
\[ \text{MS}{\text{Between}} = \frac{10}{2} = 5.0 \qquad\qquad \text{MS}{\text{Within}} = \frac{6}{12} = 0.5 \]
Step 5 — The F statistic and the complete table, in the standard format used by any report or software package:
| Source of variation | SS | df | MS | F |
|---|---|---|---|---|
| Between groups (format) | 10 | 2 | 5.0 | 10.0 |
| Within groups (residual) | 6 | 12 | 0.5 | |
| Total | 16 | 14 |
The signal is ten times the noise. Is that enough to reject?
The F distribution and how to read its table
Under \( H_0 \), the ratio \( F = \text{MS}{\text{Between}}/\text{MS}{\text{Within}} \) follows the F distribution (Fisher's), which is new to the course, so let's introduce it properly:
- It is the distribution of a ratio of two variances: it only takes positive values and is right-skewed.
- It depends on two degrees of freedom: the numerator's (between groups, here 2) and the denominator's (within, here 12). It is written \( F_{2;,12} \).
- The test is always right-tailed: only large F values (lots of signal relative to noise) contradict \( H_0 \).
How to read the F table: unlike the t table (one row per df), the F table is a two-way grid for a fixed α (there is one table for α = 0.05, another for 0.01…). In the α = 0.05 table: find the column for the numerator df (2) and the row for the denominator df (12); the intersection gives the critical value. An excerpt:
| Denominator df \ Numerator df | 1 | 2 | 3 | 4 |
|---|---|---|---|---|
| 10 | 4.96 | 4.10 | 3.71 | 3.48 |
| 12 | 4.75 | 3.89 | 3.49 | 3.26 |
| 14 | 4.60 | 3.74 | 3.34 | 3.11 |
The critical value is \( F_{0.05;,2;,12} = 3.89 \). Since our \( F = 10.0 > 3.89 \), we reject \( H_0 \) at the 5% level (in fact at the 1% level too: \( F_{0.01;,2;,12} = 6.93 \); the p-value is < 0.01). Conclusion for Marta: satisfaction is not the same across the three formats — store format matters. Be careful with the order of the df: \( F_{2;12} \neq F_{12;2} \); numerator always first.
ANOVA assumptions (the working version)
Like every test, ANOVA rests on assumptions. In practical terms:
| Assumption | What it requires | How to check it in practice |
|---|---|---|
| Independence | Observations do not influence one another | Guaranteed by design: random sampling, distinct stores, never measuring the same customer twice. It is the most important assumption and the only irreparable one |
| Normality | Each group's data come from ~normal populations | Histograms/boxplots per group with no wild skew. With large groups, the CLT from Module 4 relaxes this requirement considerably |
| Homogeneity of variances (homoscedasticity) | The groups' variances are similar | Quick rule: the largest standard deviation should not double the smallest. Here the three sample variances are identical (0.5): impeccable |
If normality fails badly with small samples, or the data are ordinal, there is a rank-based alternative (the Kruskal-Wallis test), which we will meet among the Non-Parametric Methods.
After rejecting: post-hoc comparisons (Tukey and Bonferroni)
A significant F says "not all the means are equal", but Marta will immediately ask: "which ones differ?". Answering requires going back to pairwise comparisons — but now with protection against Type I error inflation. The two most common strategies:
- Bonferroni correction (we already know it from Module 5): split α across the \( m \) comparisons, requiring each one to meet \( \alpha' = \alpha/m \). With 3 comparisons, \( \alpha' = 0.05/3 = 0.0167 \). Simple and always valid, though conservative.
- Tukey's HSD test: designed specifically for "all pairs after an ANOVA". It computes an honestly significant difference (HSD) from the Within MS and the number of groups: any pair of means differing by more than the HSD is declared different, keeping the overall error at 5% for the whole set of comparisons. It is the standard option in statistical software and somewhat more powerful than Bonferroni for this use.
Let's apply Bonferroni to our case. The standard error of a difference of means, using the Within MS as the common variance, is
\[ \text{SE} = \sqrt{\text{MS}_{\text{Within}}\left(\tfrac{1}{n_1}+\tfrac{1}{n_2}\right)} = \sqrt{0.5 \times \tfrac{2}{5}} = \sqrt{0.2} = 0.447 \]
| Comparison | Difference | t = diff/SE | Bonferroni critical \( t_{0.0167/2;,12} \approx 2.78 \) | Conclusion |
|---|---|---|---|---|
| Express vs Superstore | 2.0 | 4.47 | 2.78 | They differ |
| Express vs Standard | 1.0 | 2.24 | 2.78 | Inconclusive |
| Standard vs Superstore | 1.0 | 2.24 | 2.78 | Inconclusive |
(Tukey reaches the same conclusion here.) The result is very instructive: the global ANOVA is emphatic, and the post-hoc pinpoints the difference between the extreme formats (Express vs Superstore), while the 1-point differences between neighboring formats, with only 5 stores per group, do not reach corrected significance — they are suggestive, but more stores would be needed to confirm them (power, as we saw in 05-04). An honest report for Marta: "format influences satisfaction; the proven gap is between Express (7.0) and Superstore (9.0); the intermediate comparisons point in the same direction but need a larger sample".
How big is the effect? Eta-squared
Significant, yes — but important? Just as in Module 5 we distinguished significance from relevance, ANOVA has its own effect size measure: eta-squared, the proportion of the total variability attributable to the factor:
\[ \eta^2 = \frac{\text{SS}{\text{Between}}}{\text{SS}{\text{Total}}} = \frac{10}{16} = 0.625 \]
Store format explains 62.5% of the variability in satisfaction in the sample — a large effect (roughly: ~0.01 small, ~0.06 medium, ~0.14 large). Notice the exact parallel with regression's R²: same concept, categorical factor instead of continuous variable.
Where this lesson stops: two factors and non-parametric alternatives
What we have built here is one-way ANOVA (a single grouping variable). There is also two-way ANOVA, which analyzes, say, format and region simultaneously — including the possibility of an interaction (the format effect differing by region) —; its mechanics extend the same sum-of-squares decomposition and lie beyond our introductory scope. And when the assumptions do not hold, the rank-based alternative (Kruskal-Wallis) awaits among the Non-Parametric Methods.
Common Mistakes and Tips
- Running pairwise t tests without correction instead of (or after) an ANOVA. It is the false-alarm factory from Module 5; ANOVA exists precisely to avoid it.
- Reading rejection as "all the means differ". \( H_1 \) only asserts that at least one differs; without a post-hoc you don't know which. And conversely: failing to reject does not prove they are equal.
- Flipping the degrees of freedom when reading the F table. \( F_{2;12} = 3.89 \) but \( F_{12;2} = 19.4 \): numerator (between) first, always.
- Looking at the left tail. The ANOVA F test is right-tailed; an F close to 0 does not "prove equality" — it only signals abnormally similar groups (sometimes a symptom of manipulated or dependent data).
- Forgetting to check the variances. If one group's spread is much larger (quick rule: max s > 2 × min s), F loses reliability; consider transforming the data or using the non-parametric method.
- Stopping at the p-value with no effect size. With enormous samples, trivial differences produce significant F values. Always accompany the verdict with \( \eta^2 \) and the group means in real units.
Exercises
Exercise 1. Operations trials three checkout staffing configurations and measures the average waiting time (minutes) in 4 stores per configuration:
| Configuration | Times | Mean |
|---|---|---|
| A (current) | 4, 5, 6, 5 | 5 |
| B (peak-hour reinforcement) | 6, 7, 7, 8 | 7 |
| C (express lane) | 5, 6, 6, 7 | 6 |
Build the complete ANOVA table and decide at the 5% level whether the configuration matters (\( F_{0.05;,2;,9} = 4.26 \)). Also compute \( \eta^2 \).
Exercise 2. A junior analyst, instead of the ANOVA in exercise 1, runs the three possible t tests at the 5% level and presents as a finding the only one that came out significant. Explain the two methodological errors and compute the probability of at least one false alarm in his approach, assuming \( H_0 \) were true in all three comparisons.
Exercise 3. A report arrives with this incomplete ANOVA table on the average spend of Club Nova members across 4 tenure segments (6 members per segment, N = 24):
| Source | SS | df | MS | F |
|---|---|---|---|---|
| Between groups | 90 | ? | ? | ? |
| Within groups | ? | ? | ? | |
| Total | 210 | ? |
Fill in the blanks and decide at the 5% level (\( F_{0.05;,3;,20} = 3.10 \)).
Solutions
Exercise 1. Grand mean = (20+28+24)/12 = 6. \( \text{SS}{\text{Between}} = 4(5-6)^2 + 4(7-6)^2 + 4(6-6)^2 = 8 \). Within: A: 1+0+1+0 = 2; B: 1+0+0+1 = 2; C: 1+0+0+1 = 2; \( \text{SS}{\text{Within}} = 6 \). Table: SS 8 and 6; df 2 and 9; MS 4 and 0.667; \( F = 4/0.667 = 6.0 \). Since 6.0 > 4.26, we reject \( H_0 \): configuration influences waiting time. \( \eta^2 = 8/14 = 0.57 \). (For action purposes, the post-hoc would point to the A vs B difference — and mind the business reading: the "reinforcement" configuration B is the one showing MORE waiting!) Common mistake: forgetting to multiply by \( n_j = 4 \) in the Between SS, or using N − 1 = 11 as the denominator df instead of N − k = 9.
Exercise 2. Errors: (1) Type I error inflation — three tests at 5% accumulate \( 1 - 0.95^3 = 14.3% \) probability of at least one false alarm, almost triple the agreed level; (2) selective reporting (p-hacking, seen in 05-04): presenting only the significant test hides the search process and turns possible chance into a "finding". The right approach: global ANOVA first and, only if it rejects, a corrected post-hoc (Bonferroni: each comparison at 0.05/3 = 0.0167), reporting all three. Common mistake: believing the correction is optional when there are "only" 3 comparisons — 14.3% already nearly triples the nominal risk.
Exercise 3. \( \text{SS}_{\text{Within}} = 210 - 90 = 120 \). df: between = 3, within = 20, total = 23. MS: between = 90/3 = 30; within = 120/20 = 6. \( F = 30/6 = 5.0 \). Since 5.0 > 3.10, we reject \( H_0 \): average spend is not the same across the four tenure segments (with \( \eta^2 = 90/210 = 0.43 \), a large effect). A post-hoc would still be needed to know which segments differ. Common mistake: computing the between df as k = 4 instead of k − 1 = 3, which throws off all the MS values and F in cascade.
Conclusion
ANOVA settles the debt from Module 5: instead of a dangerous battery of t tests, a single global test that compares the between-group variability with the within-group variability via the F statistic — in our case, F = 10 against a critical value of 3.89: store format influences satisfaction, the effect is large (η² = 0.625) and the Bonferroni post-hoc pinpoints the proven difference between Express and Superstore. We have added the F distribution and its two-way table to our arsenal, along with the discipline of checking assumptions and pairing every p-value with its effect size. But ANOVA demands a quantitative response variable (satisfaction, times, euros). What about when what we want to cross-tabulate are two categorical variables — purchase channel and payment method, age bracket and membership churn —, where there are no means to compare but frequencies to count? For that we need the chi-square test, the star of the Categorical Data Analysis lesson that closes the module.
Statistics Course
Module 1: Introduction to Statistics
Module 2: Describing Data
- Measures of Central Tendency
- Measures of Dispersion
- Measures of Position and Outliers
- Graphical Representation of Data
Module 3: Probability
Module 4: Probability Distributions
- The Binomial Distribution
- The Normal Distribution
- Other Important Distributions
- The Central Limit Theorem
Module 5: Statistical Inference
Module 6: Data Analysis
- Correlation Analysis
- Regression Analysis
- Analysis of Variance (ANOVA)
- Categorical Data Analysis: Chi-Square
