This module has given you templates for individual variables: a receipt, a delivery, a count of incidents. But the analyst's daily work almost never revolves around an individual data point — it revolves around sample averages and proportions: the average receipt over 640 purchases, the 61% card-payment rate computed over thousands of receipts, the 8.1 satisfaction score over 640 responses. Here arises the question Marta has been dodging since Module 1, when the committee asked about the survey: if we had collected another 640 responses, would a different number have come out? How much can it swing? The answer demands a shift in perspective — treating the sample mean itself as a random variable — and culminates in the most important result in all of statistics: the central limit theorem (CLT), which explains why the normal bell shows up everywhere and why we can trust averages even when the underlying data are skewed and unruly. It is the bridge connecting everything learned so far to the inference of Module 5.
Contents
- The shift in perspective: the sample mean is a random variable
- The sampling distribution of the mean: center and standard error
- The conceptual experiment: receipt averages with growing samples
- The statement of the CLT and its intuition
- Calculating with the CLT: probabilities about means
- The normal approximation to the binomial and to proportions
- The bridge to inference: the committee's question, ready to be answered
The shift in perspective: the sample mean is a random variable
Until now, "the mean" was a number you computed, end of story (Module 2). The shift in perspective is this: before the sample is taken, the sample mean \( \bar{X} \) is uncertain — it depends on which 640 customers respond, on which 100 receipts land in the sample. And an uncertain quantity with numerical values is, exactly, a random variable like those of Module 3. It therefore has its own probability distribution, called the sampling distribution of the mean: the distribution of the values \( \bar{X} \) would take across all possible samples.
It pays to have the Module 1 vocabulary sharp, because three levels coexist here and must not be mixed:
| Level | Object | NovaMarket example | Random? |
|---|---|---|---|
| Population | Parameter \( \mu, \sigma \) | The true average receipt over all receipts: \( \mu = \text{€}32.40 \), \( \sigma = \text{€}21.50 \) | No (fixed, though unknown in practice) |
| One concrete sample | Computed statistic \( \bar{x} \) | The mean of these 100 receipts: €33.10 | No (already observed) |
| Before sampling | Random variable \( \bar{X} \) | The mean of the 100 receipts yet to come | Yes — and its distribution is today's topic |
The sampling distribution of the mean: center and standard error
Two exact results (valid whenever the sample is random and the observations independent, whatever the shape of the population):
\[ E(\bar{X}) = \mu \qquad\qquad \sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}} \]
- The center does not move: on average, the sample mean neither overestimates nor underestimates \( \mu \). The "high" and "low" samples cancel out across all possible samples.
- The spread shrinks with \( \sqrt{n} \). \( \sigma_{\bar{X}} \) is called the standard error of the mean, and it is the most important number in this lesson: it measures how much the sample mean fluctuates from one sample to another. When averaging, large and small receipts partially cancel, and the larger \( n \) is, the more effective the cancellation.
With NovaMarket's receipts (\( \mu = 32.40 \), \( \sigma = 21.50 \)):
| Sample size \( n \) | Standard error \( \sigma/\sqrt{n} \) | Reading |
|---|---|---|
| 1 | €21.50 | A single receipt is a wild data point |
| 4 | €10.75 | Doubling the precision requires quadrupling \( n \) |
| 25 | €4.30 | |
| 100 | €2.15 | The mean of 100 receipts typically fluctuates ±€2 |
| 400 | €1.08 | |
| 640 | €0.85 | The size of our satisfaction survey |
Note the square-root law: going from 100 to 400 observations (×4) only halves the error. Extra precision gets ever more expensive — a principle that governs the budget of every market study and will reappear when computing sample sizes in Module 5.
The conceptual experiment: receipt averages with growing samples
Recall something crucial from the previous lesson: the receipt amount is not normal — it is strongly right-skewed (many small receipts, a tail of big shopping carts). Now imagine a thought experiment any analyst can run on the receipt history: draw 200 random samples of size \( n \), compute the mean of each one, and look at the histogram of those 200 means. Then repeat with growing \( n \). Typical results of the experiment:
With \( n = 5 \) (theoretical standard error \( 21.50/\sqrt{5} \approx \text{€}9.60 \)): the 200 means are still scattered and right-tailed — a single €150 cart is enough to send its sample's mean flying:
| Sample mean (€) | 10–20 | 20–30 | 30–40 | 40–50 | 50–60 | 60–70 |
|---|---|---|---|---|---|---|
| Number of samples | 24 | 78 | 56 | 26 | 12 | 4 |
With \( n = 50 \) (standard error \( \approx \text{€}3.00 \)): the distribution tightens around 32.40 and the skewness has all but disappeared:
| Sample mean (€) | 26–29 | 29–31 | 31–33 | 33–35 | 35–37 | 37–40 |
|---|---|---|---|---|---|---|
| Number of samples | 8 | 42 | 74 | 52 | 20 | 4 |
With \( n = 500 \) (standard error \( \approx \text{€}0.96 \)): a narrow, symmetric bell, indistinguishable from a normal centered at 32.40:
| Sample mean (€) | 30.0–31.0 | 31.0–31.8 | 31.8–32.6 | 32.6–33.4 | 33.4–34.2 | 34.2–35.2 |
|---|---|---|---|---|---|---|
| Number of samples | 6 | 30 | 66 | 64 | 28 | 6 |
Two phenomena at once, and it is worth telling them apart: (1) the spread shrinks according to \( \sigma/\sqrt{n} \) — we already knew that from the previous section; and (2) the shape becomes normal, even though individual receipts are nowhere near normal. This second phenomenon is the content of the theorem.
The statement of the CLT and its intuition
Central limit theorem (practical statement, no proof): if \( \bar{X} \) is the mean of a random sample of \( n \) independent observations from a population with mean \( \mu \) and (finite) standard deviation \( \sigma \), then, for sufficiently large \( n \),
\[ \bar{X} ;\approx; N!\left(\mu,\ \frac{\sigma}{\sqrt{n}}\right) \]
whatever the shape of the population's distribution. The approximation improves as \( n \) grows; the usual rule of thumb is \( n \geq 30 \), though with heavily skewed populations (like the receipts) somewhat more is advisable, and if the population is already normal, \( \bar{X} \) is exactly normal for any \( n \).
The intuition connects with what we said at the start of the normal-distribution lesson: a mean is a sum of many small, independent contributions (each observation contributes \( x_i/n \)). When summing, individual oddities get diluted: for the mean of 500 receipts to come out sky-high, one big cart is not enough — it would take many receipts being large simultaneously, and independence makes that consensus ever more improbable. The result is the bell: central values reachable by a thousand different routes, extreme values demanding massive coincidences.
That is why the normal "shows up everywhere": daily sales (a sum of receipts), measurement errors (a sum of perturbations), aggregate indicators of every kind. The CLT is the theorem that manufactures normals out of almost anything.
Calculating with the CLT: probabilities about means
With the CLT, questions about sample means are answered with the previous lesson's machinery — standardize, Z table, translate — using the standard error in the denominator:
\[ z = \frac{\bar{x} - \mu}{\sigma / \sqrt{n}} \]
Example 1 — auditing a checkout. \( n = 100 \) receipts from Valencia-Ruzafa will be audited at random. If the store behaves like the chain (\( \mu = 32.40 \), \( \sigma = 21.50 \)), what is the probability that the sample mean exceeds €35?
- Standard error: \( 21.50 / \sqrt{100} = \text{€}2.15 \).
- Standardize: \( z = \dfrac{35 - 32.40}{2.15} \approx 1.21 \).
- Table and translation: \( P(\bar{X} > 35) = 1 - \Phi(1.21) = 1 - 0.8869 = 0.1131 \).
11%: a sample mean of €35 would be high but not scandalous. Notice what the CLT has handed us for free: receipts are not normal, and yet the calculation is legitimate because with \( n = 100 \) the mean is (approximately).
Example 2 — the same threshold with a bigger sample. With \( n = 400 \): standard error \( 21.50/20 = \text{€}1.08 \), \( z = 2.60/1.08 \approx 2.41 \), and \( P(\bar{X} > 35) \approx 1 - 0.9920 = 0.008 \). The same €2.60 deviation that was routine with 100 receipts is an extreme rarity with 400: the larger the sample, the rarer it is for the mean to drift away from \( \mu \) by pure chance. Hold on to that sentence: it is the exact mechanism by which the tests of Module 5 will work.
The normal approximation to the binomial and to proportions
The CLT also closes the circle opened in the first lesson of the module: a binomial is a sum of \( n \) independent Bernoulli variables, so the CLT applies to it head-on. For large \( n \),
\[ X \sim B(n, p) ;\approx; N!\left(np,\ \sqrt{np(1-p)}\right) \]
Practical validity rule: \( np \geq 10 \) and \( n(1-p) \geq 10 \) (some textbooks accept 5; with 10 you are on safe ground).
Example — card-paid receipts at store scale. Of the next \( n = 200 \) receipts at Bilbao-Casco, what is the probability that 130 or more are paid by card (\( p = 0.61 \))? Summing 71 binomial terms is unworkable by hand; the approximation solves it:
- Validity: \( np = 122 \geq 10 \) and \( n(1-p) = 78 \geq 10 \). ✔
- Parameters: mean \( 122 \), standard deviation \( \sqrt{200 \times 0.61 \times 0.39} = \sqrt{47.58} \approx 6.90 \).
- Continuity correction: the binomial is discrete (steps) and the normal continuous; the event "\( X \geq 130 \)" is better represented by the area from 129.5 (the step for 130 begins there). It is a fine-tuning adjustment that improves the approximation; knowing it and applying it when precision matters is enough.
- \( z = \dfrac{129.5 - 122}{6.90} \approx 1.09 \), so \( P(X \geq 130) \approx 1 - \Phi(1.09) = 1 - 0.8621 = 0.1379 \).
In terms of proportions, which is how these results get communicated, dividing by \( n \) directly gives the sampling distribution of the proportion \( \hat{p} \):
\[ \hat{p} ;\approx; N!\left(p,\ \sqrt{\frac{p(1-p)}{n}}\right) \]
With \( n = 200 \) and \( p = 0.61 \): standard error \( \sqrt{0.61 \times 0.39 / 200} \approx 0.0345 \). In other words: the "61% card payment" measured over 200 receipts typically swings ±3.5 percentage points from sample to sample — and over 2,000 receipts, only ±1.1 points. A proportion is a mean in disguise (a mean of ones and zeros), so it inherits this lesson's entire theory, square-root law included.
The bridge to inference: the committee's question, ready to be answered
Let us return to the question that opened the course. The satisfaction survey: 2,000 invitations, 640 responses, mean 8.1 with \( s = 1.6 \). The committee asks: "how much can we trust the 8.1?". We now have all the pieces:
- The 8.1 is one realization of the random variable \( \bar{X} \): another 640 responses would have produced a different number.
- Its standard error is \( \sigma/\sqrt{640} \approx 1.6/25.3 \approx 0.063 \) points.
- By the CLT, \( \bar{X} \) is approximately normal. With the 68-95-99.7 rule: if the true satisfaction of the customer population were \( \mu \), the mean of 640 responses would land within \( 2 \times 0.063 \approx 0.13 \) points of \( \mu \) in 95% of samples.
That is: the sampling fluctuation of the 8.1 is on the order of ±0.1 — small, but not zero. What remains to answer the committee in full is to invert the reasoning: to go from "if I knew \( \mu \), I know how \( \bar{X} \) swings" to "I have observed \( \bar{x} = 8.1 \): what can I claim about the unknown \( \mu \), and with how much confidence?". That inversion — with its subtleties, such as using \( s \) in place of the unknown \( \sigma \), and without forgetting that the 32% non-response rate can bias more than chance does (Module 1) — is exactly statistical inference, and it begins in the next lesson.
Common Mistakes and Tips
- Confusing \( \sigma \) with \( \sigma/\sqrt{n} \). The standard deviation of the data measures how much one receipt varies; the standard error measures how much the sample mean varies. Using 21.50 where 2.15 belongs multiplies any tail probability tenfold. Always ask yourself: is my question about an individual or about a mean?
- Believing the CLT normalizes the data. The CLT speaks about the sample mean, not the observations: no matter how many receipts you gather, their histogram will remain skewed. What becomes normal is the distribution of the means across samples.
- Applying \( n \geq 30 \) as dogma. It is a guideline: with nearly symmetric populations less is enough; with strong skewness or extreme outliers more is advisable. And with dependent data (receipts from the same family, sales on consecutive days) the CLT as we have stated it does not apply: independence is part of the contract.
- Forgetting to check \( np \geq 10 \) and \( n(1-p) \geq 10 \) before approximating a binomial. With \( p = 0.005 \) (fraud) and \( n = 400 \), \( np = 2 \): there, the right approximation was the Poisson of the previous lesson, not the normal.
- Ignoring the continuity correction in discrete counts when precision is needed: without it, \( P(X \geq 130) \) came out as 0.123 instead of 0.138. For communicating orders of magnitude it can be skipped; for decisions to the millimeter, it cannot.
- Tip: memorize the square-root law (×4 sample → ÷2 error) so you can react in meetings: when someone proposes "doubling the survey to double the precision", you already know it does not work that way.
Exercises
Exercise 1. With NovaMarket's receipts (\( \mu = \text{€}32.40 \), \( \sigma = \text{€}21.50 \)): (a) Compute the standard error of the mean for samples of 64 and of 256 receipts, and comment on the relationship between the two results. (b) For \( n = 256 \), compute \( P(\bar{X} > 34) \) (\( \Phi(1.19) = 0.8830 \)). (c) Why is it valid to use the normal in (b) if receipt amounts are skewed?
Exercise 2. 45% of receipts include fresh produce. In a sample of 100 receipts from Sevilla-Nervión: (a) Check whether the normal approximation to the binomial applies. (b) Compute, with continuity correction, the probability that 40 or fewer receipts include fresh produce (\( \Phi(0.90) = 0.8159 \)). (c) Express the standard error of the sample proportion in percentage points.
Exercise 3. About the satisfaction survey (\( n = 640 \), \( s = 1.6 \)): (a) Compute the standard error of the mean. (b) Assuming the population mean really were 8.1, between which two values would the sample mean fall in 95% of surveys with 640 responses? (c) Marketing proposes repeating the survey with only 160 responses "because the mean will come out the same anyway". What would you reply, with numbers?
Solutions
Solution 1. (a) \( 21.50/\sqrt{64} = 21.50/8 \approx \text{€}2.69 \) and \( 21.50/\sqrt{256} = 21.50/16 \approx \text{€}1.34 \). Quadrupling the sample has cut the error exactly in half: the square-root law. (b) \( z = \frac{34 - 32.40}{1.34} \approx 1.19 \); \( P = 1 - 0.8830 = 0.1170 \). (c) By the CLT: with \( n = 256 \), well above 30, the sampling distribution of the mean is approximately normal even though the population is not. Common mistake: in (b), using \( \sigma = 21.50 \) in the denominator, which would give \( z = 0.07 \) and an absurd probability of 47%.
Solution 2. (a) \( np = 45 \geq 10 \) and \( n(1-p) = 55 \geq 10 \): it applies. (b) Mean \( 45 \), standard deviation \( \sqrt{100 \times 0.45 \times 0.55} = \sqrt{24.75} \approx 4.97 \). With continuity correction, "\( X \leq 40 \)" is evaluated at 40.5: \( z = \frac{40.5 - 45}{4.97} \approx -0.90 \), so \( P \approx 1 - 0.8159 = 0.1841 \). (c) \( \sqrt{0.45 \times 0.55/100} \approx 0.0497 \): the sample proportion typically fluctuates about ±5 percentage points with 100 receipts. Common mistake: using 39.5 instead of 40.5 — the step for "up to and including 40" ends at 40.5.
Solution 3. (a) \( 1.6/\sqrt{640} \approx 1.6/25.3 \approx 0.063 \) points. (b) \( 8.1 \pm 2 \times 0.063 \): between 7.97 and 8.23 approximately (the 95% rule). (c) With \( n = 160 \), the standard error would be \( 1.6/\sqrt{160} \approx 0.126 \): double. The mean "would come out similar", yes, but with a typical fluctuation of ±0.25 points at 95% — enough to confuse an 8.1 with a 7.9 or an 8.3, exactly the differences the committee wants to monitor. Analyst's nuance: beyond chance, both surveys still carry the possible non-response bias from Module 1, which no increase in \( n \) corrects.
Conclusion
The central limit theorem crowns the module and gives it retrospective meaning: the sample mean is a random variable centered at \( \mu \), with standard error \( \sigma/\sqrt{n} \) and — here is the miracle — normal in shape whatever the starting population; proportions, being means in disguise, inherit the same behavior, and the large-\( n \) binomial surrenders to the bell. With that, the Module 4 toolbox closes, and something deeper happens: for the first time we know how much a statistic swings around the parameter it chases — the 8.1 satisfaction score swings ±0.06, the 61% card rate over 200 receipts swings ±3.5 points. All that remains is to turn the telescope around: to stop deducing the sample from the population and start inferring the population from the sample, which is what Marta needs to answer the committee. That turn — estimating unknown parameters and quantifying the confidence of the estimate — begins in Module 5 with Parameter Estimation.
Statistics Course
Module 1: Introduction to Statistics
Module 2: Describing Data
- Measures of Central Tendency
- Measures of Dispersion
- Measures of Position and Outliers
- Graphical Representation of Data
Module 3: Probability
Module 4: Probability Distributions
- The Binomial Distribution
- The Normal Distribution
- Other Important Distributions
- The Central Limit Theorem
Module 5: Statistical Inference
Module 6: Data Analysis
- Correlation Analysis
- Regression Analysis
- Analysis of Variance (ANOVA)
- Categorical Data Analysis: Chi-Square
