In the previous module we turned the telescope around: we stopped deducing how samples behave from a known population and got ready to do exactly the opposite — infer the population from a sample. This lesson takes the first step down that road: parameter estimation. You will learn what a point estimator is, which properties make an estimator "good" (unbiasedness, efficiency, consistency), why the standard error measures the quality of an estimate, and what changes when we do not know the population standard deviation — the moment a new protagonist takes the stage: Student's t distribution. Our running case will be a question Marta has been wanting to answer for weeks: how much does a Club Nova member spend per month, on average, and what proportion of members is actually active?
Contents
- From statistic to estimator: point estimation
- The Club Nova case: estimating average spend and the proportion of active members
- Desirable properties of an estimator: unbiasedness, efficiency and consistency
- The sampling distribution of the estimator and the standard error
- When σ is unknown: Student's t distribution
- How to read the t table
- Other estimation methods (a brief mention)
From statistic to estimator: point estimation
Let us recall the vocabulary from Module 1:
- A parameter is a numerical characteristic of the population: the population mean \( \mu \), the population proportion \( p \), the population variance \( \sigma^2 \). We almost never know it.
- A statistic is a numerical characteristic of the sample: the sample mean \( \bar{x} \), the sample proportion \( \hat{p} \), the sample variance \( s^2 \). We can always compute it.
When we use a statistic to approximate a parameter, we call it an estimator, and the concrete value it takes in our sample is a point estimate. It is a "point" estimate because it is a single number: our best guess at the unknown parameter.
| Parameter (unknown) | Usual estimator | Formula |
|---|---|---|
| Population mean \( \mu \) | Sample mean | \( \bar{x} = \dfrac{1}{n}\sum_{i=1}^{n} x_i \) |
| Population proportion \( p \) | Sample proportion | \( \hat{p} = \dfrac{\text{number of successes}}{n} \) |
| Population variance \( \sigma^2 \) | Sample variance | \( s^2 = \dfrac{1}{n-1}\sum_{i=1}^{n} (x_i - \bar{x})^2 \) |
Note an important subtlety of notation: the estimator is a calculation rule (a random variable, because its value depends on which sample comes out), whereas the estimate is the number we obtain from one concrete sample. This distinction, which may sound pedantic, is the key to all of inference.
The Club Nova case: estimating average spend and the proportion of active members
Club Nova is NovaMarket's loyalty program, with about 380,000 members across Spain. Marta puts two questions to the team:
- What is a member's average monthly spend? (parameter: \( \mu \), in euros)
- What proportion of members is active? — defining "active" as having made at least one purchase in the last 90 days (parameter: \( p \))
Cross-referencing all 380,000 complete purchase histories against the billing systems is a months-long project; a well-designed sample gives the answer in days. The team draws two simple random samples from the member database:
Sample A (spend): \( n = 60 \) members, reviewing their receipts from the past month:
\[ \bar{x} = \text{€}126.40 \qquad s = \text{€}38.50 \]
Sample B (activity): \( n = 150 \) members, checking whether they bought in the last 90 days: 96 did.
\[ \hat{p} = \frac{96}{150} = 0.64 \]
The point estimates are therefore: "a Club Nova member's average monthly spend is about €126.40" and "64% of members are active".
Can we claim that \( \mu = 126.40 \) exactly? No. If we repeated the sampling tomorrow with 60 different members, a different mean would come out — perhaps €122.80, perhaps €131.10. The point estimate is the best guess, but it needs to travel with a measure of its precision. That is what the rest of this lesson is about (and, definitively, the next one).
Desirable properties of an estimator: unbiasedness, efficiency and consistency
Why do we use \( \bar{x} \) to estimate \( \mu \) and not, say, the sample median, or the mean of the first 10 data points? Because good estimators satisfy three properties. Let us understand them intuitively, without proofs.
Unbiasedness: right "on average"
An estimator is unbiased if, averaging over all possible samples, its expected value equals the parameter:
\[ E(\hat{\theta}) = \theta \]
Think of a darts player: they may miss on any given throw, but if their darts scatter centered on the bullseye, there is no bias. If they systematically drift to the left, there is.
- \( \bar{x} \) is unbiased for \( \mu \): some samples overshoot and others undershoot, but there is no systematic drift.
- \( \hat{p} \) is unbiased for \( p \), for the same reason.
- And here, at last, we solve the mystery of the n − 1 we have been dragging along since Module 2: if we divided by \( n \) when computing the sample variance, we would get an estimator of \( \sigma^2 \) that is biased downward. The intuitive reason: we measure deviations from \( \bar{x} \), which is the center of our sample, and data are always closer to their own center than to the true center \( \mu \). Dividing by \( n - 1 \) instead of \( n \) slightly "inflates" the result and corrects exactly that shortfall. That is why \( s^2 \) with \( n - 1 \) is the standard estimator of \( \sigma^2 \).
Efficiency: missing by little
Between two unbiased estimators, the more efficient one is the one with smaller variance — that is, the one that scatters less around the parameter. Back to the darts player: two players both centered on the bullseye, but one groups their darts in a tight circle and the other sprays them everywhere. We prefer the first.
Classic example: in a normal population, both the sample mean and the sample median are unbiased estimators of \( \mu \), but the median has roughly 25% more standard error. With the same 60 Club Nova members, estimating spend with the median would be like throwing part of the information in the bin. That is why the "official" estimator of \( \mu \) is \( \bar{x} \) (although, as we saw in Module 2, the median remains valuable as a descriptive measure when there is skewness or outliers).
Consistency: improving with more data
An estimator is consistent if, as \( n \) grows, it gets ever closer to the parameter (its error tends to zero). It is the bare-minimum property we would demand: "if you give me the whole population, get it exactly right".
\( \bar{x} \), \( \hat{p} \) and \( s^2 \) are consistent. You already know the practical consequence from the central limit theorem: the standard error carries \( \sqrt{n} \) in the denominator, so more sample = more precision (though with diminishing returns: to double the precision you must quadruple the sample).
| Property | Question it answers | Darts analogy |
|---|---|---|
| Unbiasedness | Right on average? | Darts centered on the bullseye |
| Efficiency | Missing by little? | Darts tightly grouped |
| Consistency | Improving with more data? | With more throws, the average nails the bullseye |
The sampling distribution of the estimator and the standard error
Here we pick up the big result from Module 4. The sample mean \( \bar{X} \) is a random variable and, by the central limit theorem, with \( n \) large enough:
\[ \bar{X} \sim N!\left(\mu, ; \frac{\sigma}{\sqrt{n}}\right) \]
That standard deviation of the sampling distribution, \( \sigma / \sqrt{n} \), is the standard error: it measures how much the estimator "swings" from sample to sample. It is the natural unit for judging the precision of a point estimate.
Let us apply it to Club Nova spend. We do not know \( \sigma \), so we use its estimate \( s = 38.50 \):
\[ \widehat{SE}(\bar{x}) = \frac{s}{\sqrt{n}} = \frac{38.50}{\sqrt{60}} = \frac{38.50}{7.746} \approx \text{€}4.97 \]
Interpretation: if we repeated the sampling many times, the sample means would typically deviate about €5 from the true average spend. Our estimate of €126.40 therefore carries an imprecision on the order of ±€5 (per "step" of standard error).
For the proportion, the standard error is \( \sqrt{p(1-p)/n} \), which we estimate with \( \hat{p} \):
\[ \widehat{SE}(\hat{p}) = \sqrt{\frac{0.64 \times 0.36}{150}} = \sqrt{0.001536} \approx 0.0392 \]
In other words, the estimated 64% swings about ±3.9 percentage points per standard error. Compare with the two standard errors we already computed in the previous module and will reuse constantly: the one for mean satisfaction (8.1 with SE ≈ 0.063) and the one for the card-payment proportion (0.61 with SE = 0.0345 for \( n = 200 \)).
When σ is unknown: Student's t distribution
In Module 4 we worked with a luxury we almost never have in practice: knowing \( \sigma \). When standardizing the sample mean we computed
\[ Z = \frac{\bar{X} - \mu}{\sigma / \sqrt{n}} \sim N(0,1) \]
But in the Club Nova case we have replaced \( \sigma \) with \( s \), which is itself an estimate with its own error. Once we make that substitution, the ratio
\[ T = \frac{\bar{X} - \mu}{s / \sqrt{n}} \]
no longer follows a normal distribution exactly: it carries somewhat more uncertainty, because the denominator also fluctuates from sample to sample. Its exact distribution (when the population is approximately normal) was derived by W. S. Gosset — who published under the pseudonym "Student" while working at the Guinness brewery — and it is called Student's t distribution.
Characteristics of the t:
- It is bell-shaped, symmetric and centered at 0, like the standard normal, but with heavier tails: values far from 0 are more likely. It is the mathematical way of saying "since I don't know σ, I'm hedging my bets".
- It depends on a parameter called degrees of freedom (df). For the mean of one sample, \( \text{df} = n - 1 \) (the same \( n-1 \) as in the denominator of \( s^2 \): of the \( n \) deviations from \( \bar{x} \), only \( n-1 \) are free, because they sum to zero).
- As the degrees of freedom grow, the t looks more and more like the N(0,1). With df ≥ 30 they are very similar; with df ≥ 100, nearly indistinguishable. That makes sense: with lots of sample, \( s \) estimates \( \sigma \) very well and the correction is barely needed.
| Degrees of freedom | Value leaving 2.5% in the right tail | Comparison |
|---|---|---|
| 5 | 2.571 | Well above 1.96 |
| 11 | 2.201 | |
| 24 | 2.064 | |
| 59 | 2.001 | Almost 1.96 already |
| ∞ (= normal) | 1.960 | The Module 4 reference |
Rule of thumb: for inferences about a mean with \( \sigma \) unknown (that is, almost always), use the t with \( n - 1 \) degrees of freedom. With large samples, t and z give practically identical results, but using the t is never wrong.
How to read the t table
The t table is organized differently from the Z table, and it is worth mastering because we will use it in the next three lessons:
- The Z table (Module 4) gives, for each value of z, the cumulative probability \( P(Z \le z) \): you enter with the value and come out with the probability.
- The t table works the other way around: you enter with the probability and come out with the value. Since there is a different t curve for each df, the table only records the critical values of the most commonly used tail probabilities.
Typical structure:
- Rows: degrees of freedom (1, 2, 3, …, 30, 40, 60, 120, ∞).
- Columns: area in the right tail (0.10; 0.05; 0.025; 0.01; 0.005).
- Cell: the value \( t \) that leaves exactly that area to its right, denoted \( t_{\alpha,,\text{df}} \).
Step-by-step reading example. We want the t value that leaves 2.5% in the right tail for the Club Nova spend sample (\( n = 60 \), so df = 59):
- Find the row df = 59 (if your table jumps from 40 to 60, use df = 60 or interpolate; to be safe, the convention is to take the nearest lower row — here it would make almost no difference).
- Find the column "right tail = 0.025".
- The cell gives \( t_{0.025,,59} \approx 2.001 \).
By symmetry, the value leaving 2.5% in the left tail is −2.001, and between −2.001 and +2.001 lies the central 95%. Compare: with the normal, 1.96 played that role. With df = 11, on the other hand, \( t_{0.025,,11} = 2.201 \) — the small-sample penalty is plain to see.
Two warnings when using the table:
- Some tables label the columns with the two-tailed area (0.05 instead of 0.025). Always check the diagram in the header.
- The last row (df = ∞) reproduces the familiar z values: 1.282; 1.645; 1.960; 2.326; 2.576. It is a good self-check.
Other estimation methods (a brief mention)
Where do estimators come from? Two classic factories we will only name: the method of moments builds estimators by equating the sample moments (mean, variance…) to the theoretical ones of the distribution, and the maximum likelihood method chooses as the estimate the parameter value that makes it most probable to have observed exactly the sample we have. For the cases in this course (\( \mu \), \( p \), \( \sigma^2 \)), both lead essentially to the estimators we already use, so we will not develop them further.
Common Mistakes and Tips
- Confusing the estimate with the parameter. Saying flatly "members' average spend is €126.40" is incorrect: €126.40 is the estimate; \( \mu \) remains unknown. In reports for management, always accompany the figure with its standard error or (better, after the next lesson) with an interval.
- Using z when t is called for. If \( \sigma \) is unknown and the sample is not large, standardizing with the normal understates the uncertainty. The safe rule: σ unknown → t with \( n-1 \) df.
- Getting the t table's tails wrong. Check whether your table gives one-tailed or two-tailed areas. A one-tailed 0.05 equals a two-tailed 0.10; mixing them up changes the critical value (1.671 versus 2.001 with df = 59, for example).
- Believing that unbiased = exact. An unbiased estimator can miss badly in any concrete sample; it only guarantees that it does not miss systematically in the same direction.
- Forgetting where the n − 1 comes from. If you are asked in a meeting, the short answer is: "dividing by n would underestimate the variance, because data are always closer to their own mean than to the true mean; n − 1 corrects for that".
Exercises
Exercise 1. From sample B (150 members, 96 active): (a) give the point estimate of the proportion of inactive members; (b) compute its estimated standard error; (c) reason whether it is the same as that of \( \hat{p} = 0.64 \), and why.
Exercise 2. Marta's team is considering expanding the spend sample from 60 to 240 members. Assuming \( s \) stayed at €38.50, what would the new standard error of the mean be? Which property of the estimator does this result illustrate?
Exercise 3. With a pilot sample of \( n = 12 \) members from the Cuenca store, the mean is to be standardized using \( s \). (a) Which distribution should be used, and with how many degrees of freedom? (b) Look up in the table the value that leaves 2.5% in the right tail. (c) Compare it with 1.96 and explain in one sentence why it is larger.
Solutions
Solution 1. (a) \( \hat{q} = 1 - 0.64 = 0.36 \) (or directly 54/150). (b) \( \widehat{SE} = \sqrt{0.36 \times 0.64 / 150} = \sqrt{0.001536} \approx 0.0392 \). (c) It is exactly the same, because \( p(1-p) = (1-p)p \): the uncertainty about "active" and about "inactive" is the same information seen from the two sides. Common mistake: recomputing with \( 0.36 \times 0.36 \) — the formula carries \( \hat{p}(1-\hat{p}) \), not \( \hat{p}^2 \).
Solution 2. \( \widehat{SE} = 38.50/\sqrt{240} = 38.50/15.49 \approx \text{€}2.49 \), half of the original (€4.97). Quadrupling the sample divides the standard error by \( \sqrt{4} = 2 \). It illustrates consistency: more data, a more precise estimate — but with diminishing returns (×4 the cost for ×2 the precision). Common mistake: expecting the error to be divided by 4; the square root prevents it.
Solution 3. (a) Student's t with df \( = 12 - 1 = 11 \), because \( \sigma \) is unknown and the sample is small. (b) \( t_{0.025,,11} = 2.201 \). (c) It is larger than 1.96 because the t has heavier tails: by estimating \( \sigma \) with only 12 data points we add uncertainty, and the critical value widens to compensate for it. Common mistake: using df = 12; the degrees of freedom are \( n - 1 \).
Conclusion
We have taken the first step of inference: using sample statistics as point estimators of the parameters (\( \bar{x} \to \mu \), \( \hat{p} \to p \), \( s^2 \to \sigma^2 \)), demanding good properties of them — unbiasedness (the reason behind the n − 1), efficiency and consistency — and measuring their precision with the standard error. And we have met Student's t, the wide-tailed bell that replaces the normal when \( \sigma \) is estimated with \( s \), along with its table and its degrees of freedom.
But a point estimate with its standard error is only half an answer: Marta does not want to hear "€126.40 with a swing of about €5" — she wants something like "between this and that, with 95% confidence". Turning the standard error into a range with guarantees — and finally answering the question the committee left hanging about the 8.1 satisfaction score — is the goal of the next lesson: Confidence Intervals.
Statistics Course
Module 1: Introduction to Statistics
Module 2: Describing Data
- Measures of Central Tendency
- Measures of Dispersion
- Measures of Position and Outliers
- Graphical Representation of Data
Module 3: Probability
Module 4: Probability Distributions
- The Binomial Distribution
- The Normal Distribution
- Other Important Distributions
- The Central Limit Theorem
Module 5: Statistical Inference
Module 6: Data Analysis
- Correlation Analysis
- Regression Analysis
- Analysis of Variance (ANOVA)
- Categorical Data Analysis: Chi-Square
