In the previous module we learned to draw inferences about one variable at a time: a mean, a proportion, a difference. But, as we hinted when closing it, NovaMarket's juiciest questions involve relationships between variables: do larger stores sell more? Does satisfaction drop when delivery takes longer? In this lesson we will learn to quantify the strength of a relationship between two quantitative variables: we will start from the scatter plot we introduced in Module 2, define covariance and the Pearson correlation coefficient (computed by hand, step by step), examine its conditions and pitfalls — with the celebrated Anscombe's quartet as a cautionary tale —, present the rank-based Spearman alternative, test whether a correlation is statistically significant and, above all, burn into memory the most important rule in data analysis: correlation does not imply causation.
Contents
- From the scatter plot to a single number
- Covariance: sign, interpretation, and the units problem
- The Pearson correlation coefficient: computing it by hand
- Interpreting the magnitude of r
- Conditions and pitfalls: linearity and outliers (Anscombe's quartet)
- Spearman correlation: ranks for ordinal data and non-linear relationships
- Is the correlation significant? The test on r
- Correlation does not imply causation
From the scatter plot to a single number
Let's bring back the scatter plot we drew in the Graphical Representation of Data lesson: the sales-floor area and average daily sales of eight NovaMarket stores. To keep the numbers manageable we will express sales in thousands of euros:
| Store | Floor area \( x \) (m²) | Daily sales \( y \) (€ thousands) |
|---|---|---|
| Cuenca | 850 | 19.2 |
| Bilbao-Casco | 1,100 | 44.0 |
| Zaragoza-Centro | 1,150 | 44.0 |
| Valencia-Ruzafa | 1,200 | 44.1 |
| Barcelona-Gràcia | 1,250 | 46.1 |
| Madrid-Chamberí | 1,300 | 47.5 |
| Sevilla-Nervión | 1,400 | 52.3 |
| Madrid-Centro | 2,000 | 68.4 |
Once you plot the points, the pattern jumps out: the bigger the floor area, the higher the sales. But Marta, the head of analytics, wants more than a visual impression: "How strong is the relationship? Give me a number I can compare against other variables." That number exists, and getting to it takes two steps: first the covariance, then the correlation.
Covariance: sign, interpretation, and the units problem
The idea is simple. We compute the means of both variables and ask, store by store, whether the two variables deviate from their means in the same direction or in opposite directions:
\[ s_{xy} = \frac{\sum_{i=1}^{n} (x_i - \bar{x})(y_i - \bar{y})}{n-1} \]
- If a store is above the mean in floor area and above it in sales (or below in both), the product \( (x_i-\bar{x})(y_i-\bar{y}) \) is positive.
- If it is above in one and below in the other, the product is negative.
- The covariance is (almost) the average of those products: its sign tells us the direction of the relationship.
Let's compute the means: \( \bar{x} = 10{,}250/8 = 1{,}281.25 \) m² and \( \bar{y} = 365.6/8 = 45.7 \) € thousands. We build the worksheet, which we will reuse throughout the lesson (and in the next one as well, so it is worth doing carefully):
| Store | \( x_i \) | \( y_i \) | \( x_i-\bar{x} \) | \( y_i-\bar{y} \) | \( (x_i-\bar{x})(y_i-\bar{y}) \) | \( (x_i-\bar{x})^2 \) | \( (y_i-\bar{y})^2 \) |
|---|---|---|---|---|---|---|---|
| Cuenca | 850 | 19.2 | −431.25 | −26.5 | 11,428.13 | 185,976.56 | 702.25 |
| Bilbao-Casco | 1,100 | 44.0 | −181.25 | −1.7 | 308.13 | 32,851.56 | 2.89 |
| Zaragoza-Centro | 1,150 | 44.0 | −131.25 | −1.7 | 223.13 | 17,226.56 | 2.89 |
| Valencia-Ruzafa | 1,200 | 44.1 | −81.25 | −1.6 | 130.00 | 6,601.56 | 2.56 |
| Barcelona-Gràcia | 1,250 | 46.1 | −31.25 | 0.4 | −12.50 | 976.56 | 0.16 |
| Madrid-Chamberí | 1,300 | 47.5 | 18.75 | 1.8 | 33.75 | 351.56 | 3.24 |
| Sevilla-Nervión | 1,400 | 52.3 | 118.75 | 6.6 | 783.75 | 14,101.56 | 43.56 |
| Madrid-Centro | 2,000 | 68.4 | 718.75 | 22.7 | 16,315.63 | 516,601.56 | 515.29 |
| Sum | 0 | 0 | 29,210.0 | 774,687.5 | 1,272.84 |
Notice two sanity checks: the deviation columns add up to zero (they always do; if not, there is an arithmetic error), and seven of the eight cross-products are positive — only Barcelona-Gràcia, slightly below the mean in floor area but a touch above it in sales, contributes a negative (and tiny) product.
\[ s_{xy} = \frac{29{,}210}{8-1} = 4{,}172.9 ; \text{m²·€ thousands} \]
Interpretation: the covariance is positive, so floor area and sales move in the same direction. But is 4,172.9 a strong relationship? Here comes the units problem: covariance is measured in "m² times thousands of euros", an uninterpretable unit, and its value changes if we change scale. Had we expressed sales in euros instead of thousands, the covariance would be 4,172,900 — a thousand times larger — without the relationship having changed one bit! Covariance is good for the sign, but not for measuring strength. We need to standardize it.
The Pearson correlation coefficient: computing it by hand
The solution is the same one we used with z-scores in Module 2: divide by the standard deviations to strip away the units. The result is the Pearson correlation coefficient:
\[ r = \frac{s_{xy}}{s_x , s_y} = \frac{\sum (x_i-\bar{x})(y_i-\bar{y})}{\sqrt{\sum (x_i-\bar{x})^2 ; \sum (y_i-\bar{y})^2}} \]
The second form is the handiest for manual calculation, because it uses the worksheet sums directly (the \( n-1 \) terms cancel between numerator and denominator). With our data:
\[ r = \frac{29{,}210}{\sqrt{774{,}687.5 \times 1{,}272.84}} = \frac{29{,}210}{\sqrt{986{,}053{,}238}} = \frac{29{,}210}{31{,}401.5} = 0.93 \]
Properties of \( r \) worth memorizing:
- It always lies between −1 and +1. It is dimensionless: it makes no difference whether you measure in m² or hectares, in euros or thousands of euros.
- \( r = +1 \): perfect positive linear relationship (all points on an upward-sloping line). \( r = -1 \): perfect negative linear. \( r = 0 \): no linear relationship.
- It is symmetric: the correlation of floor area with sales is the same as that of sales with floor area. It does not distinguish cause from effect (we will come back to this).
Interpreting the magnitude of r
Is 0.93 a lot or a little? There are no universal boundaries — they depend on the field —, but for business data this rough guide works well:
| \( \lvert r \rvert \) | Rough interpretation |
|---|---|
| 0.00 – 0.20 | Very weak or non-existent relationship |
| 0.20 – 0.40 | Weak |
| 0.40 – 0.60 | Moderate |
| 0.60 – 0.80 | Strong |
| 0.80 – 1.00 | Very strong |
Our \( r = 0.93 \) indicates a very strong linear relationship: bigger stores consistently sell more. Watch out for two caveats:
- The table is a guideline, not dogma. In behavioral surveys a 0.4 can be a remarkable finding; in a controlled physical process, a 0.9 can be disappointing.
- \( r \) measures the strength of the relationship, not its economic relevance or the size of the effect ("how many more euros per additional m²?"). That question will be answered by the regression slope in the next lesson.
Conditions and pitfalls: linearity and outliers (Anscombe's quartet)
Pearson measures linear relationships, and it is very sensitive to outliers. In 1973 the statistician Francis Anscombe constructed four sets of eleven points — the famous Anscombe's quartet — designed so that they all share virtually the same statistics: same mean of \( x \), same mean of \( y \), same correlation (\( r \approx 0.82 \) in all four)… and yet their plots could not look more different:
| Dataset | What you see when you plot it | Is \( r = 0.82 \) a good summary? |
|---|---|---|
| I | Classic linear cloud with moderate scatter | Yes: this is the "textbook" case |
| II | A perfect curve (parabola), with no noise | No: the relationship is extremely strong but not linear; \( r \) understates it and describes it poorly |
| III | Near-perfect line with one outlier off to the side | No: without that point, \( r \) would be ≈ 1; the outlier drags it down |
| IV | Every point at the same \( x \), except one far away | No: the "correlation" is manufactured by a single point; without it there is no relationship at all |
The moral applies daily at NovaMarket: before computing r, always draw the scatter plot. The same number can hide a clean relationship, a curve, or the effect of a single store. In our example it pays to look at Madrid-Centro (2,000 m², €68.4 thousand): it is the most extreme observation and contributes more than half of the total cross-product (16,315.6 out of 29,210). It is not an error — it is a real store, consistent with the pattern —, but if its figure had been misrecorded, it would badly distort the result. Practical rule: recompute \( r \) without the suspicious observation; if the value changes drastically, your conclusion hangs on a single point.
Spearman correlation: ranks for ordinal data and non-linear relationships
What do we do when the relationship is clearly monotonic but not linear (always rising or falling, but along a curve), or when one of the variables is ordinal (like a 0-to-10 satisfaction score, where the numbers order things but the gaps are not guaranteed to be equal), or when we fear outliers? Use the Spearman correlation \( r_s \): replace each value with its rank (ordered position: 1 for the smallest, 2 for the next…) and compute the correlation on the ranks. With no tied ranks there is a quick formula:
\[ r_s = 1 - \frac{6 \sum d_i^2}{n(n^2-1)} \]
where \( d_i \) is the difference between each observation's ranks on the two variables.
NovaMarket case: Marta suspects that online customers' satisfaction (recall: a mean of 7.6 versus 8.3 in-store) depends on delivery time. We take eight online orders with their delivery time and the customer's satisfaction score:
| Order | Delivery (h) | Satisfaction | Delivery rank | Satisfaction rank | \( d_i \) | \( d_i^2 \) |
|---|---|---|---|---|---|---|
| A | 14 | 9 | 1 | 7 | −6 | 36 |
| B | 17 | 10 | 2 | 8 | −6 | 36 |
| C | 20 | 8 | 3 | 6 | −3 | 9 |
| D | 23 | 6 | 4 | 4 | 0 | 0 |
| E | 26 | 7 | 5 | 5 | 0 | 0 |
| F | 32 | 5 | 6 | 3 | 3 | 9 |
| G | 44 | 3 | 7 | 2 | 5 | 25 |
| H | 55 | 2 | 8 | 1 | 7 | 49 |
| Sum | 0 | 164 |
Sanity check: the \( d_i \) column sums to zero. Applying the formula:
\[ r_s = 1 - \frac{6 \times 164}{8 \times (64-1)} = 1 - \frac{984}{504} = 1 - 1.952 = -0.95 \]
Interpretation: a very strong negative monotonic relationship — the longer the delivery takes, the lower the satisfaction, almost perfectly consistently. Notice that Spearman only asks "does the ordering of one variable replicate (or invert) the ordering of the other?": that is why it is immune to monotonic curvature and far more robust to outliers (an order that took 200 h would still simply be "rank 8"). If there are ties, the tied values are assigned the average of the ranks they would occupy (two orders tied for 3rd and 4th place each receive rank 3.5), and it is then preferable to compute \( r_s \) as the Pearson correlation of the ranks.
When should you use each one?
| Situation | Recommended coefficient |
|---|---|
| Two quantitative variables, linear relationship, no serious outliers | Pearson \( r \) |
| Ordinal variable(s) (grades, rankings, satisfaction scales) | Spearman \( r_s \) |
| Monotonic but curved relationship | Spearman \( r_s \) |
| Outliers that can be neither justified nor removed | Spearman \( r_s \) (as a robustness check) |
Is the correlation significant? The test on r
Our \( r = 0.93 \) was computed from only 8 stores. With small samples, couldn't an apparent correlation show up by sheer chance, even if no relationship exists in the population? It is the same question we settled in Module 5, and it is answered with a hypothesis test:
- \( H_0: \rho = 0 \) (no linear correlation in the population; \( \rho \), "rho", is the population correlation)
- \( H_1: \rho \neq 0 \)
Under \( H_0 \), the statistic
\[ t = \frac{r\sqrt{n-2}}{\sqrt{1-r^2}} \]
follows a Student's t distribution with \( n-2 \) degrees of freedom (we lose two: one for each estimated mean). With our data:
\[ t = \frac{0.93 \times \sqrt{6}}{\sqrt{1-0.865}} = \frac{0.93 \times 2.449}{\sqrt{0.135}} = \frac{2.278}{0.367} = 6.21 \]
In the t table (which we learned to read in Parameter Estimation), the two-tailed critical value at 5% with 6 df is \( t_{0.025;,6} = 2.447 \). Since \( 6.21 > 2.447 \), we reject \( H_0 \): the correlation is statistically significant (p < 0.001). Even with only 8 stores, such a strong relationship is very unlikely to arise by chance.
Two warnings carried over from Module 5: with large samples, minuscule correlations (r = 0.05) come out "significant" without being relevant — significance ≠ relevance —; and if you correlate many pairs of variables at once, remember the multiple comparisons problem: some pair will come out significant by chance.
Correlation does not imply causation
It is the most repeated sentence in statistics, and yet it is violated daily in business. That two variables move together does not prove that one causes the other. There are three possible explanations for an observed correlation:
- X causes Y (or Y causes X — remember that r is symmetric and does not distinguish direction).
- A third, lurking variable (a confounder) causes both.
- Coincidence (especially with small samples or after scanning many pairs of variables).
Two business examples:
- Ice cream and sunscreen. In NovaMarket's data, weekly sales of ice cream and sunscreen are extremely highly correlated. Does selling ice cream make people buy sunscreen? Obviously not: the lurking variable is summer heat, which drives both up. Promoting ice cream in January won't sell a single tube of sunscreen.
- Store size as a lurking variable. A junior analyst discovers that stores with more self-checkouts sell more, and proposes installing self-checkouts in Cuenca to lift its sales. The mistake: large stores have both more self-checkouts and higher sales — floor area (and the location that usually comes with it) causes both. Installing self-checkouts does not turn Cuenca into Madrid-Centro.
Even our star correlation (floor area–sales, r = 0.93) deserves causal caution: NovaMarket's large stores also sit in more central locations with heavier customer traffic. Part of the effect attributed to the square meters may be due to location. How do you actually establish causation? With controlled experiments (A/B tests, random assignment — we sketched this when discussing data collection) or with statistical designs that control for confounders, such as the multiple regression we will preview in Multivariate Analysis. Correlation, on its own, only suggests hypotheses; it never proves them.
Common Mistakes and Tips
- Computing r without drawing the scatter plot. This is mistake number one: Anscombe's quartet exists precisely to cure it. Plot first, number second.
- Reading r = 0 as "there is no relationship". It only means "there is no linear relationship". A perfect U shape (e.g., sales versus price: they fall when it is too expensive and also when it is so cheap it breeds distrust) can yield r ≈ 0.
- Confusing a large covariance with a strong relationship. Covariance depends on the units; to compare strengths, always use r.
- Leaping from correlation to causation. Before recommending an action ("let's enlarge the store to sell more"), ask yourself what lurking variable might explain the relationship.
- Ignoring outliers. A single point can manufacture or destroy a correlation. Recompute without it and compare; if the story changes, investigate that point before concluding anything.
- Applying Pearson to ordinal scales without a second thought. With rankings or opinion scales, Spearman is the defensible choice.
- Correlating everything with everything. With 20 variables there are 190 pairs: at the 5% significance level, expect ~9 or 10 false "significant" correlations. State your hypotheses before you look.
Exercises
Exercise 1. Using the floor area–sales table from the lesson, an intern proposes redoing the analysis with floor area measured in hectares (1 ha = 10,000 m²) "so the numbers are smaller". (a) Will the covariance change? And r? Reason it out without calculating. (b) Marta asks: "does that 0.93 mean floor area explains 93% of sales?" What would you answer?
Exercise 2. The Club Nova manager ranks 8 email campaigns by their spend and by the members they recruited:
| Campaign | Spend rank | Recruitment rank |
|---|---|---|
| C1 | 1 | 2 |
| C2 | 2 | 1 |
| C3 | 3 | 4 |
| C4 | 4 | 3 |
| C5 | 5 | 6 |
| C6 | 6 | 5 |
| C7 | 7 | 8 |
| C8 | 8 | 7 |
Compute the Spearman correlation and interpret it in one sentence for the manager.
Exercise 3. Analyzing 6 stores, another analyst obtains r = 0.75 between foot traffic and sales. (a) Test at the 5% level whether the correlation is significant (two-tailed critical value: \( t_{0.025;,4} = 2.776 \)). (b) The same r = 0.75 with n = 30 gives t = 6.00. What practical lesson do you draw from comparing the two cases?
Solutions
Exercise 1. (a) The covariance does change: it gets divided by 10,000, because it depends on the units of both variables. The correlation does not change at all: r is dimensionless and invariant to changes of scale; it will still be 0.93. This is exactly why we prefer r to the covariance. (b) No: r is not a percentage explained. The proportion of the variance in sales associated with floor area is \( r^2 = 0.93^2 \approx 0.865 \), that is, 86.5% — a concept (R²) that we will develop in the next lesson. Common mistake: reading r directly as a percentage; r = 0.5 is not "half explained" (that would be r² = 0.25, i.e. 25%).
Exercise 2. Rank differences: −1, 1, −1, 1, −1, 1, −1, 1, so \( \sum d_i^2 = 8 \). Then \( r_s = 1 - \frac{6 \times 8}{8 \times 63} = 1 - \frac{48}{504} = 1 - 0.095 = 0.90 \). Interpretation: "the recruitment order replicates the spend order almost perfectly: the campaigns with the biggest budgets consistently recruit the most members". Common mistake: forgetting to square the differences (with these d values, \( \sum d_i \) equals 0 and would look like a perfect correlation) or misusing n² − 1 = 63 (remember: it is n·(n²−1) = 8 × 63 = 504).
Exercise 3. (a) \( t = \frac{0.75\sqrt{4}}{\sqrt{1-0.5625}} = \frac{1.5}{0.661} = 2.27 \). Since 2.27 < 2.776, we cannot reject \( H_0 \): with only 6 stores, a correlation of 0.75 is compatible with chance. (b) The same correlation with n = 30 is indeed significant. Lesson: significance depends jointly on the strength (r) and the sample size (n). A high r with a tiny n proves nothing, and — conversely, as we saw in Module 5 — an irrelevant r with an enormous n can come out "significant". Common mistake: using n instead of n − 2 in the formula, or concluding from part (a) that "there is no relationship": there is only insufficient evidence.
Conclusion
In this lesson we made the leap from describing variables separately to measuring relationships: covariance gives us the sign (but is a slave to the units), the Pearson coefficient r turns it into a universal measure between −1 and +1, Spearman covers us when there are ordinal scales, monotonic curves or outliers, and the t test on r protects us from correlations manufactured by chance in small samples. We take away two disciplines: always plot before you compute (Anscombe) and never confuse correlation with causation (ice cream and sunscreen). But Marta is already asking for the next step: "Fine, floor area and sales go hand in hand with r = 0.93. If we open a 1,500 m² store, how much will it sell?" Answering "how much" requires turning the relationship into a prediction equation — a line with a slope and an intercept —, and that is exactly the Regression Analysis we tackle next, reusing the worksheet we have built here.
Statistics Course
Module 1: Introduction to Statistics
Module 2: Describing Data
- Measures of Central Tendency
- Measures of Dispersion
- Measures of Position and Outliers
- Graphical Representation of Data
Module 3: Probability
Module 4: Probability Distributions
- The Binomial Distribution
- The Normal Distribution
- Other Important Distributions
- The Central Limit Theorem
Module 5: Statistical Inference
Module 6: Data Analysis
- Correlation Analysis
- Regression Analysis
- Analysis of Variance (ANOVA)
- Categorical Data Analysis: Chi-Square
