In the regression of Module 6 we predicted a store's sales from a single variable: its floor area. It worked well (R² = 0.865), but it left questions open — Cuenca sells about €10,200/day less than its floor area predicts, and we suspect the culprit is not the square meters but the street it sits on. In the real world almost no business phenomenon depends on a single variable: sales depend on floor area and foot traffic and the area's income; a member's churn depends on their age and their spend and their activity. Multivariate analysis is the toolbox for working with several variables at once. In this lesson you will learn multiple regression (the natural extension of what you already know), logistic regression (when what you predict is a yes/no), and, at a conceptual level, two exploratory techniques: principal components and clustering. The emphasis is on interpretation: software will always fit these models; your job as an analyst is to know how to read them and explain them to Marta.

Contents

  1. From one variable to several: the multiple regression model
  2. Reading the coefficients: the "all else being equal" clause
  3. Interpreting the software output column by column
  4. Adjusted R²: why R² always goes up
  5. Multicollinearity: when predictors step on each other
  6. Dummy variables and interactions
  7. Logistic regression: predicting a yes/no (Club Nova churn)
  8. Principal component analysis: condensing many variables into a few
  9. Clustering (k-means): segmenting the members

From one variable to several: the multiple regression model

Multiple linear regression generalizes the least-squares line to \(k\) predictors:

\[ \hat{y} = b_0 + b_1 x_1 + b_2 x_2 + \dots + b_k x_k \]

It is no longer a line but a "plane" (or hyperplane) fitted to the data, yet the fitting criterion is the same — minimize the sum of squared residuals — and all the inference machinery (standard errors, t tests, p-values, R²) works exactly as in 06-02, only now there is one coefficient per variable.

The NovaMarket case. The expansion team extends the sales model with two new variables measured across the 42 stores:

  • \(x_1\): sales-floor area (m²)
  • \(x_2\): foot traffic on the street (thousands of pedestrians/day)
  • \(x_3\): average income of the area (€ thousands/year per capita)

The response variable is still daily sales in € thousands (recall: mean €43,983 ≈ 44.0 thousand).

Reading the coefficients: the "all else being equal" clause

The fitted model comes out as:

\[ \hat{y} = -4.10 + 0.0312,x_1 + 0.85,x_2 + 0.42,x_3 \]

Each coefficient is read with a mandatory clause: "all else being equal" (economists say ceteris paribus):

  • \(b_1 = 0.0312\): between two stores with the same foot traffic and the same income, each additional m² is associated with €31.20 more in daily sales.
  • \(b_2 = 0.85\): at equal floor area and income, each additional thousand pedestrians per day is associated with €850 more in daily sales.
  • \(b_3 = 0.42\): at equal floor area and traffic, each additional thousand euros of per-capita income in the area is associated with €420 more in daily sales.

Notice an important detail: in the simple regression the floor-area slope was 0.0377 (€37.70/m²); now it is 0.0312. Has the effect "shrunk"? Not exactly: in the simple model, floor area took the credit for everything that travels with it — NovaMarket's large stores also tend to sit on busy streets. Once traffic enters as its own variable, each coefficient reflects only its net contribution. That is the great virtue of multiple regression: it separates effects that arrived tangled together in two-variable analysis (the confounding phenomenon that already surfaced when we warned that correlation is not causation).

And it solves the pending mystery: plugging in Cuenca's data (very low foot traffic for its floor area), its residual goes from −10.2 thousand € in the simple model to −1.8 in the multiple one. Cuenca was not a badly run store: it was a big store on an empty street.

Interpreting the software output column by column

This is how a typical statistical package presents the fit (n = 42 stores):

Variable Coefficient Standard error t p-value
Constant −4.10 2.05 −2.00 0.053
Floor area (m²) 0.0312 0.0031 10.06 < 0.001
Traffic (thousands of pedestrians/day) 0.85 0.24 3.54 0.001
Area income (€ thousands/year) 0.42 0.35 1.20 0.238

R² = 0.912 · Adjusted R² = 0.905 · n = 42

Column by column, with what you already know from Modules 5 and 6:

  • Coefficient: the point estimate of the net effect, in units of the response per unit of the predictor.
  • Standard error (SE): the uncertainty of that estimate; with it you build the 95% CI (≈ coefficient ± 2·SE). For floor area: \(0.0312 \pm 2 \times 0.0031 \Rightarrow (0.025;, 0.037)\).
  • t: the test statistic \(t = \text{coefficient} / SE\) for the null hypothesis "this coefficient is 0 (the variable adds nothing, given that the others are in the model)". For traffic: \(0.85/0.24 = 3.54\).
  • p-value: the probability of a t that extreme if the variable added nothing. Floor area and traffic are clearly significant; income is not (p = 0.238): once floor area and traffic are known, knowing the neighborhood's income barely improves the prediction. A reasonable final model would drop it.

The italicized tail of the t test matters: each variable's p-value is conditional on the rest of the model. Removing or adding a variable can change the p-values of all the others.

Adjusted R²: why R² always goes up

The multiple model's R² is 0.912, against 0.865 for the floor-area-only model. Does that prove the model is better? Careful: R² never falls when you add variables, even if they are pure noise. If you added "number of letters in the store's name", R² would tick up a fraction of a percentage point, because least squares always finds some chance covariation to squeeze.

The adjusted R² corrects for this by penalizing each additional variable:

\[ R^2_{adj} = 1 - (1 - R^2),\frac{n-1}{n-k-1} \]

where \(n\) is the number of observations and \(k\) the number of predictors. It only rises if the new variable contributes more than it "costs". Practical rules:

  • To compare models with different numbers of variables, always use the adjusted R² (here: 0.905 with three variables against 0.862 for the simple model — the improvement is real).
  • If adding a variable raises R² but lowers the adjusted R², the variable is probably dispensable. That is what happens with income: without it, the adjusted R² barely changes.
  • A model with many variables and few data points can "memorize" the sample (overfitting) and predict new stores badly — the multivariate version of a danger you already know from extrapolation.

Multicollinearity: when predictors step on each other

Multiple regression divides the credit among predictors. What if two predictors are nearly the same information? For example, if we added "total premises area" to the model on top of "sales-floor area" (correlation between the two: 0.98). That is multicollinearity, and its symptoms are recognizable:

  • The model's R² is high, but no individual coefficient comes out significant: the joint contribution is clear, but splitting it between twin variables is impossible to pin down.
  • The standard errors inflate and the coefficients become unstable: they change a lot (even in sign) when a few observations are added or removed.
  • Coefficients with absurd signs ("more floor area, fewer sales") offset by the twin variable.

The intuition: asking the data "how much more does a store sell with more sales floor at equal total premises area?" is a question the sample can barely answer, because there are almost no stores where the two differ. The standard diagnostic is the VIF (variance inflation factor): it measures how many times the variance of each coefficient is inflated by its correlation with the other predictors; as a guideline, VIF > 5 calls for review and VIF > 10 signals a serious problem. In our model: floor area 1.9, traffic 2.1, income 1.3 — moderate correlations, no problem at all. The usual fix when there is one: keep one of the twin variables (the more interpretable, or the cheaper to measure), or combine them — something the principal component analysis of section 8 does systematically.

Dummy variables and interactions

Including a categorical variable: the store format

How do you put a variable like format (Express / Standard / Superstore), which is not a number, into the equation? With dummy variables (indicators): 0/1 variables that flag membership in each category, leaving one out as the reference category. With Express as the reference:

Format \(D_{Std}\) \(D_{Sup}\)
Express 0 0
Standard 1 0
Superstore 0 1

A store-satisfaction model with these dummies:

\[ \widehat{\text{satisfaction}} = 7.0 + 1.0,D_{Std} + 2.0,D_{Sup} \]

reads as follows: the constant (7.0) is the mean of the reference category (Express); each dummy coefficient is the difference relative to the reference (Standard: 7.0 + 1.0 = 8.0; Superstore: 7.0 + 2.0 = 9.0 — exactly the ANOVA means from 06-03, which reappear here as a special case of regression). Mechanical rule: a categorical variable with \(c\) levels enters with \(c-1\) dummies; including all \(c\) would be redundant (perfect multicollinearity).

Interactions (a conceptual mention)

Sometimes the effect of one variable depends on the value of another: perhaps each extra thousand pedestrians is worth more at an Express store (impulse, walk-by purchases) than at a Superstore (a planned destination). That is modeled by adding the product of the two variables (\(x_2 \times D_{Sup}\)) as an extra predictor; if its coefficient is significant, the traffic slope differs by format. Take away the idea and the warning: with interactions, individual coefficients can no longer be read in isolation, and it pays to lean on plots of predicted values.

Logistic regression: predicting a yes/no (Club Nova churn)

Marta raises the star problem of the year: predicting which Club Nova members will cancel their membership (the annual sample churn rate is 14%, concentrated among members under 35). The response is binary (churn: yes/no), and linear regression will not do: it would predict "probabilities" below 0 or above 1.

Logistic regression models the probability through the odds (the for/against ratio \(p/(1-p)\), which we already used in probability: p = 0.14 means odds of 0.163, "1 cancellation for every 6.1 who stay"). The model is linear in the logarithm of the odds:

\[ \ln!\left(\frac{p}{1-p}\right) = b_0 + b_1 x_1 + \dots + b_k x_k \]

Fitted on a sample of members (variables: age in years, monthly spend in €, activity — the "active member" dummy; remember they are 64%):

Variable Coefficient p-value \(e^{coef}\) (odds ratio)
Constant 0.90 0.004 —
Age (years) −0.045 < 0.001 0.956
Monthly spend (€) −0.006 0.010 0.994
Active member (1 = yes) −1.20 < 0.001 0.301

How to read it, without needing to know how it is estimated (the software does it via a method called maximum likelihood, which you do not need to master):

  • The sign comes first: a negative coefficient = that variable reduces the probability of churn. Being older, spending more and being active all protect — consistent with churn concentrating among the young.
  • The magnitude is interpreted via \(e^{coef}\), the odds ratio: for each additional year of age, the odds of churn are multiplied by 0.956 (down 4.4%); over 10 years, \(e^{-0.45} \approx 0.64\), 36% lower odds. Being an active member multiplies the odds of churn by 0.30: it divides them by more than 3.
  • The p-values work as always: here all three variables are significant.

The model also yields individual probabilities. Two members with opposite profiles:

  • A 28-year-old, inactive, spending €60/month: \(\ln(odds) = 0.90 - 0.045(28) - 0.006(60) = -0.72\), so \(p = \frac{1}{1+e^{0.72}} \approx 0.33\). A 33% probability of churning: a clear candidate for a retention campaign.
  • A 52-year-old, active, spending €150/month: \(\ln(odds) = 0.90 - 2.34 - 0.90 - 1.20 = -3.54\), so \(p \approx 0.03\). A 3% probability: do not spend retention budget on her.

That is exactly the business application: rank the 380,000 members by churn probability and concentrate the campaign on the top of the list.

Principal component analysis: condensing many variables into a few

The two remaining techniques predict nothing: they explore. Principal component analysis (PCA) answers this question: I have many variables correlated with each other; can I condense them into a few new variables while losing as little information as possible?

Conceptually: PCA builds components — combinations of the original variables — such that the first captures as much of the data's variance as possible, the second captures as much as possible of what remains (while being independent of the first), and so on. If the original variables are highly correlated, two or three components are enough to retain 70-90% of the information in ten or fifteen variables.

Mini-case: on 12 behavioral variables of Club Nova members (spend by category, number of visits, basket size, online channel use, coupon use…), PCA reports that two components retain 78% of the variance, and by looking at which variables load on each one you can give them names:

  • Component 1, "purchase volume" (52%): total spend, basket size and spend on Fresh food and Pantry load heavily.
  • Component 2, "digital-ness" (26%): online channel use, in-app coupons and the frequency of short visits load heavily.

Typical uses: plotting the members on a plane (component 1 vs component 2) instead of in 12 dimensions impossible to draw, and replacing groups of twin variables with a single component — the systematic remedy for multicollinearity. The price: components are constructs, and their interpretation ("digital-ness") is supplied by the analyst, not by the mathematics.

Clustering (k-means): segmenting the members

The final question is different: do natural groups of members exist that resemble each other and differ from the rest? Nobody has labeled anything in advance — this is not predicting a known response, it is discovering structure. That is clustering (cluster analysis), and its most popular algorithm is k-means: the analyst fixes the number of groups \(k\), and the algorithm assigns each member to the group with the nearest center and recomputes the centers, iterating until things settle. Each group ends up described by its centroid (the group's "average" member).

The NovaMarket team segmented the 380,000 Club Nova members with k-means and \(k = 3\). Your job is to interpret the result:

Cluster 1: "Full-basket families" Cluster 2: "Convenience shoppers" Cluster 3: "At-risk occasionals"
% of members 38% 41% 21%
Average monthly spend €182 €96 €58
Visits/month 6.1 8.3 (short) 1.8
Dominant profile Big weekly shop, strong on Fresh food Express stores and online, small baskets Under 35, mostly inactive
Annual churn rate 6% 12% 27%

The business reading: the overall average spend (€126.40) hides three very different realities; the overall 14% churn is driven almost entirely by cluster 3 — which ties in with what the logistic regression was already flagging (young, inactive, low spend). Each cluster calls for a different action: build Fresh food loyalty in cluster 1, push the app in cluster 2, an urgent retention campaign in cluster 3.

Two conceptual warnings: the choice of \(k\) is not given by the data alone (there are guiding criteria, but business judgment weighs in: three actionable segments, or seven unmanageable ones?); and k-means works with distances, so the variables must be standardized first (remember the z-scores of Module 2) — otherwise, the variable on the largest scale (spend in euros) would dominate the segmentation on its own.

Common Mistakes and Tips

  • Reading multiple regression coefficients without the "all else being equal" clause. Saying "each m² contributes €31.20" flat out invites a bad comparison with the simple model's 37.7; they are answers to different questions.
  • Boasting about R² by piling up variables. Compare models with the adjusted R², and distrust models with many predictors and few data points.
  • Interpreting individual coefficients under severe multicollinearity. If two predictors are near-twins, their separate coefficients are not reliable even if the model predicts well.
  • Forgetting the reference category with dummies. Each dummy coefficient is a difference relative to the reference, not a group mean.
  • Reading logistic coefficients as probabilities. A coefficient of −1.20 does not mean "−120% probability": it acts on the logarithm of the odds; convert it to an odds ratio (\(e^{-1.20} = 0.30\)) or compute probabilities for specific profiles.
  • Believing the clusters "exist". k-means always returns \(k\) groups, whether or not there is real structure; check that the segments are stable and actionable before building campaigns on them.
  • Confusing prediction with causation. Everything above is observational association; to claim that changing foot traffic (or activating a member) causes higher sales or lower churn, you would need an experiment — we will talk about that in the applications of Module 8.

Exercises

Exercise 1

The expansion team is evaluating premises in Sevilla with 1,100 m², on a street with 14,000 pedestrians/day and an area income of €24,000/year. Using the model \(\hat{y} = -4.10 + 0.0312,x_1 + 0.85,x_2 + 0.42,x_3\) (x₂ and x₃ in thousands), compute the predicted daily sales. Is this prediction more or less reliable than the one we made for Alicante with the simple model? Why?

Exercise 2

In the Club Nova churn logistic model, a colleague claims: "monthly spend hardly matters: its coefficient is −0.006, laughable next to activity's −1.20". Assess the claim by computing the odds ratio associated with a €100 difference in monthly spend, and explain why comparing raw coefficients is misleading.

Exercise 3

Marta forwards you the 3-cluster table and asks: "which segment do I invest the retention budget in first, and what do I offer them?" Answer using the table, and add a statistical check from earlier modules with which you would verify that the difference in churn rates across clusters is not down to chance.

Solutions

Exercise 1.

\(\hat{y} = -4.10 + 0.0312 \times 1100 + 0.85 \times 14 + 0.42 \times 24 = -4.10 + 34.32 + 11.90 + 10.08 = 52.2\) thousand € ≈ €52,200/day. It is in principle more reliable than the Alicante prediction from the simple model: the adjusted R² is higher (0.905 vs 0.862) and the model separates out the traffic effect, which in Alicante was baked into floor area. With two caveats, though: Sevilla's values must lie within the range observed across the 42 stores (no extrapolating), and the income variable adds little (p = 0.238) — with the model without income, the prediction would barely change. Common mistake: forgetting that x₂ and x₃ enter in thousands and multiplying by 14,000 and 24,000, which produces absurd figures; always check each variable's units before substituting.

Exercise 2.

The claim confuses coefficient magnitude with importance: coefficients depend on the units. The −0.006 is per euro; a realistic difference between members runs to tens or hundreds of euros. For €100: \(e^{-0.006 \times 100} = e^{-0.6} \approx 0.55\) — a member spending €100 more per month has nearly half the odds of churning, an effect comparable to activity's (\(0.30\)). Common mistake: comparing coefficients of variables measured on different scales (years, euros, a 0/1 dummy) as if they were commensurable; always compare odds ratios associated with realistic differences in each variable.

Exercise 3.

Cluster 3 ("At-risk occasionals") first: it concentrates the highest churn rate (27%, nearly double the 14% average) and is where an improvement has the most headroom; its members are young, inactive and low-spending, so the natural offer targets activation (app perks, repeat-purchase coupons, subsidized online delivery) rather than indiscriminate discounting. Statistical check: a chi-square test of independence (06-04) on the cluster × churn/no-churn table — with these sample sizes, significance is guaranteed, so accompany it with an effect-size measure (Cramér's V) and with the standardized residuals to locate where the excess churn sits. Common mistake: picking cluster 2 "because it has more members"; volume matters, but the retention decision is driven by risk and by room for improvement, not by segment size.

Conclusion

You have made the leap from one variable to several: multiple regression separates effects that used to travel tangled together (and it rehabilitated Cuenca: its shortfall was about foot traffic, not management), the software output is read column by column with the inference tools you already commanded, the adjusted R² guards against the temptation to pile up variables, dummies bring in categorical variables (and revealed ANOVA as a special case of regression), logistic regression carries the whole scheme over to yes/no responses via odds, and PCA and k-means explore structure with no response variable — with the three Club Nova segments as an actionable result.

But this whole edifice — like the t test, ANOVA or the regression intervals — still rests on distributional assumptions: normal residuals, reasonably large samples, no wild outliers. And when the data refuse to cooperate — eight observations, ordinal satisfaction scales, one outlier that distorts everything? For that there is a family of tests that give up the assumptions in exchange for working with ranks: the Non-Parametric Methods, with which we will close the module.

© Copyright 2026. All rights reserved