In 04-05, while studying K-NN, a threat was noted by name: the curse of dimensionality. In this module the threat is direct, because clustering means distances, and distances degrade as dimensions multiply. Principal Component Analysis (PCA) is the classic answer: an unsupervised technique that compresses many features into a few new components, preserving as much as possible of the original variance — that is, of the information. In this lesson you will understand the problem it solves, the geometric intuition of "rotating the axes toward where the data lies", the concepts of explained and cumulative variance, its connection to the covariance matrix from 02-03, and its practical use at MercaFresh: visualizing the segments from 05-01 in 2D and interpreting what each component means through its loadings.
Contents
- The curse of dimensionality
- Geometric intuition: rotating the axes toward maximum variance
- Principal components, explained variance and the scree plot
- The connection to covariance (02-03) and why to standardize first (03-05)
- PCA in scikit-learn on the MercaFresh customers
- 2D visualization of the K-means segments
- Interpreting components: the loadings
- PCA as a preprocessing step for models
- Limitations, and PCA vs. feature selection (03-06)
The curse of dimensionality
What happens when the features go from 4 to 40, or to 400? Three ills combine:
- The space empties out. To cover a 1D interval with 10 evenly spread points, 10 are enough; to cover a 10-dimensional cube at the same density you would need $10^{10}$. With dimensions to spare, every dataset is a speck of dust in a deserted space: every point is far from everything.
- Distances lose contrast. In high dimensions, the distance to the nearest neighbor and to the farthest one tend to look alike. And if "near" and "far" are barely distinguishable, K-NN (04-05), K-means (05-01) and hierarchical clustering (05-02) — all built on distances — lose their raw material.
- Accumulated noise. Every irrelevant feature adds its noise to the Euclidean distance; with many of them, the noise drowns out the signal from the few that matter.
The obvious way out would be to drop irrelevant features (the selection filters from 03-06). But what if many features are relevant yet redundant with each other — like total_spend and num_orders, correlated back in 02-03? There you don't want to delete columns: you want to merge the information into fewer dimensions. That is dimensionality reduction, and PCA is its flagship technique.
Geometric intuition: rotating the axes toward maximum variance
Picture the scatter of two correlated MercaFresh features: total_spend against num_orders (standardized). The point cloud is a tilted ellipse: whoever places more orders spends more in total. Notice two things:
- The diagonal direction of the ellipse concentrates almost all the variation between customers: it is the "customer size" axis (a lot/a little of both things at once).
- The perpendicular direction barely varies: it only captures the nuance of "spends more/less than their number of orders would suggest".
PCA formalizes exactly this observation: it finds new axes, by rotating the original ones, so that the first axis points where the data varies most; the second, perpendicular to the first, toward the maximum remaining variance; and so on.
flowchart LR
A["Original axes:<br/>total_spend, num_orders<br/>(correlated, redundant)"] -->|"PCA rotation"| B["New axes:<br/>PC1 = customer size (95% var.)<br/>PC2 = residual nuance (5% var.)"]
B --> C["Compression: keeping<br/>only PC1 loses<br/>barely 5% of the information"]
The new axes are the principal components (PC1, PC2, ...). Three properties worth committing to memory:
- Each component is a linear combination of the original features (for example, PC1 ≈ 0.71·spend + 0.71·orders): it does not pick columns, it blends them.
- The components are mutually perpendicular and, by construction, uncorrelated: PCA turns redundant features into independent directions.
- They are ordered by variance: PC1 captures more than PC2, which captures more than PC3... Compressing is simply keeping the first ones and discarding the rest.
In this sense, "information" for PCA means variance: the directions in which customers differ a lot from each other are the ones that allow you to tell them apart; a direction with near-zero variance says almost nothing about anybody (the same argument as the variance filter in 03-06).
Principal components, explained variance and the scree plot
How many components to keep? The criterion is explained variance: what fraction of the total variance each component captures.
explained_variance_ratio_in scikit-learn: for example[0.55, 0.25, 0.12, 0.05, 0.03]— PC1 explains 55%, PC2 25%...- Cumulative variance adds them up in order: with 2 components, 80%; with 3, 92%. The most common practical rule: keep as many components as needed to accumulate 90-95%.
- The scree plot graphs the explained variance per component. Just like the elbow in 05-01, you look for the point where the curve flattens: the components past the elbow contribute crumbs (often, noise).
| Tool | What it shows | Decision it supports |
|---|---|---|
explained_variance_ratio_ |
% of variance of each PC | Is PC3 worth anything? |
| Cumulative variance | Total % with the first k PCs | How many do I keep for 90%? |
| Scree plot | The curve of the previous two | Locate the elbow visually |
The connection to covariance (02-03) and why to standardize first (03-05)
Where do those magical directions come from? From the covariance matrix you studied in 02-03: the table recording how each pair of features co-varies. PCA analyzes that matrix and extracts its characteristic directions from it — each principal component is one of those directions, and its explained variance is the associated magnitude (in linear algebra they are called eigenvectors and eigenvalues; we don't need the full derivation, only the consequence):
PCA is the covariance matrix turned into axes. Where the correlation heatmap from 02-03 told you "spend and orders go together", PCA gives you the concrete axis that summarizes that going-together.
From here also comes the practical golden rule: standardize before PCA (03-05, StandardScaler). Covariance depends on units: if total_spend is measured in euros (variance in the thousands) and orders_per_month in single units (variance in the tens), the direction of maximum variance will be "the spend axis" by sheer accident of units, and PC1 will be an echo of the largest column. Standardizing equalizes all variances to 1, so PCA effectively works on the correlation matrix and the directions reflect structure, not units. It is the same argument as with K-means; in PCA it is even more critical because variance is not just the internal metric but the very construction criterion.
PCA in scikit-learn on the MercaFresh customers
Let's widen the customer matrix with more features from module 3 — RFM, ratios and trend — so the reduction has a point:
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt
import numpy as np
features = ["recency_days", "orders_per_month", "avg_order_spend",
"total_spend", "num_orders", "inactivity_ratio", "trend"]
X = rfm[features]
X_esc = StandardScaler().fit_transform(X) # ALWAYS before PCA
pca = PCA() # no limit: all the components
X_pca = pca.fit_transform(X_esc)
print(pca.explained_variance_ratio_.round(3))
# e.g.: [0.46 0.27 0.12 0.07 0.05 0.02 0.01]
print(np.cumsum(pca.explained_variance_ratio_).round(3))
# e.g.: [0.46 0.73 0.85 0.92 0.97 0.99 1.00]
# Scree plot
plt.plot(range(1, 8), pca.explained_variance_ratio_, "o-")
plt.xlabel("Principal component")
plt.ylabel("Explained variance")
plt.title("Scree plot — MercaFresh customers")
plt.show()Reading the result:
fit_transformprojects each customer onto the new axes:X_pcahas one column per component, ordered from most to least variance.- Two components accumulate 73% and four accumulate 92%: the 7 original features were largely redundant (total spend, number of orders and frequency tell overlapping stories — we knew that since the heatmap of 02-03).
- Convenient constructor alternatives:
PCA(n_components=2)(a fixed number) orPCA(n_components=0.90)(scikit-learn picks however many are needed to accumulate 90%).
2D visualization of the K-means segments
PCA's most immediate use: our segments from 05-01 live in a space we cannot draw; projected onto PC1-PC2 we can.
pca2 = PCA(n_components=2)
X_2d = pca2.fit_transform(X_esc)
plt.scatter(X_2d[:, 0], X_2d[:, 1], c=rfm["segment"], cmap="tab10", s=15)
plt.xlabel(f"PC1 ({pca2.explained_variance_ratio_[0]:.0%} var.)")
plt.ylabel(f"PC2 ({pca2.explained_variance_ratio_[1]:.0%} var.)")
plt.title("MercaFresh K-means segments on the PCA plane")
plt.colorbar(label="segment")
plt.show()We color each point by the segment K-means assigned it (c=rfm["segment"]). If the four colors show up in reasonably distinct regions of the plane, we have visual confirmation that the segments are real regions of the customer space. With one honest caveat: the plane only shows 73% of the variance — two clusters that overlap in the drawing may be separated along the dimension the plane doesn't show. The projection is a map, not the territory.
Interpreting components: the loadings
PC1 and PC2 are blends of features. What do they mean? The answer lies in the loadings: the weights with which each original feature enters each component, available in pca.components_.
import pandas as pd
loadings = pd.DataFrame(pca2.components_.T,
columns=["PC1", "PC2"], index=features).round(2)
print(loadings)| Feature | PC1 | PC2 |
|---|---|---|
| recency_days | 0.45 | 0.21 |
| orders_per_month | −0.44 | 0.25 |
| avg_order_spend | −0.12 | 0.62 |
| total_spend | −0.41 | 0.48 |
| num_orders | −0.43 | 0.18 |
| inactivity_ratio | 0.47 | 0.15 |
| trend | −0.11 | −0.50 |
How to read them (signs and values are illustrative):
- PC1 pits recency and inactivity ratio (positive weights) against frequency and volume (negative): it is an activity axis, from "live customer" (very negative PC1) to "faded customer" (very positive PC1). No surprise that the "VIP" and "dormant" segments from 05-01 occupy opposite ends of this axis.
- PC2 loads on average and total spend with a positive sign and on trend with a negative one: a ticket value axis, separating high-ticket regulars from cheap occasionals.
- The overall sign of a component is arbitrary (the axis can point either way); what is interpretable are the relative signs between features and the magnitude of the weights.
With the loadings, the scatter's axes stop being abstract: the customer map has a horizontal "activity" axis and a vertical "value" axis — vocabulary the business understands.
PCA as a preprocessing step for models
Besides visualizing, PCA is used as preprocessing: training the model on the first k components instead of the original features. Benefits: fewer dimensions (relieves the curse for K-NN or K-means), no collinearity (the PCs are uncorrelated) and less noise (the last components are usually noise and get discarded). It slots into the Pipeline you already know from module 4:
from sklearn.pipeline import Pipeline
from sklearn.cluster import KMeans
pipe = Pipeline([
("scale", StandardScaler()),
("pca", PCA(n_components=0.90)), # PCs up to 90% of the variance
("kmeans", KMeans(n_clusters=4, random_state=42)),
])
labels = pipe.fit_predict(X)The order matters, and you can now justify all of it: scale (so PCA does not inherit the units) → PCA (so K-means works with a few informative dimensions) → clustering. The trade-off: the centroids come out expressed in components, and to profile them in business units you have to undo both transformations (inverse_transform of the PCA and of the scaler, chained).
Limitations, and PCA vs. feature selection (03-06)
Limitations of PCA:
- It is linear. It only finds rotations: if the data's structure is curved (a spiral, a twisted manifold), no rotation unfolds it. For visualizing nonlinear structure there are t-SNE and UMAP (05-05).
- Interpretability. "PC1 = 0.45·recency − 0.44·frequency + ..." will never be as transparent as an original feature. The loadings help, but explaining a model trained on PCs to a business committee takes more effort.
- Variance ≠ relevance. PCA keeps the directions of greatest variance without knowing what you will use the data for: in a supervised problem, the signal separating the classes could live in a low-variance component you discarded. It is unsupervised in that sense too: it ignores any
y. - It inherits variance's sensitivity to outliers (02-01): one extreme customer can bend an entire component. Cleaning (03-01) or log (03-03) first.
And an important conceptual distinction from what you saw in 03-06:
| Feature selection (03-06) | PCA (compression) | |
|---|---|---|
| What it does | Picks a subset of the original columns | Creates new columns by blending them all |
| The discarded ones | Vanish from the model | Their information can survive blended into the PCs |
| Interpretability | Total: the columns are still recency_days, etc. |
Partial: abstract axes, interpreted via loadings |
| When to prefer | Irrelevant features, or interpretability comes first | Features redundant/correlated with each other |
They don't compete: often you first select the features that make sense and then compress the remaining redundancy.
Common Mistakes and Tips
- PCA without standardizing. The feature with the highest variance (through its units) hijacks PC1 and the reduction is a mirage.
StandardScalerfirst, always — except for the rare case of features already in naturally comparable units. - Keeping 2 components because the scatter looks nice. Two PCs are for visualizing; for modeling, decide using cumulative variance or the scree plot. Visualization and compression are different uses with different criteria.
- Interpreting a component's absolute sign. Re-run on another machine and PC1 may come out with all its signs flipped: it is the same axis. Interpret relative weights.
- Over-reading the 2D scatter. Two segments touching on the plane does not prove they touch in the full space: the dimension separating them may be exactly the one that is missing.
- Tip: always label the scatter's axes with their % of variance (
PC1 (46%)); it disciplines you and warns the reader how much information the drawing leaves out.
Exercises
Exercise 1. Without code: you have two standardized features with a correlation of 0.95. (a) Roughly what fraction of the variance will PC1 explain, and why? (b) What if the correlation were 0? (c) What does each case imply for compressing down to 1 dimension? Hint: with two standardized variables, the variance explained by PC1 is $(1+|r|)/2$.
Exercise 2. Apply PCA to the lesson's 7 customer features (or generate a correlated synthetic dataset). Draw the scree plot and the cumulative variance, and decide how many components you would keep in order to (a) visualize and (b) feed a K-means while retaining 90% of the variance. Justify both answers.
Exercise 3. Take the PC1 loadings from your run in exercise 2 and write, in one sentence of MercaFresh business language, what that axis measures. Then compute the Pearson correlation (02-03) between each customer's PC1 projection and their inactivity_ratio, and explain the result.
Solutions
Exercise 1
(a) With $r = 0.95$: PC1 explains $(1+0.95)/2 = 97.5%$. The cloud is an ellipse nearly degenerated into a line; almost all the variation happens along the diagonal. (b) With $r = 0$: each PC explains 50% — the cloud is circular and no direction is privileged; PCA has nothing to compress. (c) In the first case, reducing to 1D loses 2.5% of the information: excellent compression. In the second, it loses half: unacceptable. Moral: PCA compresses redundancy; without correlations there is no compression to be had — which is why the heatmap of 02-03 is a good prior diagnostic of whether PCA will pay off.
Exercise 2
pca = PCA().fit(X_esc)
var = pca.explained_variance_ratio_
plt.subplot(1, 2, 1); plt.plot(range(1, len(var)+1), var, "o-")
plt.xlabel("PC"); plt.ylabel("Explained variance")
plt.subplot(1, 2, 2); plt.plot(range(1, len(var)+1), np.cumsum(var), "o-")
plt.axhline(0.90, ls="--", c="gray")
plt.xlabel("Cumulative PCs"); plt.ylabel("Cumulative variance")
plt.show()(a) For visualizing: 2 components, by necessity — a scatter has two axes; the relevant question is how much variance they accumulate (here ~73%: acceptable, with the lesson's caveat). (b) For K-means: wherever the cumulative curve crosses 0.90 — typically 4 components on this dataset. PCA(n_components=0.90) automates it. Note that the answers differ: the criterion depends on the use.
Exercise 3
A sentence along these lines: "PC1 measures how switched-off the customer is: high values = a long time without buying and high relative inactivity; low values = a frequent, high-volume customer". The correlation between PC1 and inactivity_ratio comes out strongly positive (around +0.8/+0.9): that is coherent, because inactivity_ratio carries one of PC1's largest positive loadings, and each customer's projection onto the axis inherits that relationship. Checking loadings against simple correlations is a quick way to validate your reading of the axis.
Conclusion
You have added to your arsenal the tool that tames dimensionality: PCA rotates the axes toward the directions of maximum variance, turns correlated features into independent components ordered by importance, and lets you choose how many to keep using explained variance and the scree plot. You know why it demands standardization (it inherits the units through the covariance of 02-03), you know how to read loadings to give each axis a business name, and you have seen its two jobs at MercaFresh: a 2D map of the K-means segments and a compressor ahead of modeling — never confusing it with the feature selection of 03-06, because PCA does not pick columns: it blends them.
But PCA carries its built-in limit: it is a rotation, and rotations are linear. In the next lesson you will meet an algorithm that needs neither spherical groups nor straight structure: DBSCAN clusters by density, finds clusters of any shape and — what interests MercaFresh most — explicitly flags the points that fit into none of them: the anomalous orders.
Machine Learning Course
Module 1: Introduction to Machine Learning
- What is Machine Learning?
- History and evolution of Machine Learning
- Types of Machine Learning
- Applications of Machine Learning
- The Machine Learning project workflow
Module 2: Foundations of Statistics and Probability
- Basic statistics concepts
- Probability distributions
- Correlation and covariance
- Statistical inference
- Bayes' theorem
Module 3: Data Preprocessing
- Data cleaning
- Handling missing data
- Data transformation
- Encoding categorical variables
- Normalization and standardization
- Feature engineering
Module 4: Supervised Machine Learning Algorithms
- Linear regression
- Logistic regression
- Decision trees
- Support Vector Machines (SVM)
- K-Nearest Neighbors (K-NN)
- Naive Bayes
- Neural networks
Module 5: Unsupervised Machine Learning Algorithms
- Clustering: K-means
- Hierarchical clustering
- Principal Component Analysis (PCA)
- DBSCAN clustering
- Data visualization with t-SNE and UMAP
Module 6: Model Evaluation and Validation
- Data splitting: training, validation and test
- Evaluation metrics
- Cross-validation
- ROC curve and AUC
- Overfitting and underfitting
Module 7: Advanced Techniques and Optimization
- Regularization: Ridge, Lasso and Elastic Net
- Ensemble Learning
- Gradient Boosting
- Deep neural networks (Deep Learning)
- Hyperparameter optimization
Module 8: Model Implementation and Deployment
- Popular frameworks and libraries
- Deploying models to production
- Model maintenance and monitoring
- Ethical and privacy considerations
Module 9: Hands-On Projects
- Project 1: Housing price prediction
- Project 2: Image classification
- Project 3: Sentiment analysis on social media
- Project 4: Fraud detection
- Project 5: Customer segmentation
