The MercaFresh churn model is now in production: nightly batch, API, Docker, gradual deployment. It would be tempting to call the project finished — and it would be a mistake. A machine learning model is one of the few software components that degrades without anyone touching the code: the model is a snapshot of the patterns of the past, and the world that generates the data keeps moving. In this lesson you will learn why this happens (data drift and concept drift), how to detect it with tools you already know (histograms, percentiles and the statistical tests from module 2), what to monitor on a dashboard, and how to organize retraining and versioning so the model stays useful year after year.
Contents
- Why models degrade
- Data drift: what comes in changes
- Concept drift: the relationship with the target changes
- Detecting drift in the inputs
- Monitoring predictions and metrics (with labels that arrive late)
- Dashboards and alerts: what to watch
- Retraining: when and how
- Model registry and versioning
- The complete lifecycle in production
Why models degrade
When you trained the churn model, module 6 insisted on one condition for the test metrics to be credible: production data must resemble training data. That condition holds on deployment day… and starts eroding the day after. The causes fall into two phenomena with names of their own:
| Phenomenon | What changes | Formally | MercaFresh example |
|---|---|---|---|
| Data drift | The distribution of the inputs X | P(X) changes; the X→y relationship may stay the same | MercaFresh opens in a new city: customers arrive with short tenure and different habits |
| Concept drift | The relationship between inputs and target | P(y|X) changes, even if P(X) does not | An aggressive competitor campaign makes previously loyal customers start leaving |
The distinction matters because they are detected and handled differently, as we will see. And they share a dangerous trait: the service keeps working. The API responds in milliseconds, the nightly batch finishes without errors, the probabilities look fine… and they get worse and worse. Without dedicated monitoring, a model's degradation is invisible until the business notices it.
Data drift: what comes in changes
Data drift (also called covariate shift) is the most common case: the population you predict on stops resembling the training population.
Concrete examples at MercaFresh:
- Expansion to a new city: the model was trained on customers from five cities; now 15% of requests come from Zaragoza, a category the pipeline's encoder never even knew about. The distribution of
months_tenurecollapses (everyone is a new customer). - Changing habits: after launching the monthly subscription, the distribution of
frequency_90dshifts upward and that ofrecency_daysdownward. The model sees values in regions where it had few examples to learn from. - Upstream technical changes: someone modifies the source system and
avg_order_spendswitches from euros to cents. Brutal, instantaneous drift — and surprisingly frequent: many "model degradations" are actually data bugs.
The model does not fail all at once: it simply extrapolates further and further from what it knew, and its reliability drops gradually and silently.
Concept drift: the relationship with the target changes
With concept drift, the inputs may look the same as ever, but what they mean for the target has changed: the function the model learned is no longer the one governing the world.
Examples at MercaFresh:
- A campaign changes who leaves: the retention team systematically contacts high-risk customers (using our model!) and many of them stay. Now a profile that historically meant "almost certain churn" no longer does. Note the irony: the model's success alters the reality the model was describing — ML systems that act on the world generate this loop naturally.
- A new competitor with free delivery: customers with a "loyal" profile (high frequency, low recency) start leaving. The X's have not changed; the meaning of the X's has.
- Seasonality: the Christmas buying pattern does not mean the same thing in January. If training did not cover full annual cycles, every new season is a small concept drift.
Concept drift is harder to detect than data drift, because looking at the inputs is not enough: you need to compare predictions with real outcomes — and those arrive late, as we will see.
Detecting drift in the inputs
The good news: to detect data drift you do not need labels, only to compare the current distribution of each variable with the reference one (the training data). And the tools are the ones from module 2.
Level 1 — Descriptive statistics and percentiles (02-01): alongside the model, store the mean, standard deviation and percentiles (p5, p25, p50, p75, p95) of each variable at training time. Every night, compute the same numbers on the scoring data and compare. It is simple, cheap and catches the gross cases (the euros/cents bug jumps out in the median).
Level 2 — Statistical tests (02-04): to detect finer shifts, formalize the comparison as a hypothesis test: H0 = "both samples come from the same distribution".
- Numeric variables → Kolmogorov-Smirnov test (compares the full cumulative distributions).
- Categorical variables → chi-square test on the frequency of each category.
import numpy as np
import pandas as pd
from scipy import stats
# Reference: the data the model was trained on (stored next to the model)
X_ref = pd.read_parquet("training_reference.parquet")
# Current: tonight's scoring data
X_cur = pd.read_parquet("customer_features_today.parquet")
# --- Numeric: Kolmogorov-Smirnov ---
for col in ["recency_days", "frequency_90d", "avg_order_spend", "months_tenure"]:
ks = stats.ks_2samp(X_ref[col], X_cur[col])
flag = " <-- DRIFT" if ks.pvalue < 0.01 else ""
print(f"{col:20s} KS={ks.statistic:.3f} p={ks.pvalue:.4f}{flag}")
# --- Categorical: chi-square on the frequency table ---
freq_table = pd.concat([
X_ref["city"].value_counts(), X_cur["city"].value_counts()
], axis=1, keys=["ref", "cur"]).fillna(0)
chi2, pvalue, _, _ = stats.chi2_contingency(freq_table.T)
print(f"city: chi2={chi2:.1f} p={pvalue:.4f}")Two nuances of professional use:
- With large samples, everything comes out "significant" (you saw this in 02-04: with a huge n, trivial differences yield minuscule p-values). That is why monitoring also looks at the effect size — the KS statistic (the maximum separation between the cumulative distributions) more than its p-value, with agreed practical thresholds (e.g. investigate if KS > 0.1).
- The test signals that there is change, not whether it matters. A large drift in a variable of little importance to the model is less urgent than a moderate one in the main variable. Crossing drift with feature importance (07-03) prioritizes the alerts.
Monitoring predictions and metrics (with labels that arrive late)
Watching the inputs is not enough; you also have to watch what comes out of the model, on two levels depending on what is available:
Level A — The distribution of the predictions (always available, instantly). Every night, store the histogram and percentiles of prob_churn over the whole customer base. If the average churn probability goes from 0.18 to 0.31 in two weeks, or the percentage of customers above the action threshold doubles, something has changed — in the data, in the world, or in the pipeline itself — and it must be investigated before knowing whether the predictions are right. This is the canary in the coal mine: cheap and with zero delay.
Level B — The real metrics (available with a delay). Here appears the difficulty specific to churn: the true label takes months to become known. If you define churn as "no orders in 90 days", the prediction you make today can only be evaluated 90 days from now. Practical consequences:
- Organize evaluation by cohorts: "the scores issued in May" are evaluated in August, when their labels mature, computing precision, recall, AUC and the confusion matrix from 06-02 on that cohort.
- Plot each metric per cohort over time: a sustained downward slope is the signature of degradation (and, if the inputs show no drift, it points to concept drift).
- Beware of intervention bias: if the retention team acts on the flagged customers, their labels no longer reflect what would have happened without acting. Keeping a small uncontacted control group (as in the A/B test of 08-02) keeps the evaluation honest.
Complete monitoring combines both levels: A gives early warnings without confirmation; B gives the truth, months late.
Dashboards and alerts: what to watch
Everything above materializes in a dashboard anyone on the team can read, with automatic alerts when something crosses a threshold. What it must contain, in two blocks:
| Block | Indicator | Frequency | Typical alert |
|---|---|---|---|
| Technical (service) | API latency (p95), error rate, requests/min, nightly batch duration | Real time / daily | p95 > 200 ms; batch not finished by 6:00 |
| Technical (data) | % missing values per variable, new categories, KS/chi² vs. reference | Daily | KS > 0.1 on an important variable; unknown category > 1% |
| Technical (model) | Distribution of prob_churn (mean, percentiles), % above the threshold |
Daily | Mean outside the agreed band |
| Business | Precision/recall per mature cohort, AUC per cohort | Monthly | Cohort recall drops 5 points vs. historical average |
| Business | Retention rate of contacted customers vs. control group, cost per retained customer | Monthly | The difference from the control stops being significant |
Two design principles:
- Every alert must have an owner and an action. An alert nobody attends to trains the team to ignore alerts. Few, well calibrated, actionable.
- The business dashboard rules. The model can have a stable AUC and still be failing the business (or vice versa). The goal was never AUC: it was retaining customers within a budget (06-04).
Retraining: when and how
Once deterioration is detected (or before it arrives), the natural response is retraining on recent data. The two key questions:
When to retrain? Two strategies, not mutually exclusive:
| Strategy | How it works | Advantages | Drawbacks |
|---|---|---|---|
| On a schedule | Every N months, retrain on the most recent data window | Predictable, simple to operate, keeps the process "well-oiled" | May retrain without need, or arrive late to an abrupt change |
| Drift-triggered | Drift/metric alerts launch (or request) the retraining | Reacts to the real need | Requires mature monitoring and well-calibrated thresholds |
In practice, the combination is the norm: a scheduled quarterly retraining, plus the option of bringing it forward if the alerts justify it. For MercaFresh: quarterly retraining of the churn model, movable up in response to events such as opening in a new city.
How to retrain? It is not "run fit again and off you go":
- Data window: decide which history to train on. The full history (more data, but drags in old patterns) or a rolling window of the last 18–24 months (fresher, less volume)? With confirmed concept drift, the recent window usually wins; it can be compared empirically.
- The same rigor as the original: retraining repeats the complete process of modules 6 and 7 — proper temporal split, CV, hyperparameter search over the pipeline — in an automated way (a script or training pipeline, not a hand-run notebook).
- Validation before replacing: the candidate must beat the production model on a common, recent evaluation set neither saw during training. If it does not, it is not deployed — retraining does not guarantee improvement.
- Gradual deployment and rollback: the approved candidate goes in through the same path you saw in 08-02 (shadow → A/B → rollout), with the previous version ready to come back. Retraining is not a special event: it is another deployment.
Model registry and versioning
With periodic retraining, within a year you will have v3, v4, v5… and questions like "which model generated the scores of March 12?" or "what data was the one that failed trained on?". Answering them demands record-keeping discipline. For each version, store at minimum:
- The artifact (serialized pipeline) and its checksum.
- Data: the time window and the query/snapshot it was trained on.
- Code and environment: the repository commit and the exact
requirements.txt. - Validation metrics and the comparison with the previous version.
- Deployment history: when it entered production, when it left, and why.
This can start as a folder convention plus a metadata file (as in 08-02). When the volume grows, a dedicated model registry is used: tools like MLflow offer exactly this — experiment tracking (parameters and metrics of each training run), a central model registry with stages (staging/production/archived) and traceability between data, code and artifact. Conceptually it adds nothing you have not seen: it systematizes this section's discipline so it does not depend on anyone's memory. We do not develop it here; its documentation is approachable with what you already know.
The complete lifecycle in production
The diagram that summarizes the module so far — and explains why this practice is sometimes called MLOps: treating the entire cycle as a continuous engineering process, not a project with an end.
graph TB
A[Training and validation<br/>modules 3-7] --> B[Gradual deployment<br/>08-02: shadow, A/B, rollout]
B --> C[Service in production<br/>batch + API]
C --> D[Continuous monitoring<br/>data, predictions, metrics, business]
D -- all stable --> C
D -- drift or degradation --> E{Diagnosis}
E -- data bug --> F[Fix upstream] --> C
E -- real drift --> G[Retraining<br/>window + validation]
G -- better candidate --> B
G -- candidate falls short --> H[Investigate further<br/>new features, redesign] --> A
C -. rollback on incident .-> B
Notice that it is a cycle, not a line: the normal state of a useful model is going around this graph for years.
Common Mistakes and Tips
- "Deployed = done". The mindset error all the others are born from: without monitoring, the first news of degradation will come from the business, months late.
- Monitoring only the service (latency, errors) and not the model. A 100% healthy API can be serving ever-worse predictions. They are two different kinds of monitoring and both are needed.
- Trusting the p-value with huge samples. With a hundred thousand rows, a KS with p < 0.001 can reflect an irrelevant difference. Look at the effect size and set practical thresholds.
- Forgetting label delay. Evaluating "this week's recall" with immature churn labels underestimates real churn and gives a false sense of accuracy. Evaluate on mature cohorts.
- Retraining and deploying without validating against the current model. Retraining on drifted data does not guarantee a better model; without the prior comparison, you may replace a mediocre model with a worse one.
- Ignoring the effect of your own interventions. If you act on the customers the model flags, the subsequent labels are contaminated by your intervention; without a control group, the model will seem to "get it wrong" precisely when it works.
- Tip: from day one, store a reference snapshot (training data with its statistics) next to the artifact. Without a reference there is no comparison, and rebuilding it months later is usually impossible.
Exercises
Exercise 1. Classify each situation as data drift, concept drift or data bug, and justify it: (a) after integrating a new logistics provider, delivery_incidents starts arriving almost always as 0 because the provider does not report incidents; (b) MercaFresh launches a loyalty card with free delivery and, with the same RFM profiles as ever, the real churn rate of high-frequency customers drops by half; (c) an acquisition campaign on social media brings a wave of young customers with small baskets, a minority segment in the training data.
Exercise 2. MercaFresh's nightly monitoring applies KS to avg_order_spend with the reference data (n = 80,000) versus the day's data (n = 75,000) and gets KS = 0.012 with p = 0.00004. The person in charge proposes retraining immediately. What would you tell them?
Exercise 3. Design in 5-7 lines the continuous evaluation plan for the churn model, knowing that the label is defined as "no orders in 90 days": what is monitored every day, what is computed every month, and what role the control group plays.
Solutions
Solution 1. (a) Data bug (even though it shows up as drift): the variable has not changed in the world, it has stopped being measured properly. The action is to fix the integration upstream, not to retrain — retraining would learn that "0 incidents" means nothing. (b) Concept drift: P(X) barely changes (same RFM profiles), but the X→y relationship does — the same profile now churns less. It will only be detected in the per-cohort metrics, not in the tests on the inputs. (c) Data drift: the composition of the input population (P(X)) changes; the profile→churn relationship may remain the same, but the model is now frequently predicting in a region where it had few examples. Watch that segment's metrics and consider retraining with it better represented.
Solution 2. That they should distinguish significance from relevance (02-04): with n ≈ 80,000 per sample, the KS test flags any minuscule difference as "significant"; the effect size is KS = 0.012 — the cumulative distributions separate by at most 1.2%, a change almost certainly irrelevant to the model. Before retraining: compare percentiles to see where the difference lies, check whether the variable is important to the model, and review the other indicators (prediction distribution, cohort metrics). Retraining over a p-value with a negligible effect is cost and risk with no benefit.
Solution 3. Sample plan: Daily — statistics and KS/chi² of every variable against the reference; distribution of prob_churn (mean and percentiles) and % of customers above the threshold; batch and API health. Monthly — the cohort of scores issued 90 days ago is "closed" (its labels have matured) and its precision, recall and AUC are computed, adding them to the metric time series to see the trend. Control group — a small random fraction of at-risk customers is not contacted; their labels make it possible to measure churn "without intervention", so the model's evaluation is not distorted by the campaigns' success, and as a bonus it measures the real causal benefit of retention against the control.
Conclusion
You now know why a model that is perfect on deployment day stops being so: data drift moves the inputs (the new city, the changing habits) and concept drift rewrites the relationship with the target (the campaign that changes who leaves). And you know how to keep watch with a layered system: statistics and KS/chi-square tests on the inputs every night, the prediction distribution as the canary, the real metrics by cohort once the labels mature, and a dashboard where business and engineering read together — all feeding a validated, versioned retraining cycle with rollback, which turns the model into a living system. One dimension remains that no dashboard measures on its own: the churn model decides which people get contacted, with which offers, using which data — and that raises questions of bias, fairness, transparency and privacy. We devote the final lesson of the module to them: the ethical considerations of putting machine learning in front of real people.
Machine Learning Course
Module 1: Introduction to Machine Learning
- What is Machine Learning?
- History and evolution of Machine Learning
- Types of Machine Learning
- Applications of Machine Learning
- The Machine Learning project workflow
Module 2: Foundations of Statistics and Probability
- Basic statistics concepts
- Probability distributions
- Correlation and covariance
- Statistical inference
- Bayes' theorem
Module 3: Data Preprocessing
- Data cleaning
- Handling missing data
- Data transformation
- Encoding categorical variables
- Normalization and standardization
- Feature engineering
Module 4: Supervised Machine Learning Algorithms
- Linear regression
- Logistic regression
- Decision trees
- Support Vector Machines (SVM)
- K-Nearest Neighbors (K-NN)
- Naive Bayes
- Neural networks
Module 5: Unsupervised Machine Learning Algorithms
- Clustering: K-means
- Hierarchical clustering
- Principal Component Analysis (PCA)
- DBSCAN clustering
- Data visualization with t-SNE and UMAP
Module 6: Model Evaluation and Validation
- Data splitting: training, validation and test
- Evaluation metrics
- Cross-validation
- ROC curve and AUC
- Overfitting and underfitting
Module 7: Advanced Techniques and Optimization
- Regularization: Ridge, Lasso and Elastic Net
- Ensemble Learning
- Gradient Boosting
- Deep neural networks (Deep Learning)
- Hyperparameter optimization
Module 8: Model Implementation and Deployment
- Popular frameworks and libraries
- Deploying models to production
- Model maintenance and monitoring
- Ethical and privacy considerations
Module 9: Hands-On Projects
- Project 1: Housing price prediction
- Project 2: Image classification
- Project 3: Sentiment analysis on social media
- Project 4: Fraud detection
- Project 5: Customer segmentation
