In the previous lesson we drew the histogram of MercaFresh's order amounts and talked about its "shape": a central mass with a tail on the right. Probability distributions are the mathematical models of those shapes: they describe how the values a random variable can take are spread out. Understanding them is essential in Machine Learning because many algorithms assume the data follow certain distributions, and because comparing real data against an expected distribution is the foundation of anomaly detection. In this lesson you will meet the most important discrete and continuous distributions, study the normal in depth, and learn to generate and visualize them with NumPy and scipy.
Contents
- Random variables and probability functions
- Discrete distributions: Bernoulli, binomial, and Poisson
- Continuous distributions: uniform, normal, and exponential
- The normal distribution in detail: the 68-95-99.7 rule and z-scores
- Why distributions matter in Machine Learning
- Generation and visualization with NumPy and scipy
Random variables and probability functions
A random variable is a variable whose value depends on chance: we don't know what it will be in the next observation, but we can describe the probability of it taking each possible value.
Examples at MercaFresh:
- X = "will the next customer abandon their cart?" → takes values 0 or 1.
- Y = "number of orders arriving in the next hour" → 0, 1, 2, 3...
- Z = "amount of the next order" → any positive real value.
How we describe those probabilities depends on the type of variable (remember the classification from 02-01?):
| Discrete variable | Continuous variable | |
|---|---|---|
| Function | Probability mass function (PMF): P(X = k) | Probability density function (PDF): f(x) |
| Interpretation | Exact probability of each value | Density; the probability is the area under the curve over an interval |
| P(X = exact value) | Can be > 0 | Always 0 (a single point has no area) |
| Total sum/integral | Σ P(X=k) = 1 | ∫ f(x) dx = 1 |
One subtlety that trips people up at first: for a continuous variable, the question "what is the probability that the order is worth exactly €50.00?" has the answer 0. What does have an answer is "what is the probability that it is worth between €45 and €55?": the area under the density curve over that interval. To compute cumulative areas we use the cumulative distribution function (CDF): F(x) = P(X ≤ x), which scipy gives us out of the box.
Discrete distributions: Bernoulli, binomial, and Poisson
Bernoulli: a single yes/no experiment
It models one trial with two outcomes: success (1) with probability p, failure (0) with probability 1−p.
- MercaFresh: "a specific customer churns this month" with p = 0.08.
- It is the basic building block: the binomial and many classification models (such as the logistic regression in module 4) are built on top of it.
Binomial: counting successes in n trials
If we repeat n independent Bernoulli trials with the same p, the total number of successes follows a binomial B(n, p).
- MercaFresh: out of 500 customers, each with a churn probability of p = 0.08, how many will churn this month? → B(500, 0.08), with mean n·p = 40 customers.
from scipy import stats
# Probability that exactly 40 out of 500 customers churn?
print(stats.binom.pmf(k=40, n=500, p=0.08).round(4)) # 0.0657
# Probability that 50 or more churn? (1 - P(X <= 49))
print(1 - stats.binom.cdf(k=49, n=500, p=0.08)) # ~0.064stats.binom.pmf(k, n, p)returns P(X = k): the probability of exactly k successes.stats.binom.cdf(k, n, p)returns P(X ≤ k); subtracting it from 1 gives the probability of the upper tail. This1 - cdfpattern comes up constantly.
Business reading: even though the "expected" figure is 40 lost customers, in more than 6% of months there will be 50 or more through sheer chance. Knowing how much variation is normal prevents unjustified alarms.
Poisson: counting events per unit of time
It models the number of events in a fixed interval when they occur independently at an average rate λ (lambda).
- MercaFresh: if orders arrive at an average of λ = 12 per hour, what is the probability of receiving 20 or more in one hour (and overloading the delivery fleet)?
| Distribution | Models | Parameters | MercaFresh example |
|---|---|---|---|
| Bernoulli | One yes/no trial | p | Does this customer churn? |
| Binomial | Number of successes in n trials | n, p | Churned customers among 500 |
| Poisson | Number of events per interval | λ | Orders per hour |
Mnemonic: Bernoulli = one coin, binomial = n coins, Poisson = "how many times does the doorbell ring in an hour".
Continuous distributions: uniform, normal, and exponential
Uniform
Every value in an interval [a, b] is equally likely; its density is a flat plateau. It rarely shows up in real data, but it is the foundation of simulation (random generators produce uniforms and derive everything else from them) and an honest model of "I know nothing beyond this range".
Normal (Gaussian)
The symmetric bell curve defined by its mean μ (center) and its standard deviation σ (width). It is the queen of distributions thanks to the central limit theorem (preview: sums and averages of many small, independent effects tend to be normal; we will demonstrate it with a simulation in lesson 02-04). That is why so many aggregate quantities — measurement errors, sample means — are bell-shaped.
- MercaFresh: the warehouse preparation time of an order hovers around a normal with μ = 18 min and σ = 4 min.
Exponential
It models the time between events of a Poisson process. If λ = 12 orders/hour arrive, the time between two consecutive orders follows an exponential with mean 1/λ = 5 minutes. It is skewed with a right tail and "memoryless": the probability of waiting 5 more minutes does not depend on how long you have been waiting already.
# If 12 orders/hour arrive, probability of going more than 15 min with no orders?
scale = 60 / 12 # mean: 5 minutes between orders
print(1 - stats.expon.cdf(15, scale=scale)) # ~0.05Note the Poisson/exponential pairing: they tell the same story from two angles (how many events per interval vs. how much time between events).
| Distribution | Shape | Parameters | MercaFresh example |
|---|---|---|---|
| Uniform | Flat plateau | a, b | Simulations; random discount between 5 and 15% |
| Normal | Symmetric bell | μ, σ | Warehouse preparation time |
| Exponential | Decays from 0, right tail | λ (or its inverse, the mean) | Minutes between consecutive orders |
The normal distribution in detail
The 68-95-99.7 rule
If X follows a normal with mean μ and standard deviation σ:
- 68.3% of the values fall within μ ± 1σ.
- 95.4% fall within μ ± 2σ.
- 99.7% fall within μ ± 3σ.
For MercaFresh's preparation time (μ = 18, σ = 4):
- 68% of orders are prepared in 14 to 22 minutes.
- 95% in 10 to 26 minutes.
- 99.7% in 6 to 30 minutes. An order that takes 35 minutes to prepare is more than 4σ out: something odd has happened.
Z-scores: measuring in "number of standard deviations"
The z-score of a value x is:
z = (x − μ) / σ
It answers: how many standard deviations from the mean is this value? It is a universal "unusualness" ruler that does not depend on the units:
import numpy as np
mu, sigma = 18, 4
x = 31 # one order took 31 min to prepare
z = (x - mu) / sigma
print(z) # 3.25 -> highly atypical
# What proportion of orders takes more than 31 min?
print(1 - stats.norm.cdf(x, loc=mu, scale=sigma)) # ~0.0006- A |z| < 2 is normal territory; |z| > 3 is a serious anomaly candidate.
stats.norm.cdf(x, loc=mu, scale=sigma)gives P(X ≤ x) for the normal with meanlocand deviationscale; its complement says only 0.06% of orders would take that long by pure chance.
Here we are using the z-score as an oddity detector. In module 3 (lesson 03-05) it will reappear in another role: standardizing variables as a preprocessing step for models. Same formula, different purpose.
Why distributions matter in Machine Learning
- Model assumptions. Several algorithms work better (or only have guarantees) under certain distributions: linear regression (04-01) assumes normal errors, Gaussian Naive Bayes (04-06) assumes per-class normality, K-means (05-01) performs best with "round" Gaussian-like clusters. Knowing the true distribution of your data tells you whether those assumptions are reasonable.
- Anomaly detection. The classic approach is: model the distribution of what is normal and flag as anomalous whatever is highly improbable under it. For MercaFresh: if a customer's order amounts roughly follow a known distribution, an order with an extremely low probability (a huge z-score, the tail of the Poisson, etc.) deserves review before being accepted: it could be an error or fraud.
- Synthetic data and simulation. Generating data with controlled distributions (as we did in 02-01) lets you test pipelines and models before real data exist.
- Tails and transformations. Recognizing a right tail (amounts, times) hints that transforming the variable may pay off (e.g., with logarithms), a technique developed in lesson 03-03.
Generation and visualization with NumPy and scipy
Division of labor between libraries:
- NumPy (
np.random.default_rng()): generating random samples. - scipy.stats: computing PMF/PDF, CDF, quantiles, and fitting distributions.
- matplotlib: visualizing and comparing sample vs. theoretical curve.
import numpy as np
import matplotlib.pyplot as plt
from scipy import stats
rng = np.random.default_rng(7)
fig, axes = plt.subplots(2, 2, figsize=(11, 7))
# 1) Binomial: monthly churners among 500 customers (p=0.08)
churners = rng.binomial(n=500, p=0.08, size=2000)
axes[0, 0].hist(churners, bins=range(20, 65), color="steelblue", edgecolor="white")
axes[0, 0].set_title("Binomial: churners/month among 500 customers")
# 2) Poisson: orders per hour (lambda=12)
orders_per_hour = rng.poisson(lam=12, size=2000)
axes[0, 1].hist(orders_per_hour, bins=range(0, 30), color="darkorange", edgecolor="white")
axes[0, 1].set_title("Poisson: orders per hour")
# 3) Normal: preparation time, sample vs. theoretical density
prep = rng.normal(loc=18, scale=4, size=2000)
axes[1, 0].hist(prep, bins=40, density=True, color="seagreen",
edgecolor="white", alpha=0.7)
x = np.linspace(4, 32, 300)
axes[1, 0].plot(x, stats.norm.pdf(x, loc=18, scale=4), "k--", label="Theoretical PDF")
axes[1, 0].set_title("Normal: preparation minutes")
axes[1, 0].legend()
# 4) Exponential: minutes between orders (mean = 5)
between_orders = rng.exponential(scale=5, size=2000)
axes[1, 1].hist(between_orders, bins=40, density=True, color="indianred",
edgecolor="white", alpha=0.7)
x = np.linspace(0, 30, 300)
axes[1, 1].plot(x, stats.expon.pdf(x, scale=5), "k--", label="Theoretical PDF")
axes[1, 1].set_title("Exponential: minutes between orders")
axes[1, 1].legend()
plt.tight_layout()
plt.show()Key points about the code:
rng.binomial,rng.poisson,rng.normal,rng.exponentialgenerate samples from each distribution;size=2000requests 2,000 simulated observations.density=Truein the histogram normalizes the Y axis so the total area is 1, making the histogram comparable with the theoretical density curve we draw on top withstats.<dist>.pdf.- Overlaying the sample histogram with the theoretical PDF is the simplest visual check of "do my data look like this distribution?". In lesson 02-04 we will see formal tests that answer with numbers.
- Look at the shapes: the binomial and Poisson are nearly symmetric discrete bells (with these parameters), the normal is the perfect continuous bell, and the exponential decays from zero with a right tail, just like real times between orders.
Common Mistakes and Tips
- Reading density as probability. f(x) can exceed 1 (for example, in a narrow uniform); the probability is the area under the curve over an interval, not the height.
- Assuming normality by default. Amounts and times usually have a right tail (skewness); applying the 68-95-99.7 rule to non-normal data yields wrong conclusions. Draw the histogram first.
- Mixing up scipy's parameters.
scaleinstats.exponis the mean (1/λ), not λ. Andstats.normtakes the standard deviation (scale), not the variance. Always check the parameterization in the documentation. - Forgetting independence. The binomial requires independent trials with the same p. If one customer's churn drags along their neighbors' (word-of-mouth effect), the binomial model falls short.
- Using z-scores with heavily skewed distributions. A z-score of 3 in an exponential is not as rare as in a normal. The |z| > 3 rule presupposes a bell shape.
- Tip: memorize the "phenomenon → distribution" pairings in this lesson's tables; recognizing the right pattern in a new problem is 80% of the job.
Exercises
Exercise 1
MercaFresh sends an email campaign to 200 customers, each with a 5% conversion probability. (a) Which distribution does the number of conversions follow, and with which parameters? (b) Use scipy to compute the probability of getting exactly 10 conversions and (c) of getting 15 or more.
Exercise 2
Order preparation time follows a normal with μ = 18 min and σ = 4 min. Using only the 68-95-99.7 rule (no code): (a) between which values does the central 95% of orders fall? (b) Roughly what percentage takes more than 26 minutes? (c) One order took 6 minutes: what is its z-score and how do you interpret it?
Exercise 3
MercaFresh's logistics center receives an average of 12 orders per hour. Use NumPy to simulate 10,000 hours of activity and empirically estimate the probability of receiving 20 or more orders in one hour. Compare the result with the exact value of 1 - stats.poisson.cdf(19, mu=12).
Solutions
Solution 1
(a) Binomial B(n=200, p=0.05): 200 independent Bernoulli trials. Its mean is n·p = 10 conversions.
from scipy import stats
print(stats.binom.pmf(10, n=200, p=0.05).round(4)) # (b) ~0.1288
print((1 - stats.binom.cdf(14, n=200, p=0.05)).round(4)) # (c) ~0.0744(b) A 12.9% chance of exactly 10. (c) A 7.4% chance of getting 15 or more. Note: for the "15 or more" tail you subtract the CDF at 14, not at 15.
Solution 2
- (a) μ ± 2σ = 18 ± 8 → between 10 and 26 minutes.
- (b) Outside μ ± 2σ lies 5% split between two tails; the upper one alone: ≈ 2.5% of orders take more than 26 min.
- (c) z = (6 − 18) / 4 = −3. It sits three standard deviations below the mean: an order prepared suspiciously fast (was the time logged incorrectly? a single-item order?). Anomalies exist on the left side too.
Solution 3
import numpy as np
from scipy import stats
rng = np.random.default_rng(0)
hours = rng.poisson(lam=12, size=10_000)
print((hours >= 20).mean()) # ~0.0210 (varies with the seed)
print(1 - stats.poisson.cdf(19, mu=12)) # 0.0213 (exact)(hours >= 20) creates a boolean array and .mean() computes the proportion of True values: the frequency of the event in the simulation. With 10,000 repetitions, the empirical estimate (~2.1%) matches the theoretical one almost exactly. This "simulate and count" pattern is an extremely powerful tool when the exact formula does not exist or you don't remember it.
Conclusion
You now know what a random variable is and have met the six everyday working distributions: Bernoulli, binomial, and Poisson for counting; uniform, normal, and exponential for measuring. You have gone deep into the normal — the 68-95-99.7 rule and z-scores — and seen why distributions matter in ML: they underpin model assumptions and anomaly detection in MercaFresh's orders, and they enable simulations with NumPy and exact calculations with scipy. So far we have looked at each variable in isolation; the next step is studying how they relate to each other: in the next lesson, covariance and correlation will tell us which MercaFresh variables move together and which ones warn of churn.
Machine Learning Course
Module 1: Introduction to Machine Learning
- What is Machine Learning?
- History and evolution of Machine Learning
- Types of Machine Learning
- Applications of Machine Learning
- The Machine Learning project workflow
Module 2: Foundations of Statistics and Probability
- Basic statistics concepts
- Probability distributions
- Correlation and covariance
- Statistical inference
- Bayes' theorem
Module 3: Data Preprocessing
- Data cleaning
- Handling missing data
- Data transformation
- Encoding categorical variables
- Normalization and standardization
- Feature engineering
Module 4: Supervised Machine Learning Algorithms
- Linear regression
- Logistic regression
- Decision trees
- Support Vector Machines (SVM)
- K-Nearest Neighbors (K-NN)
- Naive Bayes
- Neural networks
Module 5: Unsupervised Machine Learning Algorithms
- Clustering: K-means
- Hierarchical clustering
- Principal Component Analysis (PCA)
- DBSCAN clustering
- Data visualization with t-SNE and UMAP
Module 6: Model Evaluation and Validation
- Data splitting: training, validation and test
- Evaluation metrics
- Cross-validation
- ROC curve and AUC
- Overfitting and underfitting
Module 7: Advanced Techniques and Optimization
- Regularization: Ridge, Lasso and Elastic Net
- Ensemble Learning
- Gradient Boosting
- Deep neural networks (Deep Learning)
- Hyperparameter optimization
Module 8: Model Implementation and Deployment
- Popular frameworks and libraries
- Deploying models to production
- Model maintenance and monitoring
- Ethical and privacy considerations
Module 9: Hands-On Projects
- Project 1: Housing price prediction
- Project 2: Image classification
- Project 3: Sentiment analysis on social media
- Project 4: Fraud detection
- Project 5: Customer segmentation
