With the binomial and the normal we cover a lot of ground, but not all of it: how many customers will arrive at a checkout in the next 10 minutes? (there is no fixed \( n \) of trials); how many customers will the Club Nova promoter have to approach before signing up the first member? (the number of attempts is precisely the unknown); how many defective units will turn up when inspecting 8 boxes out of a crate of 50? (without replacement, \( p \) changes with every draw); how long will it be until the next online order? (a continuous, skewed wait, where we already saw the normal fail). Each question has its template: Poisson, geometric, hypergeometric, uniform and exponential. This lesson introduces them one by one with the analyst's mindset — recognize the situation, identify the parameters, calculate — and ends with the summary table that will let you pick the right distribution at a glance.
Contents
- Poisson distribution: events per unit of time or space
- Poisson as a limit of the binomial: the fraud case
- Geometric distribution: attempts until the first success
- Hypergeometric distribution: sampling without replacement
- Continuous uniform distribution: ignorance spread evenly
- Exponential distribution: the time between events
- Summary table: which template to use in each situation
Poisson distribution: events per unit of time or space
The Poisson distribution models the number of events occurring in a fixed interval of time (or space), when the events occur independently, at a constant average rate, one at a time. NovaMarket examples: online orders per hour, customers arriving at a checkout in 10 minutes, delivery incidents per day, typos per catalog page.
Its only parameter is \( \lambda \) (lambda): the average number of events per interval. If \( X \sim \text{Poisson}(\lambda) \):
\[ P(X = k) = \frac{e^{-\lambda} , \lambda^k}{k!} \qquad k = 0, 1, 2, \dots \]
where \( e \approx 2.71828 \). Two very characteristic traits:
- \( X \) has no maximum: it can take any integer value from 0 upward (unlike the binomial, which stops at \( n \)).
- Mean and variance coincide: \( E(X) = \lambda \) and \( \operatorname{Var}(X) = \lambda \) (hence \( \sigma = \sqrt{\lambda} \)). This is a useful signature: if in your count data the variance is much larger than the mean, the Poisson probably does not fit.
- \( \lambda \) scales with the interval: if 6 orders arrive per hour, then in half an hour \( \lambda = 3 \) and in a 24-hour day, \( \lambda = 144 \).
Worked example: delivery incidents. The e-commerce service records on average \( \lambda = 3 \) delivery incidents per day (wrong addresses, absences, breakages), stable and independent of one another. The customer service team can absorb up to 4 incidents a day; with 5 or more, reinforcements are needed. How often will reinforcements be required?
We compute the tail via the complement, term by term (with \( e^{-3} \approx 0.0498 \)):
| \( k \) | Calculation | \( P(X=k) \) |
|---|---|---|
| 0 | \( e^{-3} \) | 0.0498 |
| 1 | \( e^{-3} \cdot 3^1 / 1! = 0.0498 \times 3 \) | 0.1494 |
| 2 | \( e^{-3} \cdot 3^2 / 2! = 0.0498 \times 4.5 \) | 0.2240 |
| 3 | \( e^{-3} \cdot 3^3 / 3! = 0.0498 \times 4.5 \) | 0.2240 |
| 4 | \( e^{-3} \cdot 3^4 / 4! = 0.0498 \times 3.375 \) | 0.1680 |
| Sum \( P(X \leq 4) \): | 0.8153 |
\[ P(X \geq 5) = 1 - 0.8153 = 0.1847 \]
Almost one day in five will call for reinforcements, even though the average is only 3. Once again, the message of Module 4: the expected value is not enough; the full distribution tells you how often the bad days arrive. (Notice too that the most likely outcome is not a single value: \( k = 2 \) and \( k = 3 \) are tied — this happens whenever \( \lambda \) is an integer.)
Poisson as a limit of the binomial: the fraud case
The Poisson does not fall out of the sky: it is the binomial pushed to the extreme of large \( n \) and small \( p \). When there are a huge number of trials, each with a minuscule probability of success, the binomial \( B(n, p) \) is very well approximated by a Poisson with \( \lambda = np \). The usual rule of thumb: \( n \geq 100 \) and \( p \leq 0.05 \) (the larger \( n \) and the smaller \( p \), the better).
A check with online fraud. The prevalence of fraud is \( p = 0.005 \). On a day with \( n = 400 \) orders, the number of fraud cases is \( B(400;\ 0.005) \), with \( \lambda = np = 2 \).
- Exact binomial: \( P(X = 0) = 0.995^{400} \approx 0.1347 \).
- Poisson: \( P(X = 0) = e^{-2} \approx 0.1353 \).
Difference: 6 ten-thousandths. The gain in convenience is enormous: no \( \binom{400}{k} \) at all; the Poisson answers with \( e^{-2} \lambda^k / k! \). For example, \( P(X \geq 4) = 1 - e^{-2}(1 + 2 + 2 + \tfrac{4}{3}) = 1 - 0.1353 \times 6.3333 \approx 1 - 0.8569 = 0.1431 \): 14% of 400-order days will bring 4 or more fraud attempts. This is why the Poisson is sometimes called "the law of rare events": fraud, accidents, complaints, breakages — individually unlikely, but with a great many opportunities to occur.
Geometric distribution: attempts until the first success
A change of question: we are no longer counting successes in \( n \) trials, but how many trials it takes until the first success. If the trials are independent Bernoulli trials with probability \( p \), the variable \( X \) = "the number of the attempt on which the first success arrives" follows a geometric distribution:
\[ P(X = k) = (1-p)^{k-1} , p \qquad k = 1, 2, 3, \dots \]
The logic is pure product rule: \( k-1 \) failures in a row and then a success. Its expected value is intuitive and powerful:
\[ E(X) = \frac{1}{p} \]
If a success has probability 1/8, on average you must wait 8 attempts.
Worked example: in-store recruitment. NovaMarket sets up a Club Nova recruitment stand in Barcelona-Gràcia. Experience from previous campaigns says that 1 in every 8 customers approached ends up becoming a member (\( p = 0.125 \)).
- Expected attempts until the first member: \( E(X) = 1/0.125 = 8 \) customers approached.
- Probability of signing them up exactly on the third attempt: \( P(X = 3) = 0.875^2 \times 0.125 = 0.7656 \times 0.125 \approx 0.0957 \).
- Probability of getting the first member within the first 5 contacts: the easiest route is the complement, "the first 5 all fail": \( P(X \leq 5) = 1 - 0.875^5 = 1 - 0.5129 = 0.4871 \).
The geometric has a famous property: memorylessness. If the promoter has had 10 failures in a row, the probability that the next customer says yes is still 0.125: previous failures do not "store up" luck. It is the formal antidote to the gambler's fallacy ("I'm due one") that we dismantled in Module 3.
Hypergeometric distribution: sampling without replacement
In the previous lesson we announced this model when listing the ways the binomial fails. The hypergeometric distribution describes the number of successes when drawing a sample without replacement from a small population: each draw changes the composition of the batch, so \( p \) is not constant and the trials are not independent.
Parameters: \( N \) = population size, \( K \) = number of "successes" it contains, \( n \) = size of the sample drawn. The probability of finding exactly \( k \) successes is a count of combinations (the classical approach of 03-01: favorable cases over possible cases):
\[ P(X = k) = \frac{\dbinom{K}{k} \dbinom{N-K}{n-k}}{\dbinom{N}{n}} \]
and its mean is what intuition dictates: \( E(X) = n \cdot \dfrac{K}{N} \).
Worked example: fruit quality control. A pallet arrives at Madrid-Chamberí with \( N = 50 \) boxes of strawberries, of which — unbeknownst to us — \( K = 5 \) are spoiled. The protocol inspects \( n = 8 \) boxes at random and rejects the pallet if any defective one appears. What is the probability that the pallet passes the inspection despite being contaminated?
Passing the inspection is \( X = 0 \). Rather than wrestling with large binomial coefficients, we chain conditional probabilities (product rule, Module 3), which is equivalent and more transparent:
\[ P(X=0) = \frac{45}{50} \times \frac{44}{49} \times \frac{43}{48} \times \frac{42}{47} \times \frac{41}{46} \times \frac{40}{45} \times \frac{39}{44} \times \frac{38}{43} \approx 0.4015 \]
(each factor is "the next box inspected is also fine", with one fewer good box at each step). So \( P(\text{reject}) = 1 - 0.4015 = 0.5985 \): the inspection only catches this level of contamination about 60% of the time — a key figure when deciding whether to inspect more boxes. The expected number of defective boxes in the sample is \( E(X) = 8 \times \frac{5}{50} = 0.8 \).
What about the binomial? With \( p = 5/50 = 0.10 \) it would give \( P(X=0) = 0.9^8 \approx 0.4305 \) — close but not equal, because the sample (8) is 16% of the batch (50), above the 10% threshold we gave as a rule of thumb. With large populations (8 receipts out of thousands) the two coincide and we use the binomial for convenience.
Continuous uniform distribution: ignorance spread evenly
The continuous uniform \( U(a, b) \) models a continuous variable about which all we know is that it falls in the interval \( [a, b] \), with no regions more likely than others: its density is a flat plateau of height \( 1/(b-a) \). It is the simplest of the continuous distributions, and probabilities are computed as proportions of length:
\[ P(c \leq X \leq d) = \frac{d - c}{b - a} \qquad E(X) = \frac{a+b}{2} \qquad \sigma = \frac{b-a}{\sqrt{12}} \]
NovaMarket example. The customer chooses a 2-hour delivery window for their online order (say, 17:00–19:00), and within the window the driver's arrival time is, for practical purposes, uniform: \( T \sim U(0, 120) \) minutes.
- Probability of arriving in the first half hour: \( \frac{30 - 0}{120} = 0.25 \).
- Expected wait from the start of the window: \( E(T) = 60 \) minutes, with \( \sigma = 120/\sqrt{12} \approx 34.6 \) min.
- Probability of arriving between 18:30 and 19:00: \( \frac{120 - 90}{120} = 0.25 \) — any 30-minute stretch is worth the same: that is the uniform's signature.
Its role in practice is twofold: an honest model when there truly is no directional information, and the starting point of simulations (random number generators produce uniforms that are then transformed into any other distribution).
Exponential distribution: the time between events
If events arrive according to a Poisson process at rate \( \lambda \) (per unit of time), the waiting time until the next event follows an exponential distribution. They are two sides of the same coin: the Poisson counts how many in an interval; the exponential measures how long until the next one.
It is continuous, takes only positive values, and its CDF (the cumulative probability we learned to read in 03-03) has a closed form:
\[ F(t) = P(T \leq t) = 1 - e^{-\lambda t} \qquad E(T) = \frac{1}{\lambda} \qquad \sigma = \frac{1}{\lambda} \]
Worked example: late-night online orders. In the overnight slot, NovaMarket's website receives orders at an average rate of \( \lambda = 6 \) orders/hour (one every 10 minutes, on average: \( E(T) = 1/6 \) h = 10 min).
- Probability that the next order arrives within 5 minutes: with \( t \) in hours, \( P(T \leq \tfrac{1}{12}) = 1 - e^{-6/12} = 1 - e^{-0.5} \approx 0.3935 \).
- Probability of more than 20 minutes with no orders: \( P(T > \tfrac{1}{3}) = e^{-6/3} = e^{-2} \approx 0.1353 \).
The exponential density is strongly right-skewed: a great many short waits and a tail of long ones — exactly the silhouette that, in the previous lesson, made the normal fail with delivery times. It also inherits from the geometric (its discrete cousin) the property of memorylessness: if 10 minutes have passed with no orders, the probability that the next one takes more than another 20 is still \( e^{-2} \); the system does not "accumulate" delay. This Poisson-exponential pair is the foundation of queueing theory, used to size checkouts, phone lines and delivery fleets: how many checkouts to open so that the average wait stays below a threshold is, at bottom, a problem of \( \lambda \)s. For now it is enough to know that this door exists.
Summary table: which template to use in each situation
| Distribution | Type | Question it answers | Parameters | \( E(X) \) | NovaMarket example |
|---|---|---|---|---|---|
| Bernoulli | Discrete | Success or failure in 1 trial? | \( p \) | \( p \) | Is this receipt paid by card? |
| Binomial | Discrete | How many successes in \( n \) independent trials? | \( n, p \) | \( np \) | Card-paid receipts among the next 10 |
| Poisson | Discrete | How many events in a fixed interval? | \( \lambda \) | \( \lambda \) | Delivery incidents per day |
| Geometric | Discrete | How many attempts until the first success? | \( p \) | \( 1/p \) | Customers approached until signing up a member |
| Hypergeometric | Discrete | How many successes when sampling without replacement? | \( N, K, n \) | \( nK/N \) | Defective boxes in the pallet inspection |
| Uniform | Continuous | Where does a value fall with no preferred regions? | \( a, b \) | \( (a+b)/2 \) | Arrival minute within the delivery window |
| Exponential | Continuous | How long until the next event? | \( \lambda \) | \( 1/\lambda \) | Minutes until the next online order |
| Normal | Continuous | How is a sum of many effects distributed? | \( \mu, \sigma \) | \( \mu \) | A store's daily sales |
The analyst's method, in three questions: am I counting or measuring? (discrete vs continuous); what exactly am I counting/measuring? (successes in a fixed \( n \) → binomial; events per interval → Poisson; attempts until a success → geometric; time until an event → exponential); do the assumptions hold? (independence, constant rate or \( p \), with or without replacement).
Common Mistakes and Tips
- Confusing binomial and Poisson. The binomial has a fixed, known \( n \) ("of today's 20 orders…"); the Poisson has no ceiling ("whatever incidents come up tomorrow…"). If you cannot say how many trials there are, it is not binomial.
- Forgetting to rescale \( \lambda \) when the interval changes. With 6 orders/hour, the Poisson for a quarter of an hour uses \( \lambda = 1.5 \), not 6. And in the exponential, \( \lambda \) and \( t \) must be in the same unit of time — mixing "per hour" with minutes is the most common calculation error of all.
- Miscounting the geometric. \( P(X = 3) \) is "success exactly on the 3rd" = \( q^2 p \), with two prior failures, not three. And \( E(X) = 1/p \) counts total attempts, including the successful one.
- Using the binomial where there is sampling without replacement from a small batch. If the sample exceeds ~10% of the population, the hypergeometric and the binomial diverge; use the hypergeometric (or the chain of conditionals, which is more transparent).
- Assuming uniformity out of laziness. "I don't know anything" does not always justify the uniform: customer arrivals over the course of a day have unmistakable peak hours. The uniform is a model, not an excuse.
- Tip: memorize each template's signature — Poisson: mean = variance; geometric and exponential: memoryless; uniform: equal stretches, equal probabilities. Recognizing the signature in real data is half the modeling job.
Exercises
Exercise 1. At peak time, an average of 4 customers arrive at the Madrid-Centro express checkout every 10 minutes, independently and at a steady rate. (a) Justify the choice of model and give its parameter. (b) Compute the probability that nobody arrives in the next 10 minutes (\( e^{-4} \approx 0.0183 \)). (c) The cashier can serve up to 5 customers in 10 minutes; compute the probability that a queue forms from excess arrivals (6 or more).
Exercise 2. At the Club Nova stand (\( p = 0.125 \) per customer approached): (a) What is the probability that the day's first member arrives exactly with the fourth customer approached? (b) How many customers should the promoter expect to approach to sign up the first one? (c) The promoter has had 15 noes in a row and declares: "the next one is almost certain to say yes, I'm due one". Assess the claim.
Exercise 3. Overnight, web orders arrive at rate \( \lambda = 6 \)/hour. (a) What distribution does the time until the next order follow, and what is its mean in minutes? (b) Compute \( P(T \leq 5 \text{ min}) \). (c) The on-call technician wants to take a 20-minute break right after dispatching an order. What is the probability that no order comes in during the break? Solve it two ways: with the exponential and with the Poisson for the interval.
Solutions
Solution 1. (a) A count of events in a fixed interval, steady average rate, arrivals independent and one at a time: Poisson with \( \lambda = 4 \) (per 10-minute block). (b) \( P(X=0) = e^{-4} \approx 0.0183 \). (c) \( P(X \geq 6) = 1 - P(X \leq 5) \). Terms: 0.0183; 0.0733; 0.1465; 0.1954; 0.1954; 0.1563 → \( P(X \leq 5) \approx 0.7852 \), so \( P(X \geq 6) \approx 0.2148 \). One in five 10-minute blocks overwhelms the cashier. Common mistake: setting it up as a binomial — there is no fixed "number of candidate customers".
Solution 2. (a) Geometric: \( P(X = 4) = 0.875^3 \times 0.125 = 0.6699 \times 0.125 \approx 0.0837 \). (b) \( E(X) = 1/0.125 = 8 \) customers. (c) False: by the memorylessness of the geometric, after 15 failures the probability for the next customer is still 0.125. Past attempts do not load the future — it is the gambler's fallacy. (A different matter is that a very long streak might invite you to question whether \( p = 0.125 \) is right today; testing that formally is Module 5 material.)
Solution 3. (a) Exponential with \( \lambda = 6 \)/h; \( E(T) = 1/6 \) h = 10 minutes. (b) \( P(T \leq 5 \text{ min}) = 1 - e^{-6 \times 5/60} = 1 - e^{-0.5} \approx 0.3935 \). (c) Exponential route: \( P(T > 20 \text{ min}) = e^{-6 \times 20/60} = e^{-2} \approx 0.1353 \). Poisson route: in 20 min, \( \lambda = 6 \times \tfrac{20}{60} = 2 \), and \( P(X = 0) = e^{-2} \approx 0.1353 \). Both agree — they are the same coin seen from its two sides. Common mistake: using \( \lambda = 6 \) with \( t = 20 \) without converting to the same unit, producing the absurd \( e^{-120} \).
Conclusion
Your toolbox is complete: Bernoulli and binomial for successes in counted trials, Poisson for events per interval (and as a shortcut for the rare-event binomial), geometric for the first success, hypergeometric for sampling without replacement, uniform for flat uncertainty and exponential for waiting times — with the normal reigning over sums of many effects. The summary table is your map; each template's assumptions, your compass. One piece remains, the one that ties everything together and turns this module into the gateway to inference: why does the normal appear even when the underlying data are not normal — like receipt amounts? And what distribution does the mean of a sample follow, that number we have been using since Module 1? The answer is the most important result in all of statistics: the central limit theorem.
Statistics Course
Module 1: Introduction to Statistics
Module 2: Describing Data
- Measures of Central Tendency
- Measures of Dispersion
- Measures of Position and Outliers
- Graphical Representation of Data
Module 3: Probability
Module 4: Probability Distributions
- The Binomial Distribution
- The Normal Distribution
- Other Important Distributions
- The Central Limit Theorem
Module 5: Statistical Inference
Module 6: Data Analysis
- Correlation Analysis
- Regression Analysis
- Analysis of Variance (ANOVA)
- Categorical Data Analysis: Chi-Square
