The previous module ended with a promise: to stop learning tools and start using them the way they are used in real work — facing a complete case, with nobody whispering which method to apply. That is exactly what this lesson is. You will follow NovaMarket's analytics team through four business cases from start to finish: an A/B test on the website, inventory management under uncertain demand, statistical control of delivery times, and the rigorous measurement of a marketing campaign. In each one you will walk the full cycle — question, data, method, result, decision — calling on the tools from the previous seven modules by name. There is no new theory here: there is craft.

Contents

  1. The question → data → method → decision cycle
  2. Case 1: a complete A/B test on the NovaMarket website
  3. Case 2: inventory under uncertain demand
  4. Case 3: statistical process control for deliveries
  5. Case 4: measuring a marketing campaign properly
  6. Synthesis: the business analyst's method map

The question → data → method → decision cycle

In modules 1 to 7 the order was: "here is the tool, here is an example". In business the order is reversed, and skipping steps is expensive:

  1. Business question. Phrased in euros, customers or hours — never in statistical jargon. "Does the new checkout sell more?", not "is the difference in proportions significant?".
  2. Data. Does it exist? How is it collected? This is where the biases from Data Collection and Sampling live: a convenient but biased data point is worse than no data at all.
  3. Method. Chosen after the question and the data, never before. Half of professional judgment is this matching.
  4. Result. The number, with its uncertainty: an interval, a p-value, a forecast error.
  5. Decision. Translating the result into an action with its cost and its risk. If the analysis could not change any possible decision, it was not worth doing.

The four cases in this lesson walk the entire cycle. Notice that the "hard" part is almost never the calculation: it is getting steps 1, 2 and 5 right.

Case 1: a complete A/B test on the NovaMarket website

The question

The e-commerce team has redesigned the checkout: fewer steps, express payment for Club Nova members. Marta's question: does the new checkout raise conversion enough to justify the change? Current conversion (sessions that end in an order) is 3.2%.

An A/B test is exactly the randomized experiment from Data Collection and Sampling: sessions are randomly assigned to version A (current) or B (new), so that both groups are comparable in everything except the checkout. Without randomization, any difference could be down to the day of the week, the device, or the type of customer.

Design: how much sample and for how long

Before launching anything, the team fixes three things (as Errors, Power and Sample Size demands):

  • Minimum relevant effect: +0.4 percentage points (from 3.2% to 3.6%). Below that, the change does not pay for the development cost.
  • Significance level: \(\alpha = 0.05\) two-sided.
  • Power: 80% (\(z_{\beta} = 0.84\)).

The size per group for comparing two proportions:

\[ n = \frac{(z_{\alpha/2} + z_{\beta})^2 ,[p_1(1-p_1) + p_2(1-p_2)]}{(p_2 - p_1)^2} = \frac{(1.96 + 0.84)^2 (0.032 \cdot 0.968 + 0.036 \cdot 0.964)}{0.004^2} \]

\[ n = \frac{7.84 \times 0.0657}{0.000016} \approx 32,200 \text{ sessions per group} \]

The website gets around 330,000 sessions a week. Assigning 10% of the traffic (5% to each arm) for two weeks would be enough: \(330,000 \times 2 \times 0.05 = 33,000\) per group. The team chooses two full weeks, not ten scattered days: conversion has weekly seasonality (a lesson from Time Series) and cutting mid-cycle would bias the comparison. One more golden rule: sample size and duration are fixed before starting, and nobody peeks at the result every day to stop "as soon as it comes out significant" — that inflates the type I error, the p-hacking that the Hypothesis Testing lesson already warned you about.

Result

After two weeks:

Version Sessions Orders Conversion
A (current) 33,000 1,056 3.20%
B (new) 33,000 1,208 3.66%

Two-proportion test (\(H_0: p_A = p_B\)). Pooled proportion: \(\hat{p} = 2264/66,000 = 0.0343\).

\[ z = \frac{0.0366 - 0.0320}{\sqrt{0.0343 \cdot 0.9657 \left(\tfrac{1}{33,000} + \tfrac{1}{33,000}\right)}} = \frac{0.0046}{0.00142} \approx 3.25 \]

Two-sided p-value \(\approx 0.001 < 0.05\): \(H_0\) is rejected. And the 95% confidence interval for the difference (Confidence Intervals):

\[ 0.0046 \pm 1.96 \times 0.00142 ;\Rightarrow; (+0.18 \text{ pp};; +0.74 \text{ pp}) \]

Decision and practical relevance

Significant is not enough: it has to be translated into euros — the lesson of practical relevance versus statistical significance. With ~17 million sessions a year and an average online order of €52:

  • Central scenario (+0.46 pp): around 78,000 more orders a year, ~€4 million in additional revenue.
  • Cautious scenario (the lower bound of the CI, +0.18 pp): ~30,600 orders, ~€1.6 million.

Even the worst plausible case far exceeds the development cost. Decision: roll out checkout B to 100% of traffic, keeping the monitoring running for a few weeks in case the effect fades (novelty rather than genuine improvement).

Case 2: inventory under uncertain demand

The question

Every stockout is a lost sale and an annoyed customer; every excess is capital tied up and, in fresh produce, shrinkage. The question: how much stock should we hold to cover demand 95% of the time, without overstocking? Demand is a random variable, so the answer is a percentile, not a mean.

High-turnover item: the normal distribution and safety stock

Fresh milk at Madrid-Centro sells an average of 180 units/day with standard deviation 35, and the histogram is reasonably symmetric: the normal model fits. The supplier takes 4 days to deliver, so what matters is demand over those 4 days. Summing 4 independent days:

\[ \mu_L = 4 \times 180 = 720 \qquad \sigma_L = 35\sqrt{4} = 70 \]

(the standard deviation grows with \(\sqrt{n}\), not with \(n\) — the same square root behind the Central Limit Theorem). For a 95% service level, the order is placed when stock drops to the 95th percentile of that demand:

\[ \text{Reorder point} = \mu_L + z_{0.95},\sigma_L = 720 + 1.645 \times 70 \approx 836 \text{ units} \]

The \(1.645 \times 70 \approx 116\) units above the mean are the safety stock: the price of uncertainty. The table shows what each extra point of service costs:

Service level \(z\) Safety stock Reorder point
90% 1.282 90 810
95% 1.645 116 836
99% 2.326 163 883

Going from 95% to 99% costs 47 more units for a single item in a single store: the service level is an economic decision, not a dogma. That is why NovaMarket sets 99% for image-critical staples (milk, bread) and 90% for substitutable items.

Low-turnover item: Poisson

The premium food processor sells an average of \(\lambda = 0.7\) units per week in Cuenca, with weekly replenishment. With counts this low the normal is a poor model (it would allow negative demand!); the right model is Poisson. Accumulating probabilities:

\(k\) \(P(X = k)\) \(P(X \leq k)\)
0 0.4966 0.4966
1 0.3476 0.8442
2 0.1217 0.9659
3 0.0284 0.9943

With 2 units in store, 96.6% of weeks are covered; with 1, only 84.4%. Decision: a stock of 2. Notice how misleading the normal approximation would have been (\(0.7 + 1.645\sqrt{0.7} \approx 2.1\), almost the same here, but it diverges as \(\lambda\) gets even smaller) and, above all, how discrete the decision is: 1.4 food processors do not exist.

Case 3: statistical process control for deliveries

The question

You already know from module 7 that online deliveries have a median of 25.5 h and 13% outside the SLA. Logistics does not just want the snapshot: it wants an alarm that goes off when the process deteriorates, without going off at every normal fluctuation. That is the idea behind the control chart, invented by Shewhart: separating the common causes of variation (the noise inherent to the process) from the special causes (something has changed and must be investigated).

Building the chart for the mean

With the process stable, LogiExpress delivery time has mean \(\mu = 26.5\) h and standard deviation \(\sigma = 6.0\) h. Each day, the mean time of a sample of \(n = 36\) deliveries is taken. By the Central Limit Theorem, that daily mean fluctuates with

\[ \sigma_{\bar{x}} = \frac{6.0}{\sqrt{36}} = 1.0 \text{ h} \]

The control limits sit at \(\pm 3\sigma_{\bar{x}}\) from the centre line — so far out that, if the process has not changed, only 0.27% of points will cross them by chance (the 68-95-99.7 rule of the normal):

\[ LCL = 26.5 - 3 \times 1.0 = 23.5 \text{ h} \qquad UCL = 26.5 + 3 \times 1.0 = 29.5 \text{ h} \]

It is, at heart, a hypothesis test repeated every day with \(\alpha \approx 0.0027\): that demanding precisely because it is tested daily and false alarms burn the team out.

A two-week reading

Daily means (h): 26.1, 27.3, 25.8, 26.9, 27.8, 28.2, 28.6, 28.9, 29.1, 30.2.

  • Days 1–4: normal fluctuation around 26.5. Touch nothing.
  • Days 5–9: all above the centre line and climbing steadily — a run like that is already suspicious even though no point crosses the limit (supplementary rules such as the Western Electric rules flag 8 consecutive points on the same side, or sustained trends).
  • Day 10: 30.2 > 29.5. An unmistakable signal of a special cause.

Investigation: LogiExpress has reorganized its Zaragoza hub and the northern routes are leaving 4 h late. Decision: a meeting with the carrier with the chart on the table, and a contingency plan with RápidoSur for the affected routes.

Two textbook warnings: first, do not over-adjust — reacting to every individual point within the limits ("27.3 yesterday, do something!") adds variability instead of removing it; second, the chart watches stability, not quality: a process can be perfectly in statistical control and still miss the SLA — that gets fixed by redesigning the process, not by watching it.

Case 4: measuring a marketing campaign properly

The question and the trap

Marketing sends a €5 coupon — the same one whose expected value you estimated a priori at +€0.36/coupon — to the "Convenience shoppers" segment (average spend €96/month). After the campaign, the first triumphant analysis arrives: "members who redeemed the coupon spent €158 this month, versus €118 for those who did not redeem it: +€40 thanks to the campaign!".

This is the classic selection bias trap: someone who redeems a coupon was already, to begin with, a heavier shopper. Comparing redeemers with non-redeemers does not measure the campaign; it measures who is who. The only clean comparison is against a random control group that did not receive the coupon: incrementality.

The experiment

Picking up the sizing done with the minimum relevant effect in Errors, Power and Sample Size (~7,250 per group to detect +€1), 15,000 members of the segment are selected at random: 7,500 receive the coupon and 7,500, chosen just as randomly, receive nothing. Spend over the following 30 days:

Group \(n\) Average spend
Treatment (coupon) 7,500 €98.90
Control 7,500 €96.00

With \(s \approx \text{€}21.50\) in both groups, the two-sample t-test:

\[ SE = 21.50\sqrt{\tfrac{2}{7500}} = 0.35 \qquad t = \frac{2.90}{0.35} \approx 8.3 \qquad p < 0.001 \]

Incrementality: +€2.90/member, with a 95% CI of \((2.21;; 3.59)\) euros. Everything the control group spent (€96) would have happened anyway without the campaign: those €2.90 are the only thing the campaign has caused.

From the effect to the P&L

  • Gross margin of the segment ≈ 35% → incremental margin: \(2.90 \times 0.35 \approx \text{€}1.02\)/member.
  • 14% redeemed the coupon → cost: \(5 \times 0.14 = \text{€}0.70\)/member.
  • Net profit: ≈ +€0.32 per member contacted — remarkably close to the +€0.36 predicted by the expected value calculation in Module 3. Theory and experiment shake hands.

Decision: extend the campaign to the segment's ~155,800 members (+~€50,000 net per wave), but with nuances the CI forces you to look at: at the lower bound (+€2.21) the net drops to +€0.07/member, close to break-even. The team's recommendation to Marta: roll it out, monitor actual redemption (if it rises above 14%, the margin gets eaten) and test a coupon conditional on a minimum purchase in the next wave.

Synthesis: the business analyst's method map

The pattern across the four cases is always the same cycle, with a different tool in step 3:

Business question Data Method Lesson
Does version B convert better? Randomized experiment Two-proportion test + CI + sample size 05-03, 05-04
How much stock for 95% service? Demand history Percentiles of the normal / Poisson 04-02, 04-03
Has the process changed or is it noise? Periodic samples Control chart \(\pm 3\sigma\) (CLT) 04-04
What did the campaign cause? Random control group Incrementality: two-sample t + expected value 05-03, 03-03
How much will we sell in Q4? Historical series Decomposition + seasonal indices (already seen: €7.1 million, MAPE 1.8%) 07-01
What explains a store's sales? Cross-section of 42 stores Multiple regression (already seen: R² = 0.912) 07-02

Look at the last column: nothing in this lesson was new. What is new is the judgment to match question and method, and the discipline of always closing with a decision in euros.

Common Mistakes and Tips

  • Stopping the A/B test "once it comes out significant". Checking the p-value every day and stopping when it crosses 0.05 multiplies false victories. Size and duration are fixed beforehand; the analysis happens at the end.
  • Confusing significant with profitable. With 33,000 sessions per group, irrelevant differences come out significant. The decision is made with the CI translated into euros, not with the p-value's asterisk.
  • Sizing stock with the mean. Stock = mean demand guarantees a stockout half the time. The service level is a percentile; the safety stock is the price of variance.
  • Using the normal for small counts. With \(\lambda < 5\) sales per period, use Poisson; the normal spreads probability over negative values and misleads in the tails.
  • Reacting to points inside the control limits. Adjusting the process at every common fluctuation (tampering) makes variability worse. Only signals justify intervening.
  • Measuring campaigns without a control group. Before/after (contaminated by seasonality) and redeemers/non-redeemers (selection bias) systematically inflate the effect. Without a random control there is no incrementality.

Exercises

Exercise 1

Marketing wants to test a new email subject line to raise the open rate from 22% to 24%. Calculate the size per group with \(\alpha = 0.05\) two-sided and 80% power (\(z_{\alpha/2}=1.96\), \(z_\beta=0.84\)).

Exercise 2

An item sells \(\lambda = 1.2\) units/week in Cuenca (Poisson, weekly replenishment). What stock covers at least 95% of weeks? (Hint: \(e^{-1.2} = 0.3012\).)

Exercise 3

The daily average transaction value at Sevilla-Nervión is monitored with samples of \(n = 100\) receipts; the stable process has \(\mu = \text{€}32.40\) and \(\sigma = \text{€}21.50\). Calculate the \(\pm 3\sigma\) control limits for the sample mean and say whether each of these is a signal: (a) \(\bar{x} = \text{€}36.10\); (b) \(\bar{x} = \text{€}39.20\).

Solutions

Exercise 1.

\[ n = \frac{7.84,(0.22 \cdot 0.78 + 0.24 \cdot 0.76)}{0.02^2} = \frac{7.84 \times 0.3540}{0.0004} \approx 6,940 \text{ per group} \]

About 6,950 emails per arm (14,000 in total). Common mistake: using 0.02 without squaring it, or mixing percentage points (2) with proportions (0.02) — the effect in the denominator goes in proportions.

Exercise 2.

\(P(0)=0.3012\); \(P(1)=0.3614\) (cum. 0.6626); \(P(2)=0.2169\) (0.8795); \(P(3)=0.0867\) (0.9662). With 3 units, 96.6% of weeks are covered; with 2, only 88.0%. Tip: in Poisson the stock is always found by accumulating until you cross the target level — there is no closed-form percentile, and rounding "down" breaks the promised service.

Exercise 3.

\(\sigma_{\bar{x}} = 21.50/\sqrt{100} = 2.15\). Limits: \(32.40 \pm 3 \times 2.15 = (25.95;; 38.85)\). (a) 36.10 is inside: common variation, do not intervene (although if several consecutive days accumulate above 32.40, the runs rules would raise the alarm). (b) 39.20 > 38.85: signal — investigate (special promotion? till error? change in mix?). Common mistake: using \(\sigma = 21.50\) instead of \(\sigma/\sqrt{n}\); the limits would come out absurdly wide and the chart would never warn you.

Conclusion

You have closed four complete business cases with tools you already had: the A/B test that approved the new checkout (a two-proportion test with a pre-set sample size and a decision in euros), safety stock as a percentile of the normal and the Poisson, the control chart that caught the Zaragoza hub problem, and the coupon's incrementality that unmasked the naive redeemer analysis — confirming along the way, with a measured +€0.32 against a predicted +€0.36, the expected value calculation from Module 3. The common thread: business statistics does not end at the p-value; it ends at a decision with a cost and a risk.

But not every assignment the team receives looks inward at NovaMarket. The next one comes from outside: a trade association wants to study consumer habits across the whole country, and that means large-scale surveys, social groups, opinion scales and the particular traps of measuring a society — Statistics in Social Sciences.

© Copyright 2026. All rights reserved