The t tests, ANOVA and the regression intervals share a piece of fine print: they assume the data follow (approximately) a known distribution, nearly always the normal, or that the sample is large enough for the Central Limit Theorem to come to the rescue. But NovaMarket's day-to-day is full of situations where that fine print does not hold: eight ratings from a pilot, 0-to-10 satisfaction scales that are ordinal rather than truly numeric, one outlandish delivery time dragging the mean along. Non-parametric methods are the alternative: tests that do not require assuming any specific distribution because they work with ranks (ordered positions) instead of the original values. In this lesson you will learn when to reach for them, what price you pay, and the four essential tests — sign, Wilcoxon, Mann-Whitney and Kruskal-Wallis — with hand calculations you can actually manage. It closes the module with the guide table pairing each parametric test with its alternative.

Contents

  1. When parametric assumptions fail
  2. What you gain and what you lose: the price in power
  3. The sign test: the simplest idea
  4. Wilcoxon signed-rank test: paired data
  5. Mann-Whitney U test: two independent samples
  6. Kruskal-Wallis test: the alternative to ANOVA
  7. Spearman, revisited
  8. Guide table and decision flowchart

When parametric assumptions fail

Three typical scenarios set off the alarm:

  • Non-normality with a small sample. With large \(n\), the CLT protects the sample mean even if the population is skewed; with \(n = 8\) or \(n = 12\), no theorem will save you: if the data are skewed (like the delivery times, with their right tail — median 25.5 h and 13% outside the SLA), the t test can produce p-values that do not mean what they say.
  • Ordinal data. A satisfaction score of 8 is more than one of 4, but is it "twice as much"? Rating scales order but do not guarantee equal distances between points; computing means and standard deviations on them is, strictly speaking, an abuse. Ranks use exactly the information the scale does offer: the order.
  • Extreme outliers. A single wild value inflates \(\bar{x}\) and, above all, \(s\), and can just as easily manufacture false significance as mask a real effect. Ranks are immune: the largest value gets the last rank, whether it is 48 or 4,800.

The idea common to all the tests in this lesson: replace each data point with its position in the joint ordering and run the inference on those positions. Since the ranks of \(n\) data points are always \(1, 2, \dots, n\), their behavior under the null hypothesis is known exactly, with no assumptions about the original distribution. That is why these tests usually state hypotheses about medians (or about a "tendency toward larger values") rather than about means.

What you gain and what you lose: the price in power

There is no free lunch. Moving from values to ranks discards information (that 48 h is far beyond 29 h, not merely "after it"), and discarding information costs power (Module 5: the ability to detect an effect that exists).

Parametric test Non-parametric test
Assumptions Distribution (normality) or large \(n\) Minimal (random samples, at least ordinal scale)
What it tests Means Medians / a shift in the distribution
With genuinely normal data Maximum power Slightly less powerful (Wilcoxon and Mann-Whitney deliver ≈ 95% of the t test's power)
With skewness or outliers Can fail badly Robust, often more powerful
Result it communicates "The mean rises by €X" "The typical level rises" (no direct effect figure in €)

Two practical consequences: first, if the parametric assumptions hold, use the parametric test — the non-parametric method's power loss is small, but it is a loss; second, the non-parametric result is less rich (it says "there is a difference", but gives no interval in euros for the mean), so always accompany it with robust descriptives: medians and interquartile ranges, which you already command from Module 2.

The sign test: the simplest idea

Let's start with the most elementary non-parametric test, useful for paired data (before/after on the same units). Its logic fits in one sentence: if the treatment did nothing, each unit would have a 50% chance of improving and a 50% chance of worsening — so the number of improvements would be a binomial with \(p = 0.5\), our old acquaintance from Module 4.

Lightning example: NovaMarket pilots a seasonal-fruit promotion in 10 stores, comparing Fresh food sales in the promotion week against the previous week. 9 stores go up and 1 goes down (exact ties are discarded). Under the null hypothesis, the number of increases is \(Bin(10;, 0.5)\):

\[ P(X \geq 9) = P(X=9) + P(X=10) = \frac{10 + 1}{2^{10}} = \frac{11}{1024} \approx 0.0107 \]

Two-tailed, p-value ≈ 0.021 < 0.05: the promotion moved sales. The sign test uses only the direction of each change, ignoring its size; it is the least powerful of the family, but also the one with the weakest assumptions and the easiest to explain in a committee meeting. The natural next step is to exploit the size of the changes as well — without going back to means — by using their ranks: that is Wilcoxon.

Wilcoxon signed-rank test: paired data

The Wilcoxon signed-rank test is the non-parametric alternative to the paired t test. Case: Valencia-Ruzafa has just been refurbished, and 8 customers from the mystery shopper panel rated the store before and after:

Customer Before After Difference \(d\) \(\lvert d \rvert\) Rank of \(\lvert d \rvert\) Signed rank
1 6.8 7.6 +0.8 0.8 5 +5
2 7.2 7.5 +0.3 0.3 3 +3
3 6.5 7.4 +0.9 0.9 7 +7
4 7.0 6.8 −0.2 0.2 2 −2
5 7.4 8.1 +0.7 0.7 4 +4
6 6.9 7.8 +0.9 0.9 7 +7
7 7.1 7.0 −0.1 0.1 1 −1
8 6.6 7.5 +0.9 0.9 7 +7

Step by step:

  1. Compute the differences after − before (zeros are discarded; there are none here).
  2. Order by absolute value and assign ranks: the smallest \(\lvert d \rvert\) (0.1) gets rank 1, and so on. The three 0.9s tie for positions 6, 7 and 8, so each receives the average rank \((6+7+8)/3 = 7\).
  3. Sum separately the ranks of the positive and negative differences:
    • \(W^+ = 5 + 3 + 7 + 4 + 7 + 7 = 33\)
    • \(W^- = 2 + 1 = 3\)
    • Check: \(33 + 3 = 36 = \frac{8 \times 9}{2}\) ✓ (the sum of the ranks of \(n\) data points is always \(n(n+1)/2\)).
  4. Statistic: \(W = \min(W^+, W^-) = 3\). The logic: if the refurbishment changed nothing, positives and negatives would split the ranks evenly (each sum would hover around 18); a very small \(W\) betrays the imbalance.
  5. Decision: for \(n = 8\) and two-tailed \(\alpha = 0.05\), the tabulated critical value is 3 — \(H_0\) is rejected if \(W \leq 3\). Since \(W = 3\), we reject: the improvement in ratings after the refurbishment is statistically significant.

Note the contrast with the sign test: with 6 improvements and 2 declines, the sign test would not have reached significance (\(P(X \geq 6) \approx 0.145\) one-tailed); Wilcoxon does, because besides counting directions it notices that the declines are tiny (ranks 1 and 2) while the improvements are large. That extra use of information is power recovered.

For \(n > 20\)–25, \(W\) is well approximated by a normal distribution and software gives you a p-value directly; the rank mechanism is the same.

Mann-Whitney U test: two independent samples

The Mann-Whitney U test (equivalent to the Wilcoxon rank-sum test) is the alternative to the two-sample t test for independent samples. Case: logistics compares the delivery times (hours) of the two carriers covering the city-center zone. Small samples and a distribution we know is skewed: non-parametric territory.

  • Carrier A (\(n_1 = 5\)): 21, 23, 26, 29, 34
  • Carrier B (\(n_2 = 6\)): 28, 31, 35, 38, 41, 44

Step by step:

  1. Order all 11 observations together and assign global ranks:
Value 21 23 26 28 29 31 34 35 38 41 44
Rank 1 2 3 4 5 6 7 8 9 10 11
Group A A A B A B A B B B B
  1. Sum the ranks of one group: \(R_A = 1 + 2 + 3 + 5 + 7 = 18\).
  2. Compute U:

\[ U_A = n_1 n_2 + \frac{n_1(n_1+1)}{2} - R_A = 30 + 15 - 18 = 27 \qquad U_B = n_1 n_2 - U_A = 30 - 27 = 3 \]

\(U_B = 3\) has a very concrete reading: of the \(5 \times 6 = 30\) possible pairings between one observation from A and one from B, only in 3 does B's arrive first. The statistic is \(U = \min(U_A, U_B) = 3\).

  1. Decision: the tabulated critical value for \(n_1 = 5\), \(n_2 = 6\), two-tailed \(\alpha = 0.05\) is 3 — reject if \(U \leq 3\). Since \(U = 3\), we reject: carrier A consistently delivers earlier (median 26 h against 36.5 h). With the SLA at stake, there is a statistical argument for renegotiating the route allocation.

An added advantage in this context: if tomorrow one of B's orders took 90 hours because of an incident, its rank would still be 11 and the conclusion would not budge — the same t test would have seen its \(s\) balloon and perhaps lost significance because of the outlier.

Kruskal-Wallis test: the alternative to ANOVA

What about more than two groups? The Kruskal-Wallis test extends Mann-Whitney to \(k\) groups just as the ANOVA of 06-03 extended the t test. The null hypothesis: the \(k\) populations have the same distribution (equal medians); the mechanics: global ranks, and compare each group's rank sum with what would be expected under \(H_0\):

\[ H = \frac{12}{N(N+1)} \sum_{i=1}^{k} \frac{R_i^2}{n_i} - 3(N+1) \]

where \(N\) is the total number of observations and \(R_i\) is the rank sum of group \(i\). Under \(H_0\), \(H\) approximately follows a chi-square distribution with \(k-1\) degrees of freedom — the same distribution from 06-04 in a new role.

Operational case: instead of the ANOVA's satisfaction means by format, internal audit scores compliance with store standards (an ordinal 0–10 scale) in 4 stores of each format:

Format Scores Global ranks \(R_i\)
Express 6.6 · 7.1 · 6.9 · 7.4 1 · 3 · 2 · 4 10
Standard 7.8 · 8.2 · 7.9 · 8.4 5 · 7 · 6 · 8 26
Superstore 8.8 · 9.1 · 8.7 · 9.3 10 · 11 · 9 · 12 42

With \(N = 12\):

\[ H = \frac{12}{12 \times 13}\left(\frac{10^2}{4} + \frac{26^2}{4} + \frac{42^2}{4}\right) - 3 \times 13 = \frac{12}{156}(25 + 169 + 441) - 39 = 0.0769 \times 635 - 39 = 9.85 \]

Comparing with \(\chi^2_{0.05;,2} = 5.99\): since \(9.85 > 5.99\), \(H_0\) is rejected (p-value ≈ 0.007): compliance with standards differs across formats, with the ranks climbing cleanly from Express to Superstore — the non-parametric echo of the 7.0/8.0/9.0 gradient the ANOVA found in satisfaction. Like ANOVA, Kruskal-Wallis is an omnibus test: it says some group differs, not which one; to locate it, run pairwise Mann-Whitney comparisons with a multiple-comparison correction (Bonferroni, as in 06-03).

Spearman, revisited

You already know the Spearman correlation from 06-01: it is exactly the Pearson coefficient computed on ranks, which is why it belongs to this family in its own right. All that remains is to place it: it is the non-parametric alternative to Pearson when the relationship is monotonic but not linear, when there are outliers, or when one of the variables is ordinal. The \(r_s = -0.95\) between delivery time and satisfaction we computed back then was a full-fledged non-parametric analysis — you were already using ranks before you knew that had a name.

Guide table and decision flowchart

The table that sums up Module 5, Module 6 and this lesson — the definitive cheat sheet for choosing a test:

Business question Parametric test Non-parametric alternative When to use the non-parametric one
Does a median/mean differ from a reference value? One-sample t (05-03) Wilcoxon signed-rank (against the reference value); sign test Small \(n\) with skewness, ordinal data, outliers
Did things change after an intervention (same units)? Paired t (05-03) Sign test; Wilcoxon signed-rank Same as above; sign test if only the direction of change is trustworthy
Do two independent groups differ? Two-sample t (05-03) Mann-Whitney U Same as above
Do \(k\) independent groups differ? One-way ANOVA (06-03) Kruskal-Wallis Non-normality, ordinal data, very unequal variances with small groups
Are two numeric variables associated? Pearson correlation (06-01) Spearman correlation Monotonic non-linear relationship, outliers, ordinal data
Are two categorical variables associated? — Chi-square test of independence (06-04) It is already a test with no distributional assumptions about the data: it needs no "alternative"

And the decision flowchart:

flowchart TD
    A[What kind of data are you comparing?] --> B{Categorical variables?}
    B -- Yes --> C[Chi-square 06-04]
    B -- No --> D{Numeric data with reasonable normality or large n, no serious outliers?}
    D -- Yes --> E{How many groups?}
    E -- "1 or paired" --> F[One-sample t / paired t]
    E -- "2 independent" --> G[Two-sample t]
    E -- "3 or more" --> H[ANOVA]
    E -- "relationship between 2 variables" --> I[Pearson / regression]
    D -- "No: small n, ordinal data or outliers" --> J{How many groups?}
    J -- "paired" --> K[Sign test / Wilcoxon]
    J -- "2 independent" --> L[Mann-Whitney U]
    J -- "3 or more" --> M[Kruskal-Wallis]
    J -- "relationship between 2 variables" --> N[Spearman]

There is also a third, modern route we will only mention: the bootstrap, which replaces the assumptions with computer-intensive resampling.

Common Mistakes and Tips

  • Using non-parametric methods "just in case" with normal data and large n. You give away power for no reason; non-parametric is a well-founded plan B, not the default plan.
  • Interpreting Mann-Whitney or Kruskal-Wallis as comparisons of means. They test medians/shift; report medians and IQRs alongside the p-value, not means.
  • Forgetting how to handle ties. Repeated values receive the average rank of their positions (the three 0.9s in the example → rank 7 each); assigning ranks arbitrarily among tied values changes the result.
  • Discarding zeros and ties incorrectly in the sign/Wilcoxon tests. Differences that are exactly zero are removed before counting \(n\); forgetting this shifts the critical values.
  • Getting the decision rule backwards. In Wilcoxon and Mann-Whitney you reject when the statistic is less than or equal to the tabulated critical value (the opposite of t, F or chi-square, where you reject from above). It is the classic slip with these tables.
  • Skipping the multiple-comparison correction after Kruskal-Wallis. Running three Mann-Whitney tests at \(\alpha = 0.05\) without correcting inflates the overall Type I error, exactly as in ANOVA's post-hoc stage.
  • Giving up on descriptives. "p < 0.05 with Mann-Whitney" does not tell Marta how much better carrier A is; each group's median (26 h vs 36.5 h) does.

Exercises

Exercise 1

A new layout of the household-products section is piloted in 9 stores, comparing the category's sales before and after. 8 stores go up and 1 goes down. Apply the sign test with two-tailed \(\alpha = 0.05\). (Hint: \(P(X \geq 8)\) with \(Bin(9;,0.5)\) = \( (9 + 1 + \binom{9}{7} \cdot 0)/2^9 \)... compute \(P(X=8) + P(X=9)\).)

Exercise 2

Six employees at Bilbao-Casco rated the workplace climate (0–10) before and after a shift change: differences after − before: +1.5; +0.5; −0.3; +1.1; +0.8; −0.6. Run the Wilcoxon signed-rank test (two-tailed critical value for \(n = 6\), \(\alpha = 0.05\): reject if \(W \leq 0\)).

Exercise 3

Time (minutes) to resolve a checkout incident at two stores: Madrid-Chamberí (\(n_1 = 4\)): 5, 7, 8, 12; Zaragoza-Centro (\(n_2 = 5\)): 9, 13, 15, 18, 40. Compute the Mann-Whitney U (two-tailed critical value for \(n_1 = 4, n_2 = 5\), \(\alpha = 0.05\): reject if \(U \leq 1\)) and comment on the role of the value 40.

Solutions

Exercise 1.

\(P(X = 8) = \binom{9}{8}/2^9 = 9/512\); \(P(X = 9) = 1/512\). One tail: \(10/512 \approx 0.0195\); two-tailed: \(\approx 0.039 < 0.05\). We reject \(H_0\): the new layout moves household-products sales. Common mistake: forgetting to double for the two-tailed test and reporting 0.0195 — if the question was "does it change?" (not "does it go up?"), the p-value is twice that.

Exercise 2.

Absolute values in order: 0.3 (rank 1); 0.5 (2); 0.6 (3); 0.8 (4); 1.1 (5); 1.5 (6). Sums: \(W^- = 1 + 3 = 4\) (the differences −0.3 and −0.6); \(W^+ = 2 + 4 + 5 + 6 = 17\); check \(4 + 17 = 21 = 6 \times 7/2\) ✓. \(W = \min = 4 > 0\): we do not reject \(H_0\). With \(n = 6\), only a perfect outcome (all differences with the same sign) would reach significance: the sample is too small to conclude anything — which connects with the lesson on power: absence of evidence is not evidence of absence. Common mistake: assigning the ranks to the signed differences instead of their absolute values.

Exercise 3.

Joint ordering: 5(1), 7(2), 8(3), 9(4), 12(5), 13(6), 15(7), 18(8), 40(9). Chamberí's ranks: 1, 2, 3, 5 → \(R_1 = 11\). \(U_1 = 4 \times 5 + \frac{4 \times 5}{2} - 11 = 20 + 10 - 11 = 19\); \(U_2 = 20 - 19 = 1\). \(U = 1 \leq 1\): we reject \(H_0\): Chamberí resolves incidents faster (medians 7.5 vs 15 min). The value 40 (an incident that had to be escalated?) simply receives rank 9: the conclusion would be identical if it were 20 or 400 — with a t test, that single data point would have blown up Zaragoza's standard deviation and probably wrecked the significance. Common mistake: computing U with the wrong group's rank sum without adjusting the formula; always check that \(U_1 + U_2 = n_1 n_2\).

Conclusion

With this lesson, the course's technical arsenal is complete: you know when parametric assumptions fail (small n without normality, ordinal data, outliers), what giving them up costs (some power and the richness of an effect in units), and you command the full ladder — the sign test for direction, Wilcoxon for pairs, Mann-Whitney for two groups, Kruskal-Wallis for several, Spearman for association — along with the guide table and the flowchart that pair each parametric test with its robust double. The Valencia-Ruzafa refurbishment, the slow carrier and the standards gradient across formats have all fallen to tests that ask no permission from the bell curve.

And with Module 7, the tool-building phase of the course ends too: descriptive statistics, probability, distributions, inference, regression, ANOVA, chi-square, time series, multivariate analysis and non-parametrics. What remains is the closest thing to an analyst's real job: facing a complete case and deciding, with nobody whispering the method, which tool answers which question. That is Module 8: deploying the whole arsenal, sector by sector, starting with NovaMarket's home turf — Statistics in Business.

© Copyright 2026. All rights reserved