Non-parametric comparison of two independent groups using ranks — the go-to alternative to the independent t-test for ordinal or non-normal data.
Compare GroupsBivariatealso known as: Wilcoxon rank-sum test, Wilcoxon–Mann–Whitney test
✓ When to use
Two independent groups with an ordinal outcome (e.g., a single Likert item) or a continuous outcome that is clearly non-normal in a small sample.
Outliers or skew would distort a mean-based comparison.
You are comfortable framing the conclusion in terms of rank/distributional differences rather than means.
✗ When NOT to use
Paired data — use the Wilcoxon Signed-Rank test.
Three or more groups — use Kruskal–Wallis.
Large samples with roughly normal data — the t-test is more powerful and easier to interpret.
You specifically need to compare means (e.g., for a cost calculation) — this test compares distributions, and only compares medians if the two distributions have the same shape.
Data requirements
Dependent / outcome variable
One ordinal or continuous variable.
Independent / grouping variable
One categorical variable with two independent levels.
Design
Between-subjects; independent observations.
Sample size guidance
Works from very small samples upward; with many ties (common in Likert data) software applies a tie correction.
Assumptions
Independence of observations.
The outcome is at least ordinal.
For a 'difference in medians' interpretation, the two distributions must have similar shapes; otherwise interpret as a difference in stochastic dominance (one group tends to score higher).
Hypotheses
H₀ — The two distributions are identical — P(X > Y) = 0.5; a randomly chosen member of either group is equally likely to score higher.
H₁ — One group tends to produce higher values than the other.
The concept
All observations are pooled and ranked; U counts, across every possible cross-group pair, how often a member of group 1 outranks a member of group 2. If the groups do not differ, U falls near n₁n₂/2; extreme values in either direction indicate that one group systematically outranks the other.
A convenient effect size is r = |Z|/√N (≈ .10 small, .30 medium, .50 large), or report the probability of superiority U/(n₁n₂), which is directly interpretable: the chance that a random member of one group scores higher than a random member of the other.
Worked example
A researcher compares intention-to-recommend (single 5-point item) between customers of two service branches (n₁ = 28, n₂ = 31). The item is ordinal, so a t-test is inappropriate.
Result: U = 262.5, Z = −2.34, p = .019, r = .30 — Branch A customers give systematically higher ratings (median 4 vs 3).
How to run it
wilcox.test(rating ~ branch, data = df) # Mann–Whitney for 2 independent groups
library(effectsize)
rank_biserial(rating ~ branch, data = df) # rank-biserial correlation effect size
tapply(df$rating, df$branch, median) # medians for reporting
from scipy import stats
import pingouin as pg
a = df.loc[df.branch == "A", "rating"]
b = df.loc[df.branch == "B", "rating"]
print(stats.mannwhitneyu(a, b, alternative="two-sided"))
print(pg.mwu(a, b)) # adds rank-biserial effect size
print(a.median(), b.median())
Set the test field (Rating) and groups (Branch); choose Mann–Whitney U.
Report U, Z, exact or asymptotic p, group medians, and r = |Z|/√N (compute by hand).
Excel has no built-in Mann–Whitney. Rank all scores together with =RANK.AVG(cell, all_scores_range, 1).
Sum the ranks for group 1 (R1). Compute U1 = R1 − n1(n1+1)/2 and U2 = n1*n2 − U1; U = MIN(U1, U2).
For n1, n2 > 20 use the normal approximation: Z = (U − n1*n2/2)/SQRT(n1*n2*(n1+n2+1)/12), p = 2*(1−NORM.S.DIST(ABS(Z),TRUE)).
For small samples, compare U against a Mann–Whitney critical-value table (or use R/Python/SPSS).
Interpreting the output
U (or W in R) and the Z approximation with its p-value.
Report group medians (and IQRs), not means, as the descriptive anchor.
Effect size r or rank-biserial correlation — significance alone says nothing about magnitude.
Remember the hypothesis is about distributions: say 'tended to score higher', not 'had a higher mean'.
APA-style reporting
A Mann–Whitney U test indicated that recommendation intention was significantly higher at Branch A (Mdn = 4) than Branch B (Mdn = 3), U = 262.5, Z = −2.34, p = .019, r = .30.
Common mistakes
Describing the result as a difference in means.
Claiming a 'median difference' when the two distributions have visibly different shapes.
Ignoring the tie correction with heavily tied Likert data (use software, not hand calculation).
Using Mann–Whitney merely because n is small even when data are approximately normal — small n alone does not require a non-parametric test.