Statistical Methods Atlas
Atlas › Compare Groups › Mann–Whitney U Test

Mann–Whitney U Test

Non-parametric comparison of two independent groups using ranks — the go-to alternative to the independent t-test for ordinal or non-normal data.

Compare GroupsBivariate also known as: Wilcoxon rank-sum test, Wilcoxon–Mann–Whitney test

✓ When to use

  • Two independent groups with an ordinal outcome (e.g., a single Likert item) or a continuous outcome that is clearly non-normal in a small sample.
  • Outliers or skew would distort a mean-based comparison.
  • You are comfortable framing the conclusion in terms of rank/distributional differences rather than means.

✗ When NOT to use

  • Paired data — use the Wilcoxon Signed-Rank test.
  • Three or more groups — use Kruskal–Wallis.
  • Large samples with roughly normal data — the t-test is more powerful and easier to interpret.
  • You specifically need to compare means (e.g., for a cost calculation) — this test compares distributions, and only compares medians if the two distributions have the same shape.

Data requirements

Dependent / outcome variableOne ordinal or continuous variable.
Independent / grouping variableOne categorical variable with two independent levels.
DesignBetween-subjects; independent observations.
Sample size guidanceWorks from very small samples upward; with many ties (common in Likert data) software applies a tie correction.

Assumptions

Hypotheses

H₀ — The two distributions are identical — P(X > Y) = 0.5; a randomly chosen member of either group is equally likely to score higher.
H₁ — One group tends to produce higher values than the other.

The concept

All observations are pooled and ranked; U counts, across every possible cross-group pair, how often a member of group 1 outranks a member of group 2. If the groups do not differ, U falls near n₁n₂/2; extreme values in either direction indicate that one group systematically outranks the other.

A convenient effect size is r = |Z|/√N (≈ .10 small, .30 medium, .50 large), or report the probability of superiority U/(n₁n₂), which is directly interpretable: the chance that a random member of one group scores higher than a random member of the other.

Worked example

A researcher compares intention-to-recommend (single 5-point item) between customers of two service branches (n₁ = 28, n₂ = 31). The item is ordinal, so a t-test is inappropriate.

Result: U = 262.5, Z = −2.34, p = .019, r = .30 — Branch A customers give systematically higher ratings (median 4 vs 3).

How to run it

wilcox.test(rating ~ branch, data = df)   # Mann–Whitney for 2 independent groups

library(effectsize)
rank_biserial(rating ~ branch, data = df)  # rank-biserial correlation effect size

tapply(df$rating, df$branch, median)       # medians for reporting

Interpreting the output

APA-style reporting

A Mann–Whitney U test indicated that recommendation intention was significantly higher at Branch A (Mdn = 4) than Branch B (Mdn = 3), U = 262.5, Z = −2.34, p = .019, r = .30.

Common mistakes

Related methods

Independent-Samples t-TestParametric alternativeWilcoxon Signed-Rank TestPaired versionKruskal–Wallis TestThree or more groups
← Paired-Samples t-TestWilcoxon Signed-Rank Test →