Statistical Methods Atlas
Atlas › Compare Groups › Wilcoxon Signed-Rank Test

Wilcoxon Signed-Rank Test

Non-parametric test for paired data — compares two related measurements when difference scores are ordinal or non-normal.

Compare GroupsBivariate also known as: Wilcoxon matched-pairs test

✓ When to use

  • The same participants measured twice (pre/post) with an ordinal outcome or non-normal difference scores.
  • One sample compared against a benchmark when normality fails (non-parametric one-sample test).
  • Small paired samples where the paired t-test's normality assumption is doubtful.

✗ When NOT to use

  • Independent groups — use Mann–Whitney U.
  • Three or more related occasions — use the Friedman test.
  • Difference scores are approximately normal — the paired t-test is more powerful.
  • The outcome is purely nominal (improved/not improved) — use McNemar's test.

Data requirements

Dependent / outcome variableOne outcome measured twice per participant, at least ordinal; differences must be rankable.
Independent / grouping variableWithin-person factor with two levels (time/condition).
DesignWithin-subjects or matched pairs.
Sample size guidanceWorks from n ≈ 6 pairs upward (exact test); normal approximation for larger n. Zero differences are dropped, reducing effective n.

Assumptions

Hypotheses

H₀ — The distribution of difference scores is symmetric around zero (no systematic change).
H₁ — Differences tend to be positive (or negative) — a systematic shift between occasions.

The concept

Compute each pair's difference, discard zeros, rank the absolute differences, then re-attach the signs. The statistic W (or T) is the smaller of the positive-rank sum and negative-rank sum. If there is no systematic change, positive and negative ranks should balance; a heavily one-sided rank sum signals a real shift.

Unlike the sign test, Wilcoxon uses the magnitude of changes (via ranks), making it more powerful. Effect size: r = |Z|/√N or the matched rank-biserial correlation.

Worked example

Employees rate workload manageability (single 5-point item) before and after a workflow redesign (n = 30 pairs, 4 ties dropped).

Result: W = 87.5, Z = −2.61, p = .009, r = .36 — manageability ratings improved significantly (Mdn 3 → 4).

How to run it

wilcox.test(df$post, df$pre, paired = TRUE)

library(effectsize)
rank_biserial(df$post, df$pre, paired = TRUE)

median(df$pre); median(df$post)

Interpreting the output

APA-style reporting

A Wilcoxon signed-rank test showed a significant improvement in workload manageability from before (Mdn = 3) to after (Mdn = 4) the redesign, T = 87.5, Z = −2.61, p = .009, r = .36.

Common mistakes

Related methods

Paired-Samples t-TestParametric alternativeMann–Whitney U TestIndependent-groups versionFriedman TestThree or more occasions
← Mann–Whitney U TestOne-Way ANOVA →