Non-parametric test for paired data — compares two related measurements when difference scores are ordinal or non-normal.
Compare GroupsBivariatealso known as: Wilcoxon matched-pairs test
✓ When to use
The same participants measured twice (pre/post) with an ordinal outcome or non-normal difference scores.
One sample compared against a benchmark when normality fails (non-parametric one-sample test).
Small paired samples where the paired t-test's normality assumption is doubtful.
✗ When NOT to use
Independent groups — use Mann–Whitney U.
Three or more related occasions — use the Friedman test.
Difference scores are approximately normal — the paired t-test is more powerful.
The outcome is purely nominal (improved/not improved) — use McNemar's test.
Data requirements
Dependent / outcome variable
One outcome measured twice per participant, at least ordinal; differences must be rankable.
Independent / grouping variable
Within-person factor with two levels (time/condition).
Design
Within-subjects or matched pairs.
Sample size guidance
Works from n ≈ 6 pairs upward (exact test); normal approximation for larger n. Zero differences are dropped, reducing effective n.
Assumptions
Pairs are independent of other pairs.
Differences can be meaningfully ranked (at least ordinal).
For a symmetric-location interpretation, the distribution of differences is roughly symmetric around the median; otherwise interpret as a tendency toward positive or negative change.
Hypotheses
H₀ — The distribution of difference scores is symmetric around zero (no systematic change).
H₁ — Differences tend to be positive (or negative) — a systematic shift between occasions.
The concept
Compute each pair's difference, discard zeros, rank the absolute differences, then re-attach the signs. The statistic W (or T) is the smaller of the positive-rank sum and negative-rank sum. If there is no systematic change, positive and negative ranks should balance; a heavily one-sided rank sum signals a real shift.
Unlike the sign test, Wilcoxon uses the magnitude of changes (via ranks), making it more powerful. Effect size: r = |Z|/√N or the matched rank-biserial correlation.
Worked example
Employees rate workload manageability (single 5-point item) before and after a workflow redesign (n = 30 pairs, 4 ties dropped).
Result: W = 87.5, Z = −2.61, p = .009, r = .36 — manageability ratings improved significantly (Mdn 3 → 4).
from scipy import stats
import pingouin as pg
print(stats.wilcoxon(df["post"], df["pre"]))
print(pg.wilcoxon(df["post"], df["pre"])) # includes effect sizes
Analyze → Nonparametric Tests → Related Samples (or Legacy Dialogs → 2 Related Samples).
Select the pair (Pre, Post); choose Wilcoxon.
Report T (sum of ranks), Z, p, medians at both occasions, and r = |Z|/√N.
Compute differences D = Post − Pre; remove rows where D = 0.
Rank |D| with =RANK.AVG(ABS(D), all |D| range, 1); sum ranks separately for positive and negative D.
W = the smaller rank sum. Normal approximation (n > 20): Z = (W − n(n+1)/4)/SQRT(n(n+1)(2n+1)/24); p = 2*(1−NORM.S.DIST(ABS(Z),TRUE)).
For small n, use a Wilcoxon critical-value table (or run it in R/Python/SPSS).
Interpreting the output
W (T) and Z with the p-value; state how many zero-differences were excluded.
Report medians (and IQRs) for both occasions.
Effect size r; direction from which rank sum dominated.
Frame conclusions as a systematic shift, not a change in means.
APA-style reporting
A Wilcoxon signed-rank test showed a significant improvement in workload manageability from before (Mdn = 3) to after (Mdn = 4) the redesign, T = 87.5, Z = −2.61, p = .009, r = .36.
Common mistakes
Using it on independent groups.
Forgetting that tied (zero) differences are dropped — report the effective n.
Interpreting the result as a mean change.
Choosing Wilcoxon after seeing the t-test result — decide on the analysis before running tests.