Non-parametric test for three or more related measurements — the distribution-free counterpart of repeated-measures ANOVA.
Compare GroupsBivariatealso known as: Friedman two-way ANOVA by ranks
✓ When to use
The same participants rated/measured under 3+ conditions or occasions, with ordinal or non-normal data.
Rankings by judges (each judge ranks all objects).
Small within-subject designs where RM-ANOVA's assumptions are untenable.
✗ When NOT to use
Independent groups — Kruskal–Wallis.
Two related occasions — Wilcoxon Signed-Rank.
Approximately normal data with sphericity manageable — RM-ANOVA is more powerful.
Missing occasions for some participants — Friedman needs complete blocks (or use mixed models on ranks).
Data requirements
Dependent / outcome variable
One outcome, at least ordinal, measured k ≥ 3 times per participant (complete rows).
Independent / grouping variable
Within-subject factor with k levels.
Design
Randomized block / within-subjects; each participant is a block.
Sample size guidance
n ≥ 10 blocks for the chi-square approximation; exact tests for tiny n.
Assumptions
Blocks (participants) are independent of each other.
The outcome is at least ordinal, rankable within each participant.
No missing cells within a participant's row.
Hypotheses
H₀ — All k conditions have identical distributions — each condition is equally likely to receive any rank within a person.
H₁ — At least one condition tends to receive higher (or lower) ranks.
The concept
Within each participant, the k condition scores are converted to ranks 1…k. If conditions do not differ, each condition's average rank across participants should hover near (k+1)/2; Friedman's χ²F measures the spread of the condition rank-sums around that expectation.
Because ranking happens within person, stable individual differences are automatically removed — the same logic as RM-ANOVA, executed on ranks. Follow a significant result with pairwise Wilcoxon signed-rank tests (Holm/Bonferroni-adjusted) or the Nemenyi procedure. Effect size: Kendall's W (0 = no agreement, 1 = perfect consistency of rankings).
Worked example
25 managers rate the usefulness (5-point item) of three appraisal formats: self, peer, and 360-degree — every manager rates all three.
Result: χ²F(2) = 11.76, p = .003, Kendall's W = .24. Post-hoc Wilcoxon (Holm): 360-degree outranks self-appraisal (p = .004).
How to run it
# wide: self, peer, deg360
friedman.test(as.matrix(df[, c("self", "peer", "deg360")]))
library(rstatix) # long format: id, format, rating
friedman_effsize(dfl, rating ~ format | id) # Kendall's W
pairwise_wilcox_test(dfl, rating ~ format, paired = TRUE,
p.adjust.method = "holm")
from scipy import stats
import pingouin as pg
print(stats.friedmanchisquare(df["self"], df["peer"], df["deg360"]))
# long format follow-ups:
print(pg.pairwise_tests(dfl, dv="rating", within="format", subject="id",
parametric=False, padjust="holm"))
Analyze → Nonparametric Tests → Related Samples (or Legacy Dialogs → K Related Samples, tick Friedman).
Select the k variables (one column per condition).
The modern dialog offers adjusted pairwise comparisons after a significant omnibus result.
Report χ², df, p, condition medians, and Kendall's W (Legacy: tick Kendall's W in the same dialog).
For each participant row, rank the k condition scores with =RANK.AVG across that row (handle ties by average rank).
Sum ranks per condition (Rj). χ²F = 12/(n·k(k+1)) * ΣRj² − 3n(k+1).
p-value: =CHISQ.DIST.RT(χ²F, k−1).
Kendall's W = χ²F / (n(k−1)).
Interpreting the output
χ²F with df = k − 1 and p.
Condition medians and mean ranks describe the ordering.
Kendall's W as effect size — how consistently participants rank the conditions.
Adjusted pairwise tests locate the specific differences.
APA-style reporting
A Friedman test indicated that perceived usefulness differed across appraisal formats, χ²F(2) = 11.76, p = .003, Kendall's W = .24. Post-hoc Wilcoxon signed-rank tests with Holm correction showed 360-degree feedback (Mdn = 4) was rated more useful than self-appraisal (Mdn = 3), p = .004.
Common mistakes
Using Friedman on independent groups.
Rows with missing conditions silently dropped — check your effective n.
Uncorrected pairwise follow-ups.
Interpreting results in terms of means rather than rank tendencies.