Statistical Methods Atlas
Atlas › Compare Groups › Friedman Test

Friedman Test

Non-parametric test for three or more related measurements — the distribution-free counterpart of repeated-measures ANOVA.

Compare GroupsBivariate also known as: Friedman two-way ANOVA by ranks

✓ When to use

  • The same participants rated/measured under 3+ conditions or occasions, with ordinal or non-normal data.
  • Rankings by judges (each judge ranks all objects).
  • Small within-subject designs where RM-ANOVA's assumptions are untenable.

✗ When NOT to use

  • Independent groups — Kruskal–Wallis.
  • Two related occasions — Wilcoxon Signed-Rank.
  • Approximately normal data with sphericity manageable — RM-ANOVA is more powerful.
  • Missing occasions for some participants — Friedman needs complete blocks (or use mixed models on ranks).

Data requirements

Dependent / outcome variableOne outcome, at least ordinal, measured k ≥ 3 times per participant (complete rows).
Independent / grouping variableWithin-subject factor with k levels.
DesignRandomized block / within-subjects; each participant is a block.
Sample size guidancen ≥ 10 blocks for the chi-square approximation; exact tests for tiny n.

Assumptions

Hypotheses

H₀ — All k conditions have identical distributions — each condition is equally likely to receive any rank within a person.
H₁ — At least one condition tends to receive higher (or lower) ranks.

The concept

Within each participant, the k condition scores are converted to ranks 1…k. If conditions do not differ, each condition's average rank across participants should hover near (k+1)/2; Friedman's χ²F measures the spread of the condition rank-sums around that expectation.

Because ranking happens within person, stable individual differences are automatically removed — the same logic as RM-ANOVA, executed on ranks. Follow a significant result with pairwise Wilcoxon signed-rank tests (Holm/Bonferroni-adjusted) or the Nemenyi procedure. Effect size: Kendall's W (0 = no agreement, 1 = perfect consistency of rankings).

Worked example

25 managers rate the usefulness (5-point item) of three appraisal formats: self, peer, and 360-degree — every manager rates all three.

Result: χ²F(2) = 11.76, p = .003, Kendall's W = .24. Post-hoc Wilcoxon (Holm): 360-degree outranks self-appraisal (p = .004).

How to run it

# wide: self, peer, deg360
friedman.test(as.matrix(df[, c("self", "peer", "deg360")]))

library(rstatix)   # long format: id, format, rating
friedman_effsize(dfl, rating ~ format | id)      # Kendall's W
pairwise_wilcox_test(dfl, rating ~ format, paired = TRUE,
                     p.adjust.method = "holm")

Interpreting the output

APA-style reporting

A Friedman test indicated that perceived usefulness differed across appraisal formats, χ²F(2) = 11.76, p = .003, Kendall's W = .24. Post-hoc Wilcoxon signed-rank tests with Holm correction showed 360-degree feedback (Mdn = 4) was rated more useful than self-appraisal (Mdn = 3), p = .004.

Common mistakes

Related methods

Repeated-Measures ANOVAParametric alternativeWilcoxon Signed-Rank TestTwo related occasionsKruskal–Wallis TestIndependent-groups analogue
← Kruskal–Wallis TestPearson Correlation →