Non-parametric comparison of three or more independent groups using ranks — the distribution-free counterpart of one-way ANOVA.
Compare GroupsBivariatealso known as: Kruskal–Wallis H test, One-way ANOVA on ranks
✓ When to use
Three or more independent groups with an ordinal outcome or a continuous outcome that is skewed/outlier-ridden in modest samples.
Group sizes too small to trust normality-based ANOVA.
You can frame conclusions as 'which groups tend to score higher'.
✗ When NOT to use
Repeated measures — use the Friedman test.
Two groups — Mann–Whitney U.
Data are approximately normal with similar variances — one-way ANOVA is more powerful and yields mean-based conclusions.
You need adjusted comparisons with covariates — rank-based ANCOVA alternatives (e.g., Quade) or robust regression.
Data requirements
Dependent / outcome variable
One ordinal or continuous variable.
Independent / grouping variable
One categorical variable with 3+ independent levels.
Design
Between-subjects.
Sample size guidance
Each group ≥ 5 for the chi-square approximation of H; software corrects for ties.
Assumptions
Independence of observations.
Outcome at least ordinal.
For a 'median difference' interpretation, similar distribution shapes across groups; otherwise interpret as stochastic dominance.
Hypotheses
H₀ — All k groups come from the same distribution (no group tends to yield higher values).
H₁ — At least one group tends to yield higher (or lower) values than another.
The concept
All N observations are ranked together; H measures how far each group's average rank strays from the grand average rank, weighted by group size. Under H₀, H follows approximately a chi-square distribution with k − 1 df. Large H means the groups' rank distributions have drifted apart.
A significant H is followed by pairwise Dunn tests with Holm or Bonferroni adjustment (not repeated Mann–Whitney tests without correction). Effect size: ε² or η²H = (H − k + 1)/(N − k).
Worked example
Comparing customer-satisfaction ratings (single 5-point item) across three store formats (n = 35, 40, 38).
kruskal.test(rating ~ format, data = df)
library(rstatix)
dunn_test(df, rating ~ format, p.adjust.method = "holm")
kruskal_effsize(df, rating ~ format) # eta-squared based on H
from scipy import stats
import scikit_posthocs as sp # pip install scikit-posthocs
groups = [g["rating"].values for _, g in df.groupby("format")]
print(stats.kruskal(*groups))
print(sp.posthoc_dunn(df, val_col="rating", group_col="format", p_adjust="holm"))
In the modern dialog, SPSS offers automatic pairwise follow-ups with adjusted significance when the omnibus test is significant.
Report H (chi-square), df, p, group medians, and the adjusted pairwise pattern.
Rank all data together with =RANK.AVG(cell, full_range, 1).
For each group compute the rank sum Ri. H = 12/(N(N+1)) * Σ(Ri²/ni) − 3(N+1).
p-value: =CHISQ.DIST.RT(H, k−1).
Tie-heavy Likert data needs a tie correction — prefer R/Python/SPSS.
Interpreting the output
H with df = k − 1 and p.
Group medians/IQRs as descriptives; mean ranks help explain the ordering.
Adjusted pairwise (Dunn) results locate the differences.
Effect size ε² or η²H for magnitude.
APA-style reporting
A Kruskal–Wallis test showed a significant difference in satisfaction ratings across store formats, H(2) = 9.84, p = .007, ε² = .09. Dunn's post-hoc tests with Holm correction indicated higher ratings in flagship stores (Mdn = 4) than kiosks (Mdn = 3), p = .005.
Common mistakes
Following a significant H with uncorrected Mann–Whitney tests.
Reporting means instead of medians/mean ranks.
Claiming median differences when group distributions differ in shape.