Statistical Methods Atlas
Atlas › Compare Groups › Kruskal–Wallis Test

Kruskal–Wallis Test

Non-parametric comparison of three or more independent groups using ranks — the distribution-free counterpart of one-way ANOVA.

Compare GroupsBivariate also known as: Kruskal–Wallis H test, One-way ANOVA on ranks

✓ When to use

  • Three or more independent groups with an ordinal outcome or a continuous outcome that is skewed/outlier-ridden in modest samples.
  • Group sizes too small to trust normality-based ANOVA.
  • You can frame conclusions as 'which groups tend to score higher'.

✗ When NOT to use

  • Repeated measures — use the Friedman test.
  • Two groups — Mann–Whitney U.
  • Data are approximately normal with similar variances — one-way ANOVA is more powerful and yields mean-based conclusions.
  • You need adjusted comparisons with covariates — rank-based ANCOVA alternatives (e.g., Quade) or robust regression.

Data requirements

Dependent / outcome variableOne ordinal or continuous variable.
Independent / grouping variableOne categorical variable with 3+ independent levels.
DesignBetween-subjects.
Sample size guidanceEach group ≥ 5 for the chi-square approximation of H; software corrects for ties.

Assumptions

Hypotheses

H₀ — All k groups come from the same distribution (no group tends to yield higher values).
H₁ — At least one group tends to yield higher (or lower) values than another.

The concept

All N observations are ranked together; H measures how far each group's average rank strays from the grand average rank, weighted by group size. Under H₀, H follows approximately a chi-square distribution with k − 1 df. Large H means the groups' rank distributions have drifted apart.

A significant H is followed by pairwise Dunn tests with Holm or Bonferroni adjustment (not repeated Mann–Whitney tests without correction). Effect size: ε² or η²H = (H − k + 1)/(N − k).

Worked example

Comparing customer-satisfaction ratings (single 5-point item) across three store formats (n = 35, 40, 38).

Result: H(2) = 9.84, p = .007, ε² = .088. Dunn–Holm comparisons: flagship stores outrank kiosks (p = .005); other pairs n.s.

How to run it

kruskal.test(rating ~ format, data = df)

library(rstatix)
dunn_test(df, rating ~ format, p.adjust.method = "holm")
kruskal_effsize(df, rating ~ format)   # eta-squared based on H

Interpreting the output

APA-style reporting

A Kruskal–Wallis test showed a significant difference in satisfaction ratings across store formats, H(2) = 9.84, p = .007, ε² = .09. Dunn's post-hoc tests with Holm correction indicated higher ratings in flagship stores (Mdn = 4) than kiosks (Mdn = 3), p = .005.

Common mistakes

Related methods

One-Way ANOVAParametric alternativeMann–Whitney U TestTwo groupsFriedman TestRepeated-measures analogueWelch ANOVANormal data, unequal variances
← MANOVA (Multivariate ANOVA)Friedman Test →