Compares three or more group means without assuming equal variances — the robust cousin of one-way ANOVA.
Compare GroupsBivariatealso known as: Welch's F test
✓ When to use
Three or more independent groups, continuous outcome, but Levene's test is significant or SDs and group sizes differ visibly.
As a routine safeguard: many analysts now report Welch ANOVA by default for between-group comparisons.
✗ When NOT to use
Repeated measures — use Repeated-Measures ANOVA.
Severely non-normal outcomes with small samples — Kruskal–Wallis.
You need factorial designs (two factors) — Welch is defined for one factor; for factorial designs with heteroscedasticity consider robust methods (e.g., WRS2 in R).
Data requirements
Dependent / outcome variable
One continuous variable.
Independent / grouping variable
One categorical variable with 3+ independent levels.
Design
Between-subjects.
Sample size guidance
Handles unequal ns gracefully; keep each group ≥ 15 if possible.
Assumptions
Independence of observations.
Approximate normality within groups (robust for larger n).
Equal variances NOT assumed.
Hypotheses
H₀ — All population means are equal.
H₁ — At least one mean differs.
The concept
Welch's F weights each group by the precision of its own mean (n/s²) instead of pooling a common error variance, and adjusts the denominator degrees of freedom downward (usually a decimal). This keeps the false-positive rate near α even when variances differ several-fold — exactly the situation where classic ANOVA with unequal group sizes fails.
Follow a significant Welch F with Games–Howell post-hoc comparisons, which likewise do not assume equal variances. Effect size: ω² or η² computed from the Welch framework (report which).
Worked example
Comparing weekly overtime hours across three plants with very different spreads: Plant A (n = 25, SD = 2.1), Plant B (n = 60, SD = 5.3), Plant C (n = 33, SD = 3.8).
Result: Welch's F(2, 51.6) = 6.90, p = .002; Games–Howell shows Plant B exceeds Plant A (p = .001).
import pingouin as pg
print(pg.welch_anova(df, dv="overtime", between="plant"))
print(pg.pairwise_gameshowell(df, dv="overtime", between="plant"))
Analyze → Compare Means → One-Way ANOVA.
Options → tick 'Welch' under Robust Tests of Equality of Means.
Post Hoc → tick 'Games-Howell'.
Read the 'Robust Tests' table: report Welch's F, its two df (the second is decimal), and p.
No built-in Welch ANOVA. Practical route: run 'Anova: Single Factor' only if variances are similar; otherwise compute Welch's F by hand (weights wi = ni/si²) — laborious — or use R/Python/SPSS.
If you must stay in Excel, run pairwise Welch t-tests (=T.TEST(r1, r2, 2, 3)) with Bonferroni-adjusted α as an approximation.
Interpreting the output
Welch's F with numerator df (k − 1) and decimal denominator df.
p < α: at least one group differs; locate pairs with Games–Howell.
Report group means, SDs and ns — the unequal spreads are part of the story.
Effect size ω² (or ε²) alongside p.
APA-style reporting
Because variances were unequal (Levene's p = .004), Welch's ANOVA was conducted; it indicated significant differences in overtime across plants, Welch's F(2, 51.6) = 6.90, p = .002. Games–Howell comparisons showed Plant B (M = 9.4, SD = 5.3) exceeded Plant A (M = 5.8, SD = 2.1), p = .001.
Common mistakes
Following Welch's F with Tukey HSD — Tukey assumes equal variances; use Games–Howell.
Reporting the classic F when Levene's test was significant.
Rounding the decimal denominator df.
Assuming Welch also fixes non-normality — it addresses variances only.