Tests whether three or more independent group means differ on a continuous outcome — e.g., does job satisfaction differ across four departments?
Compare GroupsBivariatealso known as: Analysis of variance, Single-factor ANOVA
✓ When to use
One categorical factor with three or more independent levels (departments, job grades, treatment arms).
A continuous outcome, approximately normal within groups with similar variances.
You want one overall (omnibus) answer first, then planned contrasts or post-hoc tests to locate the differences.
✗ When NOT to use
Only two groups — a t-test answers the question directly.
Repeated measurements of the same people — use Repeated-Measures ANOVA.
Clearly unequal variances — use Welch ANOVA.
Ordinal or badly skewed outcomes in small samples — use Kruskal–Wallis.
You need to control a continuous covariate — use ANCOVA.
Data requirements
Dependent / outcome variable
One continuous variable (interval/ratio).
Independent / grouping variable
One categorical variable with 3+ independent levels.
Design
Between-subjects; each participant in exactly one group.
Sample size guidance
Aim for ≥ 20–30 per group and roughly balanced sizes. G*Power: f = 0.25 (medium), 4 groups, power .80 → N ≈ 180.
Assumptions
Independence of observations.
Normality of the outcome within each group (robust with larger, balanced groups).
Homogeneity of variance across groups (Levene's test); ANOVA is fairly robust when group sizes are equal, fragile when they are not.
The factor levels are mutually exclusive and exhaustive for the comparison at hand.
Hypotheses
H₀ — All population means are equal (μ₁ = μ₂ = … = μk).
H₁ — At least one population mean differs (not all means are equal — ANOVA does not say which).
The concept
ANOVA splits total variability into between-group variance (differences among group means) and within-group variance (spread inside groups). The F statistic is their ratio: F = MSbetween/MSwithin. If all groups share one population mean, F hovers around 1; group differences push F upward.
A significant F only says 'somewhere the means differ'. Follow up with post-hoc tests (Tukey HSD when variances are equal; Games–Howell when not) or, better, a small set of planned contrasts. Effect size: η² or partial η² (.01 small, .06 medium, .14 large), or ω² for a less biased estimate.
Worked example
An HR study compares job satisfaction (1–5 composite) across four departments: Sales (n = 41, M = 3.35), Operations (n = 44, M = 3.51), IT (n = 38, M = 3.88), Finance (n = 40, M = 3.60).
Result: F(3, 159) = 4.72, p = .003, η² = .082. Tukey HSD shows IT > Sales (p = .002); other pairs do not differ significantly.
How to run it
model <- aov(satisfaction ~ department, data = df)
summary(model)
car::leveneTest(satisfaction ~ department, data = df) # variance check
TukeyHSD(model) # post-hoc, equal variances
library(effectsize)
eta_squared(model)
Outcome into 'Dependent List', factor into 'Factor'.
Options: tick Descriptives, Homogeneity of variance test, and Welch (as a safeguard).
Post Hoc: tick Tukey (equal variances) and Games–Howell (unequal).
Report F(dfbetween, dfwithin), p, effect size (Analyze → General Linear Model → Univariate prints partial η²), and the post-hoc pattern.
Arrange each group's scores in its own column.
Data → Data Analysis → 'Anova: Single Factor'; select the range grouped by columns.
Read F, P-value, and F crit. η² = SSbetween/SStotal from the output table.
Excel offers no post-hoc tests — run pairwise t-tests with a Bonferroni-adjusted α, or use R/SPSS.
Interpreting the output
F with both degrees of freedom (k − 1, N − k) and p — the omnibus verdict.
η²/ω² — proportion of variance in the outcome explained by group membership.
Post-hoc results — which specific pairs differ, with adjusted p-values.
Group means and SDs — always report the descriptive pattern; a significant F with trivial mean differences may be practically irrelevant.
APA-style reporting
A one-way ANOVA revealed a significant effect of department on job satisfaction, F(3, 159) = 4.72, p = .003, η² = .08. Tukey post-hoc comparisons indicated that IT employees (M = 3.88, SD = 0.55) reported higher satisfaction than Sales employees (M = 3.35, SD = 0.61), p = .002; no other pairwise differences were significant.
Common mistakes
Running all pairwise t-tests instead of ANOVA + corrected post-hocs — Type I error balloons.
Interpreting a significant omnibus F as 'all groups differ'.
Using Tukey when variances are unequal — switch to Games–Howell.
Ignoring unequal group sizes combined with unequal variances — the classic F becomes unreliable (use Welch).
Reporting p without effect size and group descriptives.