Tests the effects of two categorical factors — and, crucially, their interaction — on one continuous outcome, e.g., do gender and job level jointly shape pay satisfaction?
Compare GroupsMultivariatealso known as: Factorial ANOVA, Two-factor ANOVA
✓ When to use
Two categorical factors (e.g., gender × job level) and one continuous outcome.
You care about the interaction: does the effect of one factor depend on the level of the other?
Experimental 2×2 (or larger) designs, or survey data with two natural grouping variables.
✗ When NOT to use
One factor only — one-way ANOVA.
A factor is within-subjects — use mixed or repeated-measures ANOVA.
The 'factor' is actually continuous — don't bin it; use regression with an interaction term.
Severely unbalanced cells with unequal variances — consider robust alternatives or regression with heteroscedasticity-consistent errors.
Data requirements
Dependent / outcome variable
One continuous variable.
Independent / grouping variable
Two categorical factors (each 2+ levels); crossed so every combination (cell) has cases.
Design
Between-subjects factorial; aim for balanced cell sizes.
Sample size guidance
≥ 20 per cell is a practical floor; a 2×3 design thus wants N ≥ 120. Power depends on the smallest effect of interest (usually the interaction).
Assumptions
Independence of observations.
Normality of residuals (check on the model residuals, not the raw DV).
Homogeneity of variance across all cells (Levene's test).
With unbalanced data, use Type III sums of squares and interpret main effects cautiously in the presence of interactions.
Hypotheses
H₀ — Three separate nulls: no main effect of factor A; no main effect of factor B; no A×B interaction.
H₁ — For each: the corresponding effect exists in the population.
The concept
Factorial ANOVA partitions variance into factor A, factor B, their interaction, and error. The interaction term is usually the scientific payoff: a significant A×B means the effect of A differs across levels of B — e.g., a training program (A) improves performance only for novices (B).
When the interaction is significant, interpret simple effects (the effect of A within each level of B) rather than the main effects, which can be misleading averages. Effect sizes: partial η² per effect. Always plot the cell means — an interaction plot communicates the pattern faster than any table.
Worked example
A 2 (work mode: remote/office) × 3 (job level: junior/mid/senior) design on pay satisfaction (N = 240, balanced).
Results: work mode F(1, 234) = 3.11, p = .079; job level F(2, 234) = 8.45, p < .001, ηp² = .067; interaction F(2, 234) = 4.02, p = .019, ηp² = .033 — the remote advantage appears only among juniors (simple effect p = .004).
How to run it
model <- aov(pay_sat ~ mode * level, data = df) # * includes interaction
car::Anova(model, type = 3) # Type III SS
library(effectsize); eta_squared(model, partial = TRUE)
library(emmeans)
emmeans(model, ~ mode | level) |> pairs() # simple effects
interaction.plot(df$level, df$mode, df$pay_sat)
import statsmodels.api as sm
from statsmodels.formula.api import ols
import pingouin as pg
model = ols("pay_sat ~ C(mode) * C(level)", data=df).fit()
print(sm.stats.anova_lm(model, typ=3))
print(pg.anova(df, dv="pay_sat", between=["mode", "level"])) # with np2
Analyze → General Linear Model → Univariate.
DV into 'Dependent Variable'; both factors into 'Fixed Factor(s)'.
Options: tick Descriptive statistics, Estimates of effect size, Homogeneity tests.
Plots: put one factor on the horizontal axis, the other as separate lines → Add.
If the interaction is significant: EM Means → select the interaction → 'Compare simple main effects' (via /EMMEANS syntax COMPARE).
Balanced designs only: Data → Data Analysis → 'Anova: Two-Factor With Replication'.
Lay out data as a grid: factor A levels as column blocks, factor B levels as row blocks with equal replicates per cell.
Read the three F tests (Sample = rows factor, Columns = columns factor, Interaction).
Unbalanced designs cannot be handled correctly in Excel — use R/Python/SPSS.
Interpreting the output
Check the interaction first; if significant, interpret simple effects and the interaction plot, not raw main effects.
Report F, df, p, and partial η² for each of the three effects.
Cell means and SDs (or an interaction plot) are essential for readers.
Non-significant interaction: interpret each main effect as in one-way ANOVA.
APA-style reporting
A 2 × 3 between-subjects ANOVA on pay satisfaction revealed a significant main effect of job level, F(2, 234) = 8.45, p < .001, ηp² = .07, qualified by a significant work mode × job level interaction, F(2, 234) = 4.02, p = .019, ηp² = .03. Simple-effects analysis showed remote juniors reported higher pay satisfaction than office juniors (p = .004), with no mode difference at mid or senior levels.
Common mistakes
Interpreting main effects while ignoring a significant interaction.
Median-splitting a continuous variable to create the second factor.
Running separate one-way ANOVAs instead of the factorial model — you lose the interaction test.
Using Type I sums of squares on unbalanced data without realizing order matters.
Empty or tiny cells (n < 5) making estimates unstable.