Compares group means on an outcome while statistically controlling one or more continuous covariates — e.g., post-training performance controlling for pre-training scores.
Compare GroupsMultivariatealso known as: Covariance analysis
✓ When to use
You want group comparisons purified of a nuisance continuous variable (baseline score, age, tenure).
Randomized experiments with a baseline measure: ANCOVA on post-scores controlling baseline is more powerful than gain-score comparisons.
Reducing error variance to sharpen the group effect.
✗ When NOT to use
In non-randomized designs where groups differ substantially on the covariate — 'controlling' cannot fix selection (Lord's paradox); interpret with great caution.
The covariate is affected by the treatment (a mediator) — controlling it removes part of the effect you want.
The covariate–outcome slope differs across groups (violated homogeneity of regression slopes) — model the interaction instead.
Categorical covariates — those are factors, not covariates.
Data requirements
Dependent / outcome variable
One continuous outcome.
Independent / grouping variable
One (or more) categorical factor(s), plus one or more continuous covariates measured before/independently of treatment.
Design
Between-subjects with covariate(s).
Sample size guidance
As for ANOVA plus a few cases per covariate; covariates that correlate with the DV raise power.
Assumptions
All ANOVA assumptions (independence, normality of residuals, homogeneity of variance).
Linearity between covariate and outcome within each group.
Homogeneity of regression slopes — the covariate's slope is the same in every group (test the factor × covariate interaction first; it should be non-significant).
Covariate measured reliably and unaffected by the treatment.
Hypotheses
H₀ — Adjusted population means (at the mean of the covariate) are equal across groups.
H₁ — At least one adjusted mean differs.
The concept
ANCOVA fits a regression of the outcome on the covariate, then compares the group means of the residual — what is left after the covariate has explained its share. Equivalently, it compares 'adjusted means': the outcome each group would show if all groups sat at the same covariate value.
Two benefits follow: the error term shrinks (more power) and baseline imbalances are partially adjusted. The critical diagnostic is homogeneity of regression slopes — if the covariate's effect differs across groups, one adjusted comparison cannot summarize the data, and the interaction is itself the finding.
Worked example
Three training formats (in-person, e-learning, blended; n = 30 each) are compared on post-test performance, controlling pre-test scores.
The format × pretest interaction is non-significant (p = .61) — slopes are homogeneous. ANCOVA: F(2, 86) = 5.61, p = .005, ηp² = .115; adjusted means favor blended (75.8) over e-learning (70.1), Bonferroni p = .004.
How to run it
# 1. Test homogeneity of slopes (interaction should be n.s.)
summary(aov(post ~ format * pre, data = df))
# 2. ANCOVA
model <- aov(post ~ pre + format, data = df)
car::Anova(model, type = 3)
library(effectsize); eta_squared(model, partial = TRUE)
library(emmeans)
emmeans(model, ~ format) |> pairs(adjust = "bonferroni") # adjusted means
import pingouin as pg
from statsmodels.formula.api import ols
import statsmodels.api as sm
# 1. Slope homogeneity
m_int = ols("post ~ C(format) * pre", data=df).fit()
print(sm.stats.anova_lm(m_int, typ=3))
# 2. ANCOVA
print(pg.ancova(df, dv="post", covar="pre", between="format"))
Preliminary: GLM → Univariate → move DV, factor, covariate; Model → Custom → include factor, covariate AND factor*covariate; check the interaction is non-significant, then remove it.
Analyze → General Linear Model → Univariate: DV in 'Dependent', factor in 'Fixed Factor(s)', covariate in 'Covariate(s)'.
Options: Descriptives, Effect size. EM Means: factor across, Compare main effects, Bonferroni.
Report the covariate F, the adjusted group F, p, ηp², and the estimated marginal (adjusted) means.
No built-in ANCOVA. Workaround via regression: create dummy variables for the factor, then Data Analysis → Regression with covariate + dummies as predictors.
The dummies' joint contribution is the adjusted group effect (compare R² with and without them via an F-change test computed by hand).
For real ANCOVA output (adjusted means, ηp²), use R/Python/SPSS.
Interpreting the output
First confirm slope homogeneity (interaction p > .05).
Report the adjusted (estimated marginal) means, not raw means — they are the ANCOVA result.
F, df, p, ηp² for the group effect after adjustment; mention the covariate's effect too.
In quasi-experiments, temper causal language: adjustment is not randomization.
APA-style reporting
After confirming homogeneity of regression slopes, a one-way ANCOVA controlling for pre-test scores revealed a significant effect of training format on post-test performance, F(2, 86) = 5.61, p = .005, ηp² = .12. Adjusted means indicated higher performance for blended learning (Madj = 75.8) than e-learning (Madj = 70.1), p = .004.
Common mistakes
Skipping the homogeneity-of-slopes check.
Using ANCOVA to 'equate' groups that differ hugely on the covariate in observational data.
Controlling a variable that the treatment itself changed (a mediator).
Reporting raw instead of adjusted means.
Piling in many weakly related covariates — each costs df and invites overfitting.