Tests whether the same participants' means differ across three or more occasions or conditions — e.g., engagement measured quarterly over a year.
Compare GroupsMultivariatealso known as: Within-subjects ANOVA, RM-ANOVA
✓ When to use
One group measured on the same continuous outcome at 3+ time points or under 3+ conditions.
Longitudinal designs, learning curves, or within-person experimental manipulations.
You want the power advantage of each person serving as their own control.
✗ When NOT to use
Different people at each occasion — one-way ANOVA.
Only two occasions — paired t-test.
Ordinal outcomes or badly non-normal data — Friedman test.
Substantial dropout or unequally spaced, person-varying measurement times — use linear mixed models, which handle missing data far better.
Data requirements
Dependent / outcome variable
One continuous outcome measured k ≥ 3 times per person.
Independent / grouping variable
The within-subjects factor (time or condition).
Design
Within-subjects; complete data per person for classic RM-ANOVA (listwise deletion otherwise).
Sample size guidance
n ≈ 30 with complete data is a reasonable floor; power rises with the correlation among occasions.
Assumptions
Independence of participants (rows).
Normality of the outcome at each occasion (of the within-person differences, more precisely).
Sphericity — equal variances of all pairwise difference scores (Mauchly's test); when violated, apply the Greenhouse–Geisser (conservative) or Huynh–Feldt correction.
No missing occasions (or switch to mixed models).
Hypotheses
H₀ — The population means are equal across all occasions (μ₁ = μ₂ = … = μk).
H₁ — At least one occasion's mean differs.
The concept
RM-ANOVA removes stable between-person variance from the error term: variability that comes from some people simply scoring high overall no longer counts against the time effect. What remains tests whether the within-person pattern across occasions is flatter or steeper than chance.
Sphericity is the price of admission: if the variance of (T1−T2) differs from that of (T1−T3), F is inflated. Mauchly's test flags it; the Greenhouse–Geisser ε multiplies both df downward to compensate. Follow a significant F with pairwise comparisons using a Bonferroni adjustment, or better, polynomial trend contrasts when time is ordered.
Worked example
Employee engagement (1–5) measured at onboarding, 3 months, 6 months and 12 months (n = 42 complete cases). Mauchly's test is significant (p = .02), Greenhouse–Geisser ε = .81.
Result: F(2.43, 99.6) = 7.28, p < .001, ηp² = .15 — engagement peaks at 3 months (M = 3.95) then declines by 12 months (M = 3.52; Bonferroni p = .01).
How to run it
library(rstatix) # long format: id, time, engagement
res <- anova_test(data = df, dv = engagement, wid = id, within = time)
get_anova_table(res) # auto-applies GG correction when needed
pairwise_t_test(df, engagement ~ time, paired = TRUE,
p.adjust.method = "bonferroni")
import pingouin as pg # long format
print(pg.sphericity(df, dv="engagement", subject="id", within="time"))
print(pg.rm_anova(df, dv="engagement", subject="id", within="time", correction=True))
print(pg.pairwise_tests(df, dv="engagement", subject="id", within="time",
padjust="bonf"))
Analyze → General Linear Model → Repeated Measures.
Define the within-subject factor (e.g., time, 4 levels) → Add → Define; assign the four columns.
Options: tick Descriptives and Estimates of effect size. EM Means: move the factor across, tick Compare main effects, Bonferroni.
Check Mauchly's Test; if p < .05 read the Greenhouse–Geisser row of 'Tests of Within-Subjects Effects'.
Report F with corrected df, p, ηp², and the pairwise pattern.
Data → Data Analysis → 'Anova: Two-Factor Without Replication' (rows = participants, columns = occasions) approximates RM-ANOVA.
Read the 'Columns' F for the occasion effect.
Excel offers no sphericity test or correction — results are approximate; use R/Python/SPSS for publication.
Interpreting the output
Mauchly's test first; if violated, report corrected df (they become decimals) and say which correction was applied.
F, df, p, ηp² for the within-subject effect.
Pairwise (Bonferroni) comparisons or trend contrasts to describe the shape over time.
Means and SDs per occasion — a plot of the trajectory helps.
APA-style reporting
Mauchly's test indicated a violation of sphericity, χ²(5) = 13.4, p = .02; degrees of freedom were corrected using Greenhouse–Geisser ε = .81. Engagement differed significantly across occasions, F(2.43, 99.6) = 7.28, p < .001, ηp² = .15; Bonferroni comparisons showed engagement at 12 months (M = 3.52, SD = 0.66) was lower than at 3 months (M = 3.95, SD = 0.58), p = .010.
Common mistakes
Ignoring sphericity and reporting uncorrected df.
Listwise-deleting participants with one missing wave without considering mixed models.
Treating occasions as independent groups.
Not adjusting pairwise follow-ups for multiple comparisons.
Interpreting a time effect as causal when no control group exists.