Tests whether groups have equal variances — the standard gatekeeper before t-tests and ANOVA decide between classic and Welch versions.
Check AssumptionsBivariatealso known as: Homogeneity of variance test, Brown–Forsythe variant
✓ When to use
Before/alongside independent t-tests and ANOVA to choose classic vs Welch procedures.
Any analysis assuming homoscedasticity across groups.
The Brown–Forsythe (median-centered) variant when data are skewed — more robust.
✗ When NOT to use
As the sole arbiter with large n (trivial variance differences flagged) or tiny n (real ones missed) — look at the actual SDs and their ratio.
For regression's homoscedasticity across CONTINUOUS fitted values — use residual plots/Breusch–Pagan.
Paired designs (variances of what? — the analysis concerns difference scores).
Data requirements
Dependent / outcome variable
One continuous variable.
Independent / grouping variable
One categorical grouping variable (2+ groups).
Design
Between-subjects.
Sample size guidance
Any; interpret jointly with group SDs and an SDmax/SDmin ratio (> 2 is a practical warning).
Assumptions
Independent observations.
Continuous outcome.
(Levene's virtue: it does NOT require normality — unlike the F-ratio/Bartlett tests it replaced.)
Hypotheses
H₀ — All group variances are equal (σ₁² = σ₂² = … = σk²).
H₁ — At least one group's variance differs.
The concept
Levene's trick converts a variance question into a means question: compute each observation's absolute deviation from its group center (mean in the original; median in Brown–Forsythe), then run a one-way ANOVA on those deviations. Groups with bigger spread have bigger average deviations, and the ANOVA detects it.
Median-centering (Brown–Forsythe) is the robust default many packages now use. Decision logic downstream: significant Levene → Welch t/Welch ANOVA and Games–Howell post-hocs; non-significant with similar ns → classic procedures are fine. Increasingly, methodologists suggest skipping the gate and defaulting to Welch, since the two-step 'test-then-choose' inflates error rates slightly — either way, report what you did.
Worked example
Before comparing overtime across three plants, Levene's test (median-centered): F(2, 115) = 5.42, p = .006; SDs 2.1, 5.3, 3.8.
Variances are heterogeneous — analysis proceeds with Welch ANOVA and Games–Howell comparisons.
How to run it
car::leveneTest(overtime ~ plant, data = df) # median-centered default
car::leveneTest(overtime ~ plant, data = df, center = mean)
tapply(df$overtime, df$plant, sd) # look at actual SDs too
from scipy import stats
groups = [g["overtime"].values for _, g in df.groupby("plant")]
print(stats.levene(*groups, center="median")) # Brown–Forsythe
print(df.groupby("plant")["overtime"].std())
Automatic with t-tests: the Independent-Samples T Test output includes Levene's test.
For ANOVA: Analyze → Compare Means → One-Way ANOVA → Options → tick 'Homogeneity of variance test' (and 'Welch' as the ready fallback).
SPSS 26+ reports Levene's based on mean and median centering; prefer the median row for skewed data.
Report F(df1, df2), p, and the group SDs.
Group SDs: =STDEV.S per group; ratio max/min as a heuristic (> 2 = caution).
Manual Levene: compute |xi − group median| per observation, then run 'Anova: Single Factor' on these deviations.
The two-group F-ratio test (=F.TEST) exists but is normality-fragile — Levene on deviations is safer.
Interpreting the output
p < .05 → variances differ → use Welch/Games–Howell downstream.
p ≥ .05 → no detected difference, but check the SD ratio and group sizes before relaxing.
State which centering (mean/median) was used.
Remember it gates OTHER tests — Levene's own result is rarely the finding of interest.
APA-style reporting
Levene's test (median-centered) indicated unequal variances across plants, F(2, 115) = 5.42, p = .006; therefore Welch's ANOVA and Games–Howell post-hoc tests were used.