Tests whether the strength or direction of an X→Y relationship depends on a third variable W — e.g., does supervisor support buffer the stress→burnout link?
Predict & ExplainMultivariatealso known as: Interaction analysis, Moderated regression
✓ When to use
Theory says an effect is conditional: 'X matters more when W is high/low'.
Boundary-condition questions that make research contributions ('for whom / when does it hold?').
Any combination of continuous/categorical X and W with a continuous outcome.
✗ When NOT to use
The third variable is supposed to transmit (not condition) the effect — that is mediation.
Median-splitting continuous moderators to run subgroup ANOVAs — keep them continuous.
Underpowered samples: interaction effects are notoriously small; N in the several hundreds is often needed.
Focal predictor X, moderator W (continuous or categorical), and their product X×W.
Design
One sample; all variables measured per case.
Sample size guidance
Large — detecting f² ≈ .02 for an interaction at power .80 needs N ≈ 400; plan accordingly.
Assumptions
All multiple-regression assumptions for the full model including the product term.
Both X and W (and their product) measured reliably — the product's reliability is roughly the product of the components' reliabilities.
Adequate variance in W across its range; sparse regions make simple slopes there meaningless.
Hypotheses
H₀ — The interaction coefficient is zero (β₃ = 0): the X→Y slope is constant across W.
H₁ — β₃ ≠ 0: the X→Y slope changes with W.
The concept
Estimate Y = b₀ + b₁X + b₂W + b₃(X×W). The interaction coefficient b₃ is the change in X's slope per one-unit increase in W. Mean-centering X and W before forming the product makes b₁ interpretable as the X effect at the average W (it does not change b₃ or its test).
A significant b₃ is only the beginning: probe it. Simple-slopes analysis re-estimates the X→Y slope at chosen W values (mean ± 1 SD, or meaningful values); the Johnson–Neyman technique finds the exact W range where the X effect is significant. Plot the simple slopes — reviewers and readers understand pictures of interactions far better than coefficients.
Worked example
Does supervisor support (W) buffer the effect of workload (X) on burnout (Y)? N = 385 nurses; X and W mean-centered.
Result: b₃ = −0.14, SE = 0.05, p = .004, ΔR² = .018. Simple slopes: at low support (−1 SD) workload → burnout b = 0.48 (p < .001); at high support (+1 SD) b = 0.20 (p = .03) — support significantly weakens the workload–burnout link.
How to run it
df$Xc <- scale(df$workload, scale = FALSE) # mean-center
df$Wc <- scale(df$support, scale = FALSE)
model <- lm(burnout ~ Xc * Wc, data = df) # * adds product term
summary(model)
library(interactions)
sim_slopes(model, pred = Xc, modx = Wc) # simple slopes + J-N
interact_plot(model, pred = Xc, modx = Wc) # the picture
import statsmodels.formula.api as smf
df["Xc"] = df.workload - df.workload.mean()
df["Wc"] = df.support - df.support.mean()
model = smf.ols("burnout ~ Xc * Wc", data=df).fit()
print(model.summary())
# simple slopes at W = -1SD, 0, +1SD: refit with shifted W or compute
# b1 + b3*w and its SE via the covariance matrix (model.cov_params()).
Model 1; Y = burnout, X = workload, W = support; tick mean-centering, Johnson-Neyman, and plots.
Native alternative: COMPUTE Xc/Wc as centered variables and Int = Xc*Wc; run Linear Regression with Xc, Wc, Int (hierarchically, Int in Block 2, with R² change).
Report b₃ (SE, p), ΔR² for the interaction, and the simple-slope estimates.
Create centered columns Xc = X − mean(X), Wc = W − mean(W), and Int = Xc*Wc.
Data Analysis → Regression with Xc, Wc, Int as X range.
Simple slopes: effect of X at W₀ equals b1 + b3*W₀ — compute at Wc = −SD, 0, +SD and plot two/three fitted lines with a scatter chart.
Significance of each simple slope needs the coefficient covariance — practical only in R/SPSS/PROCESS.
Interpreting the output
b₃ significant → the effect of X is conditional; report ΔR² of the product term.
Simple slopes at meaningful W values tell the substantive story ('significant only for low-support employees').
Johnson–Neyman gives the W threshold where the effect switches on/off.
With centered predictors, b₁ = X effect at average W — say so explicitly.
Non-significant interactions in modest samples are inconclusive, not evidence of no moderation.
APA-style reporting
The workload × support interaction was significant, b = −0.14, SE = 0.05, p = .004, ΔR² = .018. Simple-slopes analysis indicated that workload predicted burnout more strongly at low support (−1 SD; b = 0.48, p < .001) than at high support (+1 SD; b = 0.20, p = .031), consistent with a buffering effect.
Common mistakes
Interpreting b₁ and b₂ as 'main effects' when the model contains their interaction — they are conditional effects at 0 of the other variable.
Median splits on continuous moderators.
Claiming moderation from 'significant in group A, not in group B' without testing the interaction itself.
Underpowered interaction tests over-interpreted in either direction.