Models a continuous outcome from several predictors at once, giving each predictor's unique contribution — the workhorse of survey-based management research.
Predict & ExplainMultivariatealso known as: OLS regression, Multiple OLS
✓ When to use
One continuous outcome and two or more predictors (continuous or dummy-coded).
You want each predictor's effect holding the others constant, or the best linear prediction of Y.
Testing theory-driven models: do autonomy and feedback predict engagement over and above tenure?
Severe multicollinearity among predictors (VIF > 10) — combine, drop, or use ridge.
More predictors than cases can support.
Nested/clustered data — mixed models; time-series outcomes — models with autocorrelation.
Data requirements
Dependent / outcome variable
One continuous variable.
Independent / grouping variable
Two or more predictors; categorical ones entered as dummy variables (k − 1 dummies).
Design
One sample; independent observations.
Sample size guidance
Old rules: ≥ 10–15 cases per predictor; better: power analysis on expected f² (e.g., f² = .15, 5 predictors, power .80 → N ≈ 92).
Assumptions
Linearity of each predictor–outcome relation (partial regression plots).
Independence of residuals (Durbin–Watson ≈ 2 for ordered data).
Homoscedasticity of residuals.
Normality of residuals.
No severe multicollinearity (VIF < 5–10; tolerance > .1–.2).
No unduly influential cases (Cook's distance, leverage).
Hypotheses
H₀ — Model-level: all slopes are zero (R² = 0). Predictor-level: βj = 0 given the other predictors.
H₁ — Model-level: at least one slope is non-zero. Predictor-level: βj ≠ 0.
The concept
OLS finds the weights that minimize squared prediction errors using all predictors jointly. Each coefficient bj is the expected change in Y for a one-unit change in Xj holding the other predictors constant — a 'unique contribution' logic that makes multiple regression the standard tool for statistical control.
Standardized betas allow rough comparison of predictor importance on a common scale; squared semi-partial correlations give each predictor's unique slice of R². Adjusted R² corrects the optimism of adding predictors. Two different projects — explanation (theory-driven, enter variables by design) and prediction (maximize out-of-sample accuracy, use validation) — should not be mixed casually; avoid mechanical stepwise selection for explanatory work.
Worked example
Predicting work engagement (1–7) from autonomy, supervisor feedback, and tenure among 214 employees.
import statsmodels.formula.api as smf
from statsmodels.stats.outliers_influence import variance_inflation_factor
model = smf.ols("engagement ~ autonomy + feedback + tenure", data=df).fit()
print(model.summary())
X = df[["autonomy", "feedback", "tenure"]].assign(const=1)
for i, c in enumerate(X.columns):
print(c, variance_inflation_factor(X.values, i))
Analyze → Regression → Linear; DV into Dependent, predictors into Independent(s) (Method: Enter).
Statistics: Estimates, Confidence intervals, Model fit, Collinearity diagnostics, Durbin-Watson.
Plots: ZRESID vs ZPRED; Histogram + Normal P-P of residuals.
Report R², adjusted R², model F, and per predictor b (SE), β, t, p, and VIF.
Data → Data Analysis → Regression; Y range = outcome, X range = adjacent predictor columns (max 16).
Tick Residuals and residual plots.
Output: R Square/Adjusted R Square, ANOVA F, and per-predictor coefficients with t and p.
Dummy-code categorical predictors yourself (0/1 columns); Excel provides no VIF — check predictor intercorrelations with =CORREL.
Interpreting the output
Model level: R², adjusted R², F, p — does the set predict at all?
Each b in raw units ('one more feedback point ≈ .21 engagement points, holding others constant'); β for relative weight.
CIs and p per predictor; non-significance means no unique contribution — the variable may still correlate bivariately.
Diagnostics (VIF, residual plots) are part of the result, not an optional extra.
APA-style reporting
Multiple regression showed that autonomy, feedback, and tenure jointly predicted engagement, R² = .31, F(3, 210) = 31.4, p < .001. Autonomy (β = .38, p < .001) and feedback (β = .21, p = .002) contributed uniquely, whereas tenure did not (β = .06, p = .335).
Common mistakes
'Controlling for' colliders or mediators, distorting the focal effect.
Comparing raw b's of variables on different scales instead of β.
Stepwise selection presented as theory testing.
Ignoring multicollinearity, then over-interpreting unstable signs.
Dropping non-significant predictors and re-running until 'clean' — garden of forking paths.
Interpreting adjusted R² change without a formal F-change test.