Rule of thumb: ≥ 10 events (cases in the rarer category) per predictor; total n often 200+.
Assumptions
Independence of observations.
Linearity of the LOGIT: each continuous predictor relates linearly to log-odds (Box–Tidwell check).
No severe multicollinearity.
No extreme influential cases; adequate events-per-variable.
NO assumptions of normal residuals or homoscedasticity — those belong to OLS.
Hypotheses
H₀ — A predictor's coefficient is zero (OR = 1): it does not change the odds of the outcome.
H₁ — The coefficient differs from zero (OR ≠ 1).
The concept
The model fits log(p/(1−p)) = b₀ + b₁X₁ + … — a straight line in log-odds that becomes an S-shaped curve in probability. Exponentiating a coefficient gives the odds ratio: e^b is the multiplicative change in the odds of the outcome per unit of the predictor. OR > 1 raises the odds; OR < 1 lowers them; OR = 1 is no effect.
Fit is judged by the likelihood-ratio χ² against the null model, pseudo-R² (Nagelkerke, McFadden — report which; they are not OLS R²), and classification-independent discrimination via the ROC curve's AUC (.5 = chance, .7+ acceptable, .8+ good). Odds are not probabilities: an OR of 2 does not mean 'twice as likely' unless the base rate is small.
Worked example
Predicting one-year turnover (18% left) from job satisfaction, salary grade, and commute time (N = 420).
Result: model χ²(3) = 48.2, p < .001, Nagelkerke R² = .18, AUC = .74. Satisfaction OR = 0.55 (each satisfaction point nearly halves quit odds); commute OR = 1.03 per minute (p = .01).
How to run it
model <- glm(quit ~ satisfaction + salary_grade + commute,
data = df, family = binomial)
summary(model)
exp(cbind(OR = coef(model), confint(model))) # ORs with CIs
# fit and discrimination
anova(model, test = "Chisq")
library(pROC); auc(roc(df$quit, fitted(model)))
import statsmodels.formula.api as smf
import numpy as np
from sklearn.metrics import roc_auc_score
model = smf.logit("quit ~ satisfaction + salary_grade + commute", data=df).fit()
print(model.summary())
print(np.exp(model.params)) # odds ratios
print(roc_auc_score(df["quit"], model.predict()))
Analyze → Regression → Binary Logistic.
DV into Dependent; predictors into Covariates (declare categorical ones via 'Categorical…').
Options: tick CI for exp(B), Hosmer-Lemeshow goodness-of-fit, Classification plots.
Report the omnibus model χ², Nagelkerke R², and per predictor B, SE, Wald, p, Exp(B) with CI.
No native logistic regression. Possible via Solver: set up log-likelihood = Σ[y*ln(p)+(1−y)*ln(1−p)] with p = 1/(1+EXP(−(b0+b1X1+…))) and maximize over the b cells.
Workable for teaching; for real analyses use R, Python, SPSS, or jamovi.
Interpreting the output
Each OR: multiplicative change in odds per unit predictor, other predictors held constant; CI excluding 1 = significant.
Model χ² and pseudo-R² for overall fit; AUC for discrimination.
Translate to predicted probabilities for concrete scenarios — stakeholders think in probabilities, not odds.
Check the events-per-variable ratio before trusting the coefficients.
APA-style reporting
Binary logistic regression significantly predicted turnover, χ²(3, N = 420) = 48.2, p < .001, Nagelkerke R² = .18. Job satisfaction reduced the odds of quitting, OR = 0.55, 95% CI [0.42, 0.71], p < .001, while longer commutes increased them, OR = 1.03 per minute, 95% CI [1.01, 1.05], p = .010.
Common mistakes
Interpreting odds ratios as risk ratios ('twice as likely') at high base rates.
Judging the model only by accuracy at a 0.5 cut-off with imbalanced classes.
Ignoring the linearity-of-logit assumption for continuous predictors.
Too many predictors for too few events — overfitted, unstable ORs.