Predicts membership in three or more unordered categories — e.g., which benefits package (cash/insurance/leave) employees choose.
Predict & ExplainMultivariatealso known as: Polytomous logistic regression, Baseline-category logit
✓ When to use
Outcome has 3+ categories with no natural order: career track chosen, brand preferred, exit destination (competitor/retirement/other).
You want odds ratios comparing each category against a reference category, with covariates.
✗ When NOT to use
Ordered categories — ordinal logistic (uses the ordering, more power, fewer parameters).
Two categories — binary logistic.
Categories that are alternatives with attributes of their own (price, distance) — conditional logit/discrete-choice models.
Tiny categories — merge or collect more data; each contrast needs its own events.
Data requirements
Dependent / outcome variable
One nominal variable with 3+ categories; choose a meaningful reference category.
Independent / grouping variable
Continuous and/or dummy-coded predictors.
Design
Independent observations; each case in exactly one category.
Sample size guidance
Events-per-variable logic applies per outcome contrast — samples need to be larger than for binary models (each extra category adds a full coefficient set).
Assumptions
Independence of observations.
Independence of irrelevant alternatives (IIA): relative odds between two categories unaffected by presence of others — think hard when categories are close substitutes.
Linearity in the logit for continuous predictors; no perfect separation; manageable multicollinearity.
Hypotheses
H₀ — A predictor's coefficients are zero across all contrasts: it does not distinguish any category from the reference.
H₁ — At least one contrast's coefficient differs from zero.
The concept
With k categories and category k as reference, the model fits k − 1 simultaneous binary-style equations: log(P(cat j)/P(cat k)) = b₀ⱼ + b₁ⱼX₁ + … Each predictor therefore gets k − 1 odds ratios, one per contrast, and its overall contribution is tested with a likelihood-ratio χ² pooling the contrasts.
Interpretation discipline matters: every OR is relative to the reference category ('a satisfaction point multiplies the odds of choosing insurance over cash by 1.4'). Predicted probabilities across representative covariate profiles are usually the clearest way to present results.
Worked example
Employees (N = 450) choose one of three benefit packages: cash (reference, 38%), insurance (34%), extra leave (28%); predictors: age and family size.
Result: model χ²(4) = 61.3, p < .001. Age: OR = 1.06 per year for insurance-vs-cash (p < .001), n.s. for leave-vs-cash. Family size: OR = 1.52 for insurance-vs-cash, OR = 0.80 for leave-vs-cash.
How to run it
library(nnet)
df$package <- relevel(factor(df$package), ref = "cash")
model <- multinom(package ~ age + family_size, data = df)
summary(model)
exp(coef(model)) # ORs per contrast
# z and p values
z <- summary(model)$coefficients / summary(model)$standard.errors
2 * pnorm(abs(z), lower.tail = FALSE)
import statsmodels.formula.api as smf
import numpy as np
model = smf.mnlogit("package ~ age + family_size", data=df).fit()
print(model.summary())
print(np.exp(model.params)) # ORs; columns = non-reference categories
Analyze → Regression → Multinomial Logistic.
DV into Dependent (set the Reference Category button), covariates/factors accordingly.
Statistics: tick Likelihood ratio tests, Pseudo R-square, Classification table.
Report the model χ², per-predictor LR tests, and per-contrast B, Wald, p, Exp(B) with CI, naming the reference category everywhere.
Not feasible natively (k−1 simultaneous equations).
Use R (nnet::multinom), Python (statsmodels mnlogit), SPSS, or jamovi; Excel can present the resulting predicted-probability tables.
Interpreting the output
State the reference category up front; every OR is 'versus reference'.
Per-predictor likelihood-ratio test = does it matter anywhere; per-contrast Wald tests = where.
Predicted probabilities for typical profiles beat raw coefficient tables for communication.
Pseudo-R² and classification accuracy as global summaries (with base-rate context).
APA-style reporting
Multinomial logistic regression (reference: cash) showed the model significantly predicted package choice, χ²(4, N = 450) = 61.3, p < .001, Nagelkerke R² = .14. Family size increased the odds of choosing insurance over cash, OR = 1.52, 95% CI [1.24, 1.86], p < .001, and decreased the odds of choosing extra leave over cash, OR = 0.80, 95% CI [0.65, 0.98], p = .034.
Common mistakes
Using multinomial when the outcome is actually ordered.
Forgetting which category is the reference mid-interpretation.
Ignoring IIA when categories are near-substitutes.
Reading the classification table without comparing to the majority-class base rate.