Finds the weighted combinations of predictors that best separate known groups, and classifies new cases — e.g., which ratios distinguish surviving from failing firms.
Classify & GroupMultivariatealso known as: Linear discriminant analysis, Fisher's discriminant
✓ When to use
Group membership is KNOWN and you want the combination of continuous predictors that separates groups (descriptive) or a classification rule (predictive).
When LDA's assumptions hold, it is more efficient than logistic regression at small n.
✗ When NOT to use
Predictors include many categoricals — logistic regression handles them naturally.
Covariance matrices differ strongly across groups — quadratic DA or logistic regression.
Groups to be DISCOVERED, not predicted — cluster analysis.
Heavily non-normal predictors — logistic regression is assumption-lighter and usually preferred today.
Data requirements
Dependent / outcome variable
One categorical grouping variable (2+ groups).
Independent / grouping variable
Two or more continuous predictors.
Design
Between-subjects; ideally a holdout sample or cross-validation for honest accuracy.
Sample size guidance
Smallest group n > number of predictors; ≥ 20 cases per predictor overall recommended.
Assumptions
Multivariate normality of predictors within groups.
Equal group covariance matrices (Box's M; if violated, quadratic DA).
No severe multicollinearity; outliers screened.
Prior probabilities specified sensibly (group sizes or theory).
Hypotheses
H₀ — Group centroids on the discriminant function(s) are equal (no separation).
H₁ — Centroids differ (the function separates groups).
The concept
LDA derives discriminant functions — linear combinations of predictors — maximizing between-group relative to within-group variance (up to min(groups − 1, predictors) functions). Wilks' Λ tests separation; standardized coefficients and structure loadings reveal which predictors carry it.
For classification, each case's function scores place it nearest one group centroid, weighted by priors. Judge accuracy against the base rate (proportional chance criterion), and trust cross-validated (leave-one-out) accuracy over the flattering resubstitution figure. Conceptually, LDA is MANOVA in reverse: MANOVA asks whether groups differ on the variable set; LDA asks which combination does the differing.
Worked example
Distinguishing 90 surviving vs 60 failed SMEs using liquidity, profitability, and leverage ratios.
Result: Wilks' Λ = .71, χ²(3) = 49.8, p < .001, canonical R = .54. Leverage (loading .81) dominates. Leave-one-out classification: 78% correct vs 52% proportional chance — a useful early-warning rule.
from sklearn.discriminant_analysis import LinearDiscriminantAnalysis
from sklearn.model_selection import cross_val_score
X, y = df[["liquidity", "profit", "leverage"]], df["status"]
lda = LinearDiscriminantAnalysis()
print(cross_val_score(lda, X, y, cv=10).mean()) # honest accuracy
lda.fit(X, y)
print(lda.coef_, lda.means_)
Analyze → Classify → Discriminant.
Grouping Variable (define range); predictors into Independents.
Statistics: Means, Box's M, Function coefficients (standardized).
Classify: tick 'Leave-one-out classification' and set Priors (compute from group sizes if unbalanced).
Report Wilks' Λ with χ² and p, canonical correlation, structure matrix, and cross-validated classification %.
Two-group case only, laboriously: compute group mean vectors and pooled covariance, invert it (MINVERSE), and form the discriminant weights w = S⁻¹(μ₁−μ₂) with MMULT.
Score cases with SUMPRODUCT and classify by cut-off between group mean scores.
For real work (tests, cross-validation), use R/Python/SPSS.
Interpreting the output
Wilks' Λ (with χ², df, p): does separation exist? Canonical correlation² = variance in function scores explained by groups.
Structure loadings (> |.30| notable) name the function.
Cross-validated accuracy vs proportional chance — the honest usefulness metric.
Check Box's M; switch to QDA/logistic if violated.
APA-style reporting
Linear discriminant analysis significantly separated surviving and failed firms, Wilks' Λ = .71, χ²(3, N = 150) = 49.8, p < .001, canonical R = .54. Leverage dominated the function (structure loading = .81). Leave-one-out classification achieved 78.0% accuracy, exceeding the proportional chance criterion of 52%.
Common mistakes
Reporting resubstitution accuracy as if it were out-of-sample performance.
Ignoring Box's M violations.
Comparing accuracy to 50% when groups are unbalanced.
Stepwise predictor selection without validation.
Using LDA where logistic regression's weaker assumptions fit the data better — and not saying why.