Statistical Methods Atlas
Atlas › Classify & Group › Discriminant Analysis (LDA)

Discriminant Analysis (LDA)

Finds the weighted combinations of predictors that best separate known groups, and classifies new cases — e.g., which ratios distinguish surviving from failing firms.

Classify & GroupMultivariate also known as: Linear discriminant analysis, Fisher's discriminant

✓ When to use

  • Group membership is KNOWN and you want the combination of continuous predictors that separates groups (descriptive) or a classification rule (predictive).
  • Classic applications: bankruptcy prediction (Altman's Z), selection research, MANOVA follow-up.
  • When LDA's assumptions hold, it is more efficient than logistic regression at small n.

✗ When NOT to use

  • Predictors include many categoricals — logistic regression handles them naturally.
  • Covariance matrices differ strongly across groups — quadratic DA or logistic regression.
  • Groups to be DISCOVERED, not predicted — cluster analysis.
  • Heavily non-normal predictors — logistic regression is assumption-lighter and usually preferred today.

Data requirements

Dependent / outcome variableOne categorical grouping variable (2+ groups).
Independent / grouping variableTwo or more continuous predictors.
DesignBetween-subjects; ideally a holdout sample or cross-validation for honest accuracy.
Sample size guidanceSmallest group n > number of predictors; ≥ 20 cases per predictor overall recommended.

Assumptions

Hypotheses

H₀ — Group centroids on the discriminant function(s) are equal (no separation).
H₁ — Centroids differ (the function separates groups).

The concept

LDA derives discriminant functions — linear combinations of predictors — maximizing between-group relative to within-group variance (up to min(groups − 1, predictors) functions). Wilks' Λ tests separation; standardized coefficients and structure loadings reveal which predictors carry it.

For classification, each case's function scores place it nearest one group centroid, weighted by priors. Judge accuracy against the base rate (proportional chance criterion), and trust cross-validated (leave-one-out) accuracy over the flattering resubstitution figure. Conceptually, LDA is MANOVA in reverse: MANOVA asks whether groups differ on the variable set; LDA asks which combination does the differing.

Worked example

Distinguishing 90 surviving vs 60 failed SMEs using liquidity, profitability, and leverage ratios.

Result: Wilks' Λ = .71, χ²(3) = 49.8, p < .001, canonical R = .54. Leverage (loading .81) dominates. Leave-one-out classification: 78% correct vs 52% proportional chance — a useful early-warning rule.

How to run it

library(MASS)
lda_fit <- lda(status ~ liquidity + profit + leverage, data = df)
lda_fit                      # coefficients, group means, priors

# leave-one-out cross-validated accuracy
pred_cv <- lda(status ~ liquidity + profit + leverage, data = df, CV = TRUE)
mean(pred_cv$class == df$status)
table(df$status, pred_cv$class)

candisc::candisc(lm(cbind(liquidity, profit, leverage) ~ status, df)) # descriptive view

Interpreting the output

APA-style reporting

Linear discriminant analysis significantly separated surviving and failed firms, Wilks' Λ = .71, χ²(3, N = 150) = 49.8, p < .001, canonical R = .54. Leverage dominated the function (structure loading = .81). Leave-one-out classification achieved 78.0% accuracy, exceeding the proportional chance criterion of 52%.

Common mistakes

Related methods

Binary Logistic RegressionAssumption-lighter rivalMANOVA (Multivariate ANOVA)Mirror-image analysisK-Means ClusteringGroups unknown
← Hierarchical Cluster AnalysisTrend Analysis & Decomposition →