Statistical Methods Atlas
Atlas › Reduce Dimensions › Exploratory Factor Analysis (EFA)

Exploratory Factor Analysis (EFA)

Uncovers the latent factors underlying a set of items — the standard first step in developing a new measurement scale.

Reduce DimensionsMultivariate also known as: Common factor analysis, Principal axis factoring

✓ When to use

  • Scale development: which items cluster into which constructs?
  • No strong prior structure — the factor pattern is genuinely open.
  • Checking dimensionality of an item pool before CFA on a fresh sample.

✗ When NOT to use

  • A clear hypothesized structure exists — go straight to CFA.
  • Pure data compression without latent-variable claims — PCA.
  • Sample too small or correlations too weak (KMO < .60).
  • Running EFA and CFA on the SAME sample as 'validation' — that is double-dipping; split the sample or collect anew.

Data requirements

Dependent / outcome variableA set of items (ideally 4+ per expected factor), interval-like; use polychoric correlations for strongly ordinal items.
Independent / grouping variable—
DesignOne sample; items measured on comparable scales.
Sample size guidanceCommon guidance: n ≥ 200 or 10 cases per item, whichever is larger; stronger loadings tolerate smaller n.

Assumptions

Hypotheses

H₀ — — (exploratory; formal tests appear in ML extraction's fit test, but retention is the real decision).
H₁ — —

The concept

EFA models each item as loading on common factors plus a unique term, analyzing only the shared variance among items (communalities on the diagonal) — the key difference from PCA. Extraction (principal axis factoring or maximum likelihood) finds the factors; rotation then makes them interpretable: varimax forces uncorrelated factors, while oblique rotations (promax, oblimin) allow correlated factors — the realistic default in social science, where constructs correlate.

Retention should rest on parallel analysis plus interpretability, not the eigenvalue > 1 default. Read the pattern matrix: items load saliently (≥ .40) on their factor, cross-load weakly (< .30) elsewhere; violators are dropped one at a time with re-analysis. The factor correlation matrix from an oblique rotation previews discriminant validity questions that CFA will formally test.

Worked example

A researcher pilots 18 new items intended to measure workplace wellbeing on n = 260 employees. KMO = .87; parallel analysis suggests three factors.

PAF with promax rotation yields clean factors: physical (5 items, loadings .58–.81), psychological (6 items), social (4 items); three cross-loading items are dropped. Factors correlate .38–.52; total variance explained 54%.

How to run it

library(psych)
items <- df[, paste0("wb", 1:18)]

KMO(items); cortest.bartlett(cor(items), n = nrow(items))
fa.parallel(items, fa = "fa")                     # retention

efa <- fa(items, nfactors = 3, fm = "pa", rotate = "promax")
print(efa, cut = .30, sort = TRUE)                # pattern matrix
efa$Phi                                            # factor correlations

Interpreting the output

APA-style reporting

Exploratory factor analysis (principal axis factoring, promax rotation) on the 18 wellbeing items (KMO = .87; Bartlett's χ²(153) = 2,114, p < .001) supported a three-factor solution based on parallel analysis, explaining 54.3% of common variance. All retained items loaded ≥ .58 on their intended factor with cross-loadings < .30; inter-factor correlations ranged from .38 to .52.

Common mistakes

Related methods

Principal Component Analysis (PCA)Summary components, not latent factorsConfirmatory Factor Analysis (CFA)Confirmatory follow-upCronbach's AlphaReliability of resulting scalesConvergent & Discriminant ValidityValidity evidence next
← Principal Component Analysis (PCA)Confirmatory Factor Analysis (CFA) →