Tests whether your data fit a pre-specified measurement model — the standard evidence that items measure the constructs you claim.
Validate ConstructsMultivariatealso known as: Measurement model analysis
✓ When to use
A theory- or EFA-based factor structure exists and needs formal testing on (ideally new) data.
Before structural models: establish the measurement model first (two-step logic).
Comparing rival structures (one factor vs three; bifactor vs correlated factors).
Producing the ingredients for reliability/validity reporting: standardized loadings, CR, AVE.
✗ When NOT to use
No prior structure — EFA first.
Sample far too small for the parameter count.
Formative indicators (items CAUSE the construct, e.g., SES) — reflective CFA is misspecified; consider PLS-SEM or composite approaches.
Single-item measures — nothing to factor-analyze.
Data requirements
Dependent / outcome variable
Items assigned a priori to factors; ≥ 3 indicators per factor (2 possible with conditions).
Independent / grouping variable
—
Design
One sample, preferably distinct from the EFA sample.
Sample size guidance
n ≥ 200 as a floor for typical models; more with many parameters, ordinal estimators, or weak loadings.
Assumptions
Correctly specified model: each item loads only on its assigned factor; residuals uncorrelated unless justified.
Multivariate normality for ML (use robust ML — MLR — for moderate violations; WLSMV for ordinal items).
Adequate sample size for stable estimation.
Independent observations; interval-like indicators (or an ordinal estimator).
Hypotheses
H₀ — The model-implied covariance matrix equals the population covariance matrix (the model fits).
H₁ — The model does not fit. (Note the reversal: researchers usually hope NOT to reject H₀.)
The concept
CFA fixes which items load on which latent factors and estimates loadings, factor variances/covariances, and residuals so as to reproduce the observed covariance matrix as closely as possible. Misfit means your theoretical structure cannot account for how the items actually covary.
Fit is judged by a battery, not one number: χ² (sensitive to n), CFI/TLI ≥ .90 acceptable / ≥ .95 good, RMSEA ≤ .08 acceptable / ≤ .06 good, SRMR ≤ .08. Standardized loadings should exceed ~.60–.70. Modification indices can suggest improvements, but chasing them turns confirmation back into exploration — any data-driven change needs theoretical defense and ideally fresh-sample validation. CFA output feeds composite reliability and AVE for the validity story.
Worked example
Testing the three-factor wellbeing model from a prior EFA on a new sample (n = 410) with 15 items.
Result: χ²(87) = 201.5, p < .001; CFI = .953; TLI = .943; RMSEA = .057 [.047, .067]; SRMR = .048. Standardized loadings .62–.86. The three-factor model beats a one-factor alternative (Δχ² p < .001) — measurement structure confirmed.
Compare rival models with Δχ² (nested) or AIC/BIC (non-nested) — beating alternatives is stronger than fitting alone.
Any modification (correlated residuals) must be disclosed and justified.
Then compute CR and AVE for the reliability/validity narrative.
APA-style reporting
The hypothesized three-factor model showed good fit, χ²(87) = 201.5, p < .001, CFI = .95, TLI = .94, RMSEA = .057, 90% CI [.047, .067], SRMR = .048, and fit significantly better than a one-factor model, Δχ²(3) = 412.7, p < .001. Standardized loadings ranged from .62 to .86 (all p < .001).
Common mistakes
Judging fit by χ² alone in large samples (or ignoring it entirely).
Freeing correlated residuals wholesale via modification indices.
Skipping measurement-model testing and jumping to structural paths.
ML on severely ordinal/skewed items instead of MLR/WLSMV.
Presenting a re-specified model as if it were the a priori one.