Evidence that a scale's items converge on their own construct (AVE, CR) and that constructs are empirically distinct from each other (Fornell–Larcker, HTMT).
Validate ConstructsMultivariatealso known as: AVE analysis, Fornell–Larcker, HTMT
✓ When to use
Any survey study with multi-item latent constructs — reviewers expect this table.
After CFA/PLS measurement estimation, before interpreting structural results.
As a substitute for content validity — statistics cannot rescue items that never sampled the construct domain.
Cutoffs applied mechanically to reject borderline scales without judgment.
Data requirements
Dependent / outcome variable
Standardized loadings and construct correlations from a fitted CFA (or PLS) model.
Independent / grouping variable
—
Design
Same data as the measurement model.
Sample size guidance
Whatever the CFA/PLS required.
Assumptions
A reasonably fitting measurement model — validity indices from a misfitting CFA are meaningless.
Reflective indicators.
Constructs measured on comparable populations/conditions.
Hypotheses
H₀ — — (criterion-based assessment rather than hypothesis testing; HTMT can be tested against a threshold via bootstrap CI).
H₁ — —
The concept
Convergent validity asks whether items assigned to a construct share enough variance with it: standardized loadings ≥ .70 (≥ .60 tolerable), composite reliability CR ≥ .70, and Average Variance Extracted AVE ≥ .50 — the construct explains at least half its items' variance. AVE = mean of squared standardized loadings.
Discriminant validity asks whether supposedly different constructs are empirically distinguishable. Fornell–Larcker: each construct's √AVE must exceed its correlations with every other construct (it shares more variance with its own items than with other constructs). The modern, more sensitive criterion is HTMT — the ratio of between-construct to within-construct item correlations — with thresholds of .85 (strict) or .90 (lenient), ideally with a bootstrap CI excluding the threshold. Failing discriminant validity signals construct redundancy: consider merging constructs or revisiting items.
Worked example
After CFA of a model with autonomy, engagement, and satisfaction (n = 410): loadings .64–.88; CR = .86/.89/.84; AVE = .61/.58/.57.
√AVE values (.78/.76/.75) exceed all inter-construct correlations (max .62); HTMT max = .71 with 95% CI upper bound .79 < .85. Both convergent and discriminant validity supported.
How to run it
library(lavaan); library(semTools)
fit <- cfa(model, data = df, estimator = "MLR")
reliability(fit) # alpha, omega/CR, AVE per construct
# Fornell–Larcker: compare sqrt(AVE) to latent correlations
lavInspect(fit, "cor.lv")
htmt(model, data = df) # HTMT matrix (semTools)
import semopy
import numpy as np
m = semopy.Model(model); m.fit(df)
est = m.inspect(std_est=True)
# AVE per construct = mean of squared std loadings; CR from loadings/residuals
load = est[(est.op == "~") | (est.op == "=~")]
# compute per factor: AVE = mean(l^2); CR = (Σl)^2 / ((Σl)^2 + Σ(1−l^2))
# HTMT: compute from item correlation matrix (heterotrait vs monotrait means)
SPSS/Amos give standardized loadings and latent correlations; compute CR and AVE in a spreadsheet: CR = (Σλ)²/((Σλ)² + Σ(1−λ²)); AVE = Σλ²/k.
Fornell–Larcker: put √AVE on the diagonal of the latent correlation matrix and compare.
HTMT: easiest via the free semTools in R or SmartPLS; James Gaskin's Excel 'Stats Tools Package' also computes it from correlation input.
From CFA loadings: AVE =SUMSQ(loadings)/k; CR =(SUM(l))^2/((SUM(l))^2 + SUM(1−l^2)).
Build the Fornell–Larcker table: construct correlations with √AVE on the diagonal.
HTMT from the item correlation matrix: mean heterotrait-heteromethod correlation ÷ geometric mean of the two monotrait means — laborious but feasible for few constructs.
Interpreting the output
Convergent: all loadings sig. and ≥ .60–.70, CR ≥ .70, AVE ≥ .50 (AVE slightly < .50 with CR ≥ .70 is arguable — cite Fornell & Larcker's own concession).
Report a single table combining CR, AVE, correlations and √AVE diagonal — the standard journal exhibit.
Failures are findings: merge constructs, drop cross-loaders, or rethink the theory.
APA-style reporting
All constructs demonstrated convergent validity (standardized loadings .64–.88; CR = .84–.89; AVE = .57–.61). Discriminant validity was supported: each construct's √AVE exceeded its correlations with other constructs, and all HTMT values were below .85 (max = .71, 95% CI [.62, .79]).
Common mistakes
Computing AVE/CR from a badly fitting CFA.
Relying on Fornell–Larcker alone — HTMT detects problems it misses.
Applying reflective criteria to formative constructs.
Deleting items purely to push AVE over .50, gutting content coverage.
Confusing discriminant validity (constructs distinct) with low correlation (constructs can validly correlate .6).