Compresses many correlated variables into a few uncorrelated components that retain maximum variance — data reduction, not latent-construct discovery.
Reduce DimensionsMultivariatealso known as: Principal components
✓ When to use
You want a smaller set of composite variables for further analysis (indices, dashboards, fighting multicollinearity).
Building a weighted index from correlated indicators (e.g., a firm-performance index from five metrics).
Visualizing high-dimensional data in 2–3 dimensions.
✗ When NOT to use
You believe underlying latent constructs CAUSE the item responses — that is EFA's job; PCA components are formative summaries, not reflective factors.
Validating a measurement scale — use EFA then CFA.
Variables on wildly different scales without standardizing first.
Categorical items with few levels — consider categorical PCA (CATPCA) or polychoric approaches.
Data requirements
Dependent / outcome variable
A set of p correlated numeric variables (standardize unless scales are identical).
Independent / grouping variable
—
Design
One sample; complete or properly imputed data.
Sample size guidance
n > p, ideally n ≥ 5–10 per variable; KMO ≥ .60 signals the correlation matrix is worth compressing.
Assumptions
Variables are correlated enough to compress (Bartlett's test significant; KMO ≥ .60).
Linearity among variables.
Continuous (or at least interval-like) measurement; standardization when units differ.
Outliers screened — components chase variance, and outliers manufacture variance.
Hypotheses
H₀ — — (PCA is a descriptive decomposition; no hypothesis test in standard use).
H₁ — — (decisions concern how many components to retain).
The concept
PCA rotates the coordinate system: the first component is the direction through the data cloud with maximal variance, the second the maximal-variance direction orthogonal to the first, and so on. Each component is an exact weighted sum of the observed variables; eigenvalues state how much variance each captures.
How many to keep? Kaiser's eigenvalue > 1 rule is popular but crude; scree plots help; parallel analysis (compare eigenvalues against random-data eigenvalues) is the current best practice. Component loadings show each variable's contribution; component scores become new variables downstream. PCA analyzes total variance — unlike EFA, which models only shared variance — which is precisely why PCA summarizes but does not 'discover constructs'.
Worked example
An analyst compresses eight correlated branch-performance metrics (sales growth, NPS, footfall, conversion, etc.) into an index. KMO = .78; parallel analysis retains two components explaining 61% of total variance.
Component 1 (42%) loads on commercial metrics — a 'commercial performance' summary; Component 2 (19%) on service metrics. Branch scores on Component 1 feed a league table.
How to run it
vars <- df[, c("growth","nps","footfall","conversion","basket","returns","waittime","complaints")]
library(psych)
KMO(vars); cortest.bartlett(cor(vars), n = nrow(vars))
fa.parallel(vars, fa = "pc") # how many components
pc <- prcomp(vars, scale. = TRUE)
summary(pc) # variance explained
pc$rotation[, 1:2] # loadings
scores <- pc$x[, 1:2] # component scores
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
import numpy as np
X = StandardScaler().fit_transform(df[cols])
pca = PCA().fit(X)
print(pca.explained_variance_) # eigenvalues
print(pca.explained_variance_ratio_.cumsum())
loadings = pca.components_.T * np.sqrt(pca.explained_variance_)
scores = PCA(n_components=2).fit_transform(X)
Analyze → Dimension Reduction → Factor.
Move the variables in. Descriptives: tick KMO and Bartlett's test.
Extraction: Method = Principal components; check the scree plot; set the number of components after inspecting it (or eigenvalue > 1 initially).
Rotation: usually None for pure reduction (Varimax only if interpretability of components matters).
Scores: tick 'Save as variables' to create component scores for later analyses.
No native PCA. Small-scale workaround: standardize variables, compute the correlation matrix (Data Analysis → Correlation), then eigen-decompose externally.
Practical route: run PCA in R/Python/jamovi and import the component scores back into Excel for reporting/dashboards.
Interpreting the output
Eigenvalues / % variance per component and cumulative % — how much information the reduction keeps.
Loadings: which variables define each component; name components by their heavy loaders.
Retention decision justified by parallel analysis or scree, not Kaiser alone.
Component scores are composites for downstream use — remember they are sample-specific weightings.
APA-style reporting
A principal component analysis with parallel analysis supported retaining two components, which together explained 61.2% of total variance (KMO = .78; Bartlett's χ²(28) = 612.4, p < .001). The first component (42.1%) was defined by commercial indicators (loadings .68–.84) and was saved as a commercial-performance index.
Common mistakes
Reporting PCA as if it identified latent constructs — that claim needs EFA/CFA.