Statistical Methods Atlas
Atlas › Reduce Dimensions › Principal Component Analysis (PCA)

Principal Component Analysis (PCA)

Compresses many correlated variables into a few uncorrelated components that retain maximum variance — data reduction, not latent-construct discovery.

Reduce DimensionsMultivariate also known as: Principal components

✓ When to use

  • You want a smaller set of composite variables for further analysis (indices, dashboards, fighting multicollinearity).
  • Building a weighted index from correlated indicators (e.g., a firm-performance index from five metrics).
  • Visualizing high-dimensional data in 2–3 dimensions.

✗ When NOT to use

  • You believe underlying latent constructs CAUSE the item responses — that is EFA's job; PCA components are formative summaries, not reflective factors.
  • Validating a measurement scale — use EFA then CFA.
  • Variables on wildly different scales without standardizing first.
  • Categorical items with few levels — consider categorical PCA (CATPCA) or polychoric approaches.

Data requirements

Dependent / outcome variableA set of p correlated numeric variables (standardize unless scales are identical).
Independent / grouping variable—
DesignOne sample; complete or properly imputed data.
Sample size guidancen > p, ideally n ≥ 5–10 per variable; KMO ≥ .60 signals the correlation matrix is worth compressing.

Assumptions

Hypotheses

H₀ — — (PCA is a descriptive decomposition; no hypothesis test in standard use).
H₁ — — (decisions concern how many components to retain).

The concept

PCA rotates the coordinate system: the first component is the direction through the data cloud with maximal variance, the second the maximal-variance direction orthogonal to the first, and so on. Each component is an exact weighted sum of the observed variables; eigenvalues state how much variance each captures.

How many to keep? Kaiser's eigenvalue > 1 rule is popular but crude; scree plots help; parallel analysis (compare eigenvalues against random-data eigenvalues) is the current best practice. Component loadings show each variable's contribution; component scores become new variables downstream. PCA analyzes total variance — unlike EFA, which models only shared variance — which is precisely why PCA summarizes but does not 'discover constructs'.

Worked example

An analyst compresses eight correlated branch-performance metrics (sales growth, NPS, footfall, conversion, etc.) into an index. KMO = .78; parallel analysis retains two components explaining 61% of total variance.

Component 1 (42%) loads on commercial metrics — a 'commercial performance' summary; Component 2 (19%) on service metrics. Branch scores on Component 1 feed a league table.

How to run it

vars <- df[, c("growth","nps","footfall","conversion","basket","returns","waittime","complaints")]

library(psych)
KMO(vars); cortest.bartlett(cor(vars), n = nrow(vars))
fa.parallel(vars, fa = "pc")            # how many components

pc <- prcomp(vars, scale. = TRUE)
summary(pc)                              # variance explained
pc$rotation[, 1:2]                       # loadings
scores <- pc$x[, 1:2]                    # component scores

Interpreting the output

APA-style reporting

A principal component analysis with parallel analysis supported retaining two components, which together explained 61.2% of total variance (KMO = .78; Bartlett's χ²(28) = 612.4, p < .001). The first component (42.1%) was defined by commercial indicators (loadings .68–.84) and was saved as a commercial-performance index.

Common mistakes

Related methods

Exploratory Factor Analysis (EFA)Latent constructs instead of summariesConfirmatory Factor Analysis (CFA)Confirm a hypothesized structureMultiple Linear RegressionComponents as predictors
← Multinomial Logistic RegressionExploratory Factor Analysis (EFA) →