Statistical Methods Atlas
Atlas › Classify & Group › Hierarchical Cluster Analysis

Hierarchical Cluster Analysis

Builds a tree (dendrogram) of nested clusters without pre-specifying their number — cut the tree where the structure makes sense.

Classify & GroupMultivariate also known as: Agglomerative clustering, Ward's method

✓ When to use

  • You do not know how many clusters to expect and want the dendrogram to inform the choice.
  • Small-to-medium samples (the distance matrix grows with n²).
  • Nested structure is itself interesting (families of segments within segments).
  • Often paired with k-means: hierarchical to pick k and seed centroids, k-means to refine.

✗ When NOT to use

  • Very large datasets — distance matrices explode; use k-means or mini-batch variants.
  • You need cases to be re-assignable — agglomerative merges are irreversible (a case fused early stays put).
  • Mixed categorical data without an appropriate distance (use Gower distance if so).

Data requirements

Dependent / outcome variableContinuous (standardized) clustering variables, or any data with a defensible distance metric.
Independent / grouping variable—
DesignOne sample.
Sample size guidanceComfortable up to a few thousand cases; beyond that, sample or switch methods.

Assumptions

Hypotheses

H₀ — — (unsupervised exploration).
H₁ — —

The concept

Agglomerative clustering starts with every case as its own cluster and repeatedly merges the two closest clusters until one remains, recording each merge height. The dendrogram displays this history; cutting it at a chosen height yields a partition, and large vertical gaps between merges suggest natural cut points.

'Closest' depends on linkage: Ward merges the pair whose fusion least increases within-cluster variance — the natural companion to k-means' objective. Validate cuts with silhouette widths and, as with k-means, profile and externally validate the resulting groups. The dendrogram's shape is also diagnostic: long thin chains (single linkage) or one giant cluster plus crumbs signal weak structure.

Worked example

An HR team clusters 180 job roles on standardized skill-requirement scores (analytical, interpersonal, technical, physical). Ward linkage; the dendrogram shows a clear 3-cluster gap; silhouette = .52.

Clusters: knowledge roles, service roles, operational roles — used to design three distinct training tracks.

How to run it

vars <- scale(df[, c("analytical", "interpersonal", "technical", "physical")])

d  <- dist(vars, method = "euclidean")
hc <- hclust(d, method = "ward.D2")
plot(hc, labels = FALSE)                 # dendrogram

library(factoextra)
fviz_nbclust(vars, FUN = hcut, method = "silhouette")

groups <- cutree(hc, k = 3)
table(groups)
aggregate(df[, 2:5], list(groups), mean)  # profiles

Interpreting the output

APA-style reporting

Hierarchical cluster analysis (Ward's method, squared Euclidean distance, z-standardized inputs) suggested a three-cluster solution, supported by the agglomeration schedule and average silhouette width (.52). The clusters represented knowledge (n = 61), service (n = 67), and operational (n = 52) role families.

Common mistakes

Related methods

K-Means ClusteringPre-set k, scalablePrincipal Component Analysis (PCA)Visualize/reduce before clusteringDiscriminant Analysis (LDA)Explain resulting groups
← K-Means ClusteringDiscriminant Analysis (LDA) →