Builds a tree (dendrogram) of nested clusters without pre-specifying their number — cut the tree where the structure makes sense.
Classify & GroupMultivariatealso known as: Agglomerative clustering, Ward's method
✓ When to use
You do not know how many clusters to expect and want the dendrogram to inform the choice.
Small-to-medium samples (the distance matrix grows with n²).
Nested structure is itself interesting (families of segments within segments).
Often paired with k-means: hierarchical to pick k and seed centroids, k-means to refine.
✗ When NOT to use
Very large datasets — distance matrices explode; use k-means or mini-batch variants.
You need cases to be re-assignable — agglomerative merges are irreversible (a case fused early stays put).
Mixed categorical data without an appropriate distance (use Gower distance if so).
Data requirements
Dependent / outcome variable
Continuous (standardized) clustering variables, or any data with a defensible distance metric.
Independent / grouping variable
—
Design
One sample.
Sample size guidance
Comfortable up to a few thousand cases; beyond that, sample or switch methods.
Assumptions
A meaningful distance metric (Euclidean for standardized continuous data; Gower for mixed).
Linkage choice fits the goal: Ward (compact, similar-size clusters — most common in research), complete, average, single (chaining-prone).
Standardization; outlier screening.
Hypotheses
H₀ — — (unsupervised exploration).
H₁ — —
The concept
Agglomerative clustering starts with every case as its own cluster and repeatedly merges the two closest clusters until one remains, recording each merge height. The dendrogram displays this history; cutting it at a chosen height yields a partition, and large vertical gaps between merges suggest natural cut points.
'Closest' depends on linkage: Ward merges the pair whose fusion least increases within-cluster variance — the natural companion to k-means' objective. Validate cuts with silhouette widths and, as with k-means, profile and externally validate the resulting groups. The dendrogram's shape is also diagnostic: long thin chains (single linkage) or one giant cluster plus crumbs signal weak structure.
Worked example
An HR team clusters 180 job roles on standardized skill-requirement scores (analytical, interpersonal, technical, physical). Ward linkage; the dendrogram shows a clear 3-cluster gap; silhouette = .52.
Clusters: knowledge roles, service roles, operational roles — used to design three distinct training tracks.
How to run it
vars <- scale(df[, c("analytical", "interpersonal", "technical", "physical")])
d <- dist(vars, method = "euclidean")
hc <- hclust(d, method = "ward.D2")
plot(hc, labels = FALSE) # dendrogram
library(factoextra)
fviz_nbclust(vars, FUN = hcut, method = "silhouette")
groups <- cutree(hc, k = 3)
table(groups)
aggregate(df[, 2:5], list(groups), mean) # profiles
from scipy.cluster.hierarchy import linkage, dendrogram, fcluster
from sklearn.preprocessing import StandardScaler
import matplotlib.pyplot as plt
X = StandardScaler().fit_transform(df[cols])
Z = linkage(X, method="ward")
dendrogram(Z); plt.show()
df["cluster"] = fcluster(Z, t=3, criterion="maxclust")
print(df.groupby("cluster")[cols].mean())
Analyze → Classify → Hierarchical Cluster.
Move standardized variables in; Method: Ward's method, Squared Euclidean distance; under Transform Values choose Z-scores if not pre-standardized.
Plots: tick Dendrogram.
Statistics: Agglomeration schedule (large jumps in coefficients suggest the cut).
Save: Cluster membership for your chosen number(s); then profile with Compare Means.
Not feasible natively beyond toy examples (n² distance matrix + iterative merging).
Cluster in R/Python/SPSS; bring the cluster labels into Excel for pivots and charts.
Interpreting the output
Dendrogram: cut where merge heights jump; report linkage and distance used.
Silhouette/agglomeration schedule to support the chosen k.
Cluster profiles in original units; sizes per cluster.
External validation as for k-means.
Solutions are descriptive tools — different linkages can tell different stories; disclose sensitivity.
APA-style reporting
Hierarchical cluster analysis (Ward's method, squared Euclidean distance, z-standardized inputs) suggested a three-cluster solution, supported by the agglomeration schedule and average silhouette width (.52). The clusters represented knowledge (n = 61), service (n = 67), and operational (n = 52) role families.
Common mistakes
Ward linkage with a non-Euclidean distance (contradiction).
Reading cluster count off an unstable dendrogram without silhouette support.
Forgetting standardization.
Treating the first solution as final without linkage sensitivity checks.