Statistical Methods Atlas
Atlas › Classify & Group › K-Means Clustering

K-Means Clustering

Partitions cases into k groups so that members are similar within and different between clusters — e.g., segmenting customers by behavior.

Classify & GroupMultivariate also known as: Partitioning clustering

✓ When to use

  • You suspect natural segments (customer types, employee profiles) but have no labels.
  • Continuous (or standardized) clustering variables and a willingness to pre-specify k.
  • Large samples where hierarchical methods get unwieldy.

✗ When NOT to use

  • Group membership is already known and you want to predict it — discriminant analysis or logistic regression.
  • Heavily categorical data — use k-modes/k-prototypes or latent class analysis.
  • Clusters expected to be non-spherical, of very different sizes/densities — consider hierarchical or model-based (Gaussian mixture) clustering.
  • Unstandardized variables on different scales — the biggest-variance variable will dominate.

Data requirements

Dependent / outcome variableA set of continuous clustering variables, standardized; no outcome variable.
Independent / grouping variable—
DesignOne sample; variables chosen by theory (what should define the segments?).
Sample size guidancen comfortably larger than k × variables; results stabilize with hundreds of cases.

Assumptions

Hypotheses

H₀ — — (unsupervised exploration; no significance test in standard use).
H₁ — —

The concept

The algorithm alternates two steps: assign each case to its nearest centroid, then recompute centroids as cluster means, repeating until stable. It minimizes within-cluster sum of squares (WSS). Because the result depends on random starting centroids, run many starts (n_start ≥ 25) and keep the best.

Choosing k: the elbow plot (WSS by k), average silhouette width (cohesion vs separation; > .50 reasonable structure), and the gap statistic — triangulated with interpretability and actionability. Validate the solution: profile clusters on the input variables, check stability across random splits, and compare clusters on external variables NOT used in clustering (the real test of usefulness).

Worked example

A bank clusters 2,400 customers on standardized transaction frequency, average balance, digital usage, and product count. Elbow and silhouette (.46) both point to k = 4.

Segments emerge: digital-first savers (28%), traditional high-balance (22%), low-engagement (31%), credit-active (19%). Segments differ on churn (external validation) and get differentiated retention strategies.

How to run it

vars <- scale(df[, c("freq", "balance", "digital", "products")])

library(factoextra)
fviz_nbclust(vars, kmeans, method = "wss")        # elbow
fviz_nbclust(vars, kmeans, method = "silhouette") # silhouette

set.seed(42)
km <- kmeans(vars, centers = 4, nstart = 25)
km$size; km$centers                                # profile
fviz_cluster(km, data = vars)
aggregate(df$churn, list(km$cluster), mean)        # external validation

Interpreting the output

APA-style reporting

K-means clustering (z-standardized inputs; 25 random starts) with k = 4, supported by elbow and silhouette criteria (average silhouette = .46), identified four customer segments. Segments differed significantly on 12-month churn, χ²(3, N = 2,400) = 84.1, p < .001, supporting external validity of the solution.

Common mistakes

Related methods

Hierarchical Cluster AnalysisNo pre-set k; dendrogramDiscriminant Analysis (LDA)Predict known groupsPrincipal Component Analysis (PCA)Reduce dimensions before clustering
← Intraclass Correlation Coefficient (ICC)Hierarchical Cluster Analysis →