Clusters roughly spherical and similar in size (k-means' implicit geometry).
Outliers screened — single extreme cases can hijack centroids.
k chosen by diagnostics + interpretability, not convenience.
Hypotheses
H₀ — — (unsupervised exploration; no significance test in standard use).
H₁ — —
The concept
The algorithm alternates two steps: assign each case to its nearest centroid, then recompute centroids as cluster means, repeating until stable. It minimizes within-cluster sum of squares (WSS). Because the result depends on random starting centroids, run many starts (n_start ≥ 25) and keep the best.
Choosing k: the elbow plot (WSS by k), average silhouette width (cohesion vs separation; > .50 reasonable structure), and the gap statistic — triangulated with interpretability and actionability. Validate the solution: profile clusters on the input variables, check stability across random splits, and compare clusters on external variables NOT used in clustering (the real test of usefulness).
Worked example
A bank clusters 2,400 customers on standardized transaction frequency, average balance, digital usage, and product count. Elbow and silhouette (.46) both point to k = 4.
Segments emerge: digital-first savers (28%), traditional high-balance (22%), low-engagement (31%), credit-active (19%). Segments differ on churn (external validation) and get differentiated retention strategies.
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
X = StandardScaler().fit_transform(df[cols])
for k in range(2, 8):
km = KMeans(n_clusters=k, n_init=25, random_state=42).fit(X)
print(k, km.inertia_, silhouette_score(X, km.labels_))
km = KMeans(n_clusters=4, n_init=25, random_state=42).fit(X)
df["cluster"] = km.labels_
print(df.groupby("cluster")[cols + ["churn"]].mean())
Stability check (rerun on random halves) guards against artifacts.
External validation on variables not used in clustering demonstrates practical meaning.
Cluster labels are researcher-made names — the algorithm only made groups.
APA-style reporting
K-means clustering (z-standardized inputs; 25 random starts) with k = 4, supported by elbow and silhouette criteria (average silhouette = .46), identified four customer segments. Segments differed significantly on 12-month churn, χ²(3, N = 2,400) = 84.1, p < .001, supporting external validity of the solution.
Common mistakes
Forgetting to standardize.
Single random start (unstable solutions).
Choosing k solely by the elbow's ambiguity.
Validating clusters only on the variables that built them (circular).
Treating clusters as real fixed types rather than a useful partition.