Clustering
Clustering groups similar data points without predefined labels — an unsupervised technique for discovering structure.
Definition
Clustering assigns each point to a cluster so that intra-cluster points are more similar than inter-cluster points.
No "true" labels are used — the algorithm defines what a cluster is based on the metric and linkage criterion.
Intuition
The right number of clusters is often ambiguous; use domain knowledge, elbow plots (total within-cluster variance vs. ), or silhouette scores.
Clustering is sensitive to scale — always standardize features (zero mean, unit variance) before clustering, or features with large ranges dominate distance.
Worked example
Customer segmentation: group shoppers by purchase behavior to target marketing differently for each group.
Document clustering groups news articles by topic without predefined categories.
The math
K-means minimizes within-cluster sum of squares (WCSS): .
Hierarchical agglomerative clustering builds a dendrogram bottom-up, stopping when the desired number of clusters is reached.
In practice
Used for market segmentation, image compression, anomaly detection (points far from any cluster center), and exploratory data analysis.
DBSCAN is an alternative that finds clusters of arbitrary shape and labels others as noise.
Go deeper
- InteractiveK-Means Clustering Playground
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.