ML (Beginner) · Evaluate & Ship · cozy lesson
Clustering with K-Means 🔒 Premium
🔒 Premium chapter — free for early learners.
10 min · 2 min read · no scary math, promise
Finding groups nobody labeled
No churn column, no spam flags — just customers described by spending, visits, tenure. Clustering discovers natural groupings anyway: bargain hunters, loyal regulars, about-to-leave. It's exploration, not prediction — the output is insight ("oh, there are four kinds of users"), which then drives decisions, features, or targeted models per group.
How K-means thinks
Pick K center points. Repeat: assign every point to its nearest center, then move each center to its group's middle. Stop when centers stop moving. The result: K blobs minimizing within-group distance.
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
Xs = StandardScaler().fit_transform(X) # scale FIRST — see below
labels = KMeans(n_clusters=4, n_init=10, random_state=7).fit_predict(Xs)
The two things beginners get wrong
- Unscaled features. Income (tens of thousands) vs visits (single digits): raw distance is 99% income. Standardize first, always, for any distance-based method.
- Treating K as truth. K-means always returns K groups, even from pure noise. Try several K values, elbow-plot the inertia (the "bend" suggests a natural count), and — critically — validate clusters mean something in the business, not just the math. A segment nobody can act on is decoration.
Remember this
- No labels → cluster for structure; scale first; choose K by elbow + business sense.
- Clusters are hypotheses to validate, not answers to ship blindly.
Check your understanding
Correct answers earn XP (once each).
1. K-Means needs…
2. Scale features first because…
My notes (saved in this browser)
Select text above → Save selection, or write your own. Your notebook lives in this browser.
No notes yet. Your highlights will live here.
Finished reading? Seal it with a tick ✅
The checkbox in the explorer turns green too — same progress.