ML (Beginner) · Data First · cozy lesson
Features & Labels
9 min · 2 min read · no scary math, promise
Every column gets a job
Open any dataset and each column must be assigned a role. Features (X) are what the model sees at prediction time: tenure, monthly bill, support tickets, email text. The label (y) is the answer to learn: churned or not, spam or ham, price. This split sounds trivial and causes half of all beginner disasters — because a column in the wrong role poisons everything downstream.
Features must exist at prediction time
The golden rule: a feature is legitimate only if you'll actually have its value when predicting. cancel_date when predicting cancellation, refund_issued when predicting refunds — these are leakage: the answer smuggled into the inputs. Models with leaked features post miraculous 99% scores and fail completely in production, where the smuggled column doesn't exist yet.
Choosing features is the actual skill
Domain sense beats algorithms here. For churn: declining usage matters more than raw usage (trends over snapshots), support-ticket sentiment beats ticket counts, tenure interacts with everything. Two habits of strong practitioners: engineer a few smart features before trying a fancier model, and drop features that are proxies for the label in disguise.
Remember this
- Features = inputs available at prediction time. Label = the answer. Never confuse them.
- "Would I know this when predicting?" — ask it about every column.
Check your understanding
Correct answers earn XP (once each).
1. Label vs feature?
2. Leakage means…
My notes (saved in this browser)
Select text above → Save selection, or write your own. Your notebook lives in this browser.
No notes yet. Your highlights will live here.
Finished reading? Seal it with a tick ✅
The checkbox in the explorer turns green too — same progress.