SproutStack logoSproutStack
···

ML (Beginner) · Data First · cozy lesson

Features & Labels

9 min · 2 min read · no scary math, promise

🤖
You’ve got this. Read a little, play a little — I’ll wait. No rush.

Every column gets a job

Open any dataset and each column must be assigned a role. Features (X) are what the model sees at prediction time: tenure, monthly bill, support tickets, email text. The label (y) is the answer to learn: churned or not, spam or ham, price. This split sounds trivial and causes half of all beginner disasters — because a column in the wrong role poisons everything downstream.

Features must exist at prediction time

The golden rule: a feature is legitimate only if you'll actually have its value when predicting. cancel_date when predicting cancellation, refund_issued when predicting refunds — these are leakage: the answer smuggled into the inputs. Models with leaked features post miraculous 99% scores and fail completely in production, where the smuggled column doesn't exist yet.

Choosing features is the actual skill

Domain sense beats algorithms here. For churn: declining usage matters more than raw usage (trends over snapshots), support-ticket sentiment beats ticket counts, tenure interacts with everything. Two habits of strong practitioners: engineer a few smart features before trying a fancier model, and drop features that are proxies for the label in disguise.

Remember this

  • Features = inputs available at prediction time. Label = the answer. Never confuse them.
  • "Would I know this when predicting?" — ask it about every column.

Check your understanding

Correct answers earn XP (once each).

1. Label vs feature?

2. Leakage means…

My notes (saved in this browser)

Select text above → Save selection, or write your own. Your notebook lives in this browser.

No notes yet. Your highlights will live here.

Finished reading? Seal it with a tick ✅

The checkbox in the explorer turns green too — same progress.