SproutStack logoSproutStack
···

ML (Beginner) · Evaluate & Ship · cozy lesson

Overfitting & How to Fix It

11 min · 2 min read · no scary math, promise

🤖
You’ve got this. Read a little, play a little — I’ll wait. No rush.

The student who memorized the textbook

A student memorizes past papers word-for-word: 100% on those exact questions, blank stare at anything new. That's overfitting — the model learned training quirks instead of patterns. The opposite failure, underfitting (bad on train and validation), means the model is too simple to capture even the training signal.

Reading the curves

Plot train and validation score as complexity grows (deeper trees, more epochs). Healthy: both rise together, small gap. Overfitting: train keeps climbing while validation stalls then falls — the divergence point is where learning stopped and memorizing began. This picture diagnoses more ML problems than any other single image.

The fix ladder (climb in order)

  1. More data — the honest cure; quirks drown in volume.
  2. Simpler model — shallower trees, fewer features. Capacity must match data.
  3. Regularization — penalize complexity (C in logistic regression, max_depth in trees). Gentle pressure toward simpler explanations.
  4. Cross-validation — rotate which slice validates (cross_val_score) so your number isn't a lucky split.
  5. Early stopping — halt training the moment validation worsens.

Remember this

  • Gap between train and validation is the diagnosis; the ladder is the prescription.
  • Complexity must be earned by data — constrain first, enlarge only with evidence.

Check your understanding

Correct answers earn XP (once each).

1. Classic overfit signal?

2. First fixes?

My notes (saved in this browser)

Select text above → Save selection, or write your own. Your notebook lives in this browser.

No notes yet. Your highlights will live here.

Finished reading? Seal it with a tick ✅

The checkbox in the explorer turns green too — same progress.