SproutStack logoSproutStack
···

ML (Beginner) · Core Models · cozy lesson

Trees & Forests

12 min · 2 min read · no scary math, promise

🤖
You’ve got this. Read a little, play a little — I’ll wait. No rush.

Twenty questions, automated

A decision tree plays twenty questions with your data: "income above 50K?" → yes → "has mortgage?" → no → predict "will churn." Each split chosen to separate classes as cleanly as possible. Trees handle mixed data (numbers and categories), ignore feature scales, and explain themselves — you can literally draw the reasoning.

One tree memorizes; a forest generalizes

A deep tree keeps splitting until it perfectly memorizes training quirks — textbook overfitting (low bias, high variance). The fix is gloriously simple: grow hundreds of trees, each on a bootstrapped sample with random feature subsets, and let them vote. Individual mistakes average away; the signal survives. That's a random forest (bagging in one word):

from sklearn.ensemble import RandomForestClassifier
model = RandomForestClassifier(n_estimators=200, random_state=7).fit(Xtr, ytr)
print(model.score(Xte, yte))

Why forests are the default first model

No scaling needed, handles mess gracefully, rarely catastrophically wrong, and feature_importances_ tells you what actually drives predictions — free insight for stakeholders ("tenure matters 3× more than plan type"). For tabular classification, reach here before anywhere fancier.

Remember this

  • Trees ask sequential questions; forests vote across bootstrapped variants.
  • Default for tabular data; read feature importances like a business report.

Check your understanding

Correct answers earn XP (once each).

1. One deep tree usually…

2. Random forest =

My notes (saved in this browser)

Select text above → Save selection, or write your own. Your notebook lives in this browser.

No notes yet. Your highlights will live here.

Finished reading? Seal it with a tick ✅

The checkbox in the explorer turns green too — same progress.