ML (Beginner) · Core Models · cozy lesson
Trees & Forests
12 min · 2 min read · no scary math, promise
Twenty questions, automated
A decision tree plays twenty questions with your data: "income above 50K?" → yes → "has mortgage?" → no → predict "will churn." Each split chosen to separate classes as cleanly as possible. Trees handle mixed data (numbers and categories), ignore feature scales, and explain themselves — you can literally draw the reasoning.
One tree memorizes; a forest generalizes
A deep tree keeps splitting until it perfectly memorizes training quirks — textbook overfitting (low bias, high variance). The fix is gloriously simple: grow hundreds of trees, each on a bootstrapped sample with random feature subsets, and let them vote. Individual mistakes average away; the signal survives. That's a random forest (bagging in one word):
from sklearn.ensemble import RandomForestClassifier
model = RandomForestClassifier(n_estimators=200, random_state=7).fit(Xtr, ytr)
print(model.score(Xte, yte))
Why forests are the default first model
No scaling needed, handles mess gracefully, rarely catastrophically wrong, and feature_importances_ tells you what actually drives predictions — free insight for stakeholders ("tenure matters 3× more than plan type"). For tabular classification, reach here before anywhere fancier.
Remember this
- Trees ask sequential questions; forests vote across bootstrapped variants.
- Default for tabular data; read feature importances like a business report.
Check your understanding
Correct answers earn XP (once each).
1. One deep tree usually…
2. Random forest =
My notes (saved in this browser)
Select text above → Save selection, or write your own. Your notebook lives in this browser.
No notes yet. Your highlights will live here.
Finished reading? Seal it with a tick ✅
The checkbox in the explorer turns green too — same progress.