ML (Beginner) · Core Models · cozy lesson
Regression Intuition
10 min · 2 min read · no scary math, promise
Predicting numbers, starting with lines
House prices, delivery demand, customer tenure — whenever the answer is a number, you're doing regression. The simplest version draws the best line through scattered points: given square footage, the line guesses price. A line in one dimension becomes a plane in two, a hyperplane beyond — same idea, more slopes.
from sklearn.linear_model import LinearRegression
model = LinearRegression().fit(Xtr, ytr)
print("R²:", model.score(Xte, yte))
Under the hood it minimizes squared error — big misses hurt quadratically, so the line chases the bulk of the points while outliers tug it around (a preview of why data cleaning matters).
Reading R² honestly
R² = 1.0 means perfect prediction; 0 means "as good as always guessing the average." Real models live between: 0.85 on house prices is excellent, 0.3 on stock returns might still be tradeable. Context sets the bar, never the number alone. And note what's not the metric: accuracy % belongs to categories — applying it to numbers is a category error beginners make once.
When lines aren't enough
Curved relationships (price vs distance isn't a straight line) defeat linear models — the classic underfit. That's your cue for trees and forests (next chapters), which bend with the data. Start linear anyway: it's fast, interpretable ("each bedroom adds $40K"), and the baseline every fancier model must beat.
Remember this
- Numbers out → regression. R² near 1 good, near 0 baseline.
- Linear first for speed and interpretability; trees when lines underfit.
Check your understanding
Correct answers earn XP (once each).
1. Regression predicts…
2. Fit quality in one number?
My notes (saved in this browser)
Select text above → Save selection, or write your own. Your notebook lives in this browser.
No notes yet. Your highlights will live here.
Finished reading? Seal it with a tick ✅
The checkbox in the explorer turns green too — same progress.