ML (Beginner) · Start Here · cozy lesson
Your ML Setup: scikit-learn in 10 Minutes
10 min · 2 min read · no scary math, promise
One library, one pattern, everything after
scikit-learn is the standard toolkit for classical ML: datasets, dozens of models, and evaluation metrics behind a single consistent API. Learn its rhythm once — fit to train, predict to guess, score to grade — and every model in this track (and most you'll meet in jobs) works the same way.
Install and run your first model (copy this)
pip install scikit-learn pandas
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
X, y = load_iris(return_X_y=True) # 150 flowers, 4 measurements each
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=7)
model = RandomForestClassifier(n_estimators=50, random_state=7)
model.fit(Xtr, ytr)
print("accuracy:", model.score(Xte, yte)) # expect ~0.95+
That loop — load → split → fit → score — is 90% of this whole track. Run it now, in our playground or your terminal, before reading further. A running model in minute 10 beats a perfect mental model in week 3.
Why toy datasets first
Iris (150 flowers) and California housing are tiny, clean, and documented — mistakes surface as wrong numbers, not as mysterious data bugs. Graduate to messy real data only after the loop feels boring. And random_state=7 everywhere: fixed seeds make every run reproducible, so "it worked yesterday" stays true.
Remember this
fittrains,predictguesses,scoregrades. Same verbs, every model.- Toy data first, seeds fixed, loop running by minute 10.
Check your understanding
Correct answers earn XP (once each).
1. scikit-learn gives you…
2. First dataset to try?
My notes (saved in this browser)
Select text above → Save selection, or write your own. Your notebook lives in this browser.
No notes yet. Your highlights will live here.
References
Finished reading? Seal it with a tick ✅
The checkbox in the explorer turns green too — same progress.