🌱 AI Engineering · Embeddings, Vector Search & RAG · cozy lesson
Chunking Strategies
10 min · 1 min read · no scary math, promise
🤖
You’ve got this. Read a little, play a little — I’ll wait. No rush.
Why chunking decides RAG quality
Too big → fuzzy, token-hungry. Too small → context lost. Play /animations/chunking first.
Recipe
- Split by headings, then by tokens (not chars).
- 300–800 tokens, overlap 10–20%.
- Keep
{source, section, page, updated}metadata. - Tables/code: keep whole, don’t slice mid-row.
def chunk(text, n=500, overlap=80):
words = text.split()
out, i = [], 0
while i < len(words):
out.append(" ".join(words[i:i+n]))
i += n - overlap
return out
💛 Enjoying? Try 5 playful quizzes or watch it move.
Check your understanding
Correct answers earn XP (once each).
1. Sweet spot for most docs?
2. Keep with each chunk?
My notes (saved in this browser)
Select text above → Save selection, or write your own. AlgoMaster-style notebook, local-first for MVP.
No notes yet. Your highlights will live here.
Finished reading? Seal it with a tick ✅
The checkbox in the explorer turns green too — same progress.