SproutStack logoSproutStack
···

🌱 AI Engineering · Embeddings, Vector Search & RAG · cozy lesson

Chunking Strategies

10 min · 1 min read · no scary math, promise

🤖
You’ve got this. Read a little, play a little — I’ll wait. No rush.

Why chunking decides RAG quality

Too big → fuzzy, token-hungry. Too small → context lost. Play /animations/chunking first.

Recipe

  • Split by headings, then by tokens (not chars).
  • 300–800 tokens, overlap 10–20%.
  • Keep {source, section, page, updated} metadata.
  • Tables/code: keep whole, don’t slice mid-row.
def chunk(text, n=500, overlap=80):
    words = text.split()
    out, i = [], 0
    while i < len(words):
        out.append(" ".join(words[i:i+n]))
        i += n - overlap
    return out

Check your understanding

Correct answers earn XP (once each).

1. Sweet spot for most docs?

2. Keep with each chunk?

My notes (saved in this browser)

Select text above → Save selection, or write your own. AlgoMaster-style notebook, local-first for MVP.

No notes yet. Your highlights will live here.

Finished reading? Seal it with a tick ✅

The checkbox in the explorer turns green too — same progress.