Full-Text Search Basics
Read a little, play a little. No scary maths, and no rush.
Finding a word without reading everything
Your blog has 40 posts. "Which ones mention consistent hashing?" You can read all 40 titles. Now it has 400,000 posts and you need search — and LIKE '%hashing%' cannot help, because it has no idea where in the text the match might be, so it reads every row from the start. The database is doing a full scan wearing a disguise.
The flip that makes it fast
An inverted index stores things backwards. Instead of "this document contains these words", it stores "this word appears in these documents":
hashing → post_12, post_88, post_1043
consistent→ post_12, post_1043
raft → post_77, post_88, post_1043
Searching for "consistent hashing" is now two lookups and an intersection. That is the entire trick behind Elasticsearch, Lucene, and your database's built-in full-text index. You pay at write time — indexing every document as it changes — to make reads cheap.
Not every match is equal
Once you can find candidates, you have to rank them, because a word in a title matters more than the same word buried in paragraph nine. Classic scoring is TF-IDF: term frequency (how often it appears here) times inverse document frequency (how rare it is across everything). "the" appears everywhere, so it contributes almost nothing; "raft" is rare, so a single occurrence is strong evidence. Modern engines add field boosts, phrase proximity and recency on top.
Keep the obvious improvements first, because they are free: lowercase everything, drop stop words, and stem the rest so "running", "runs" and "run" collapse to one term. Without stemming, a search for "hashing" misses "hashes".
Where this fits in a design
Do not reach for a search cluster too early. A database's built-in full-text search (Postgres tsvector) handles millions of rows comfortably and keeps your data in one place with normal transactions. Move to a dedicated engine when you need typo tolerance, faceting, per-field relevance tuning, or the index has outgrown the primary database's write load.
The honest catch: two copies of your data means two things to keep in sync. Search indexes are eventually consistent by nature, so a document indexed a second after it was written is normal, not a bug. Design the UI to tolerate a brief delay.
Remember this
LIKE '%x%'cannot use an index; an inverted index can, because it stores words → documents.- Rank with frequency and rarity (TF-IDF), and treat title matches as stronger.
- Stemming is what makes "hashes" find "hashing".
- Start with your database's full-text search; move to a dedicated engine when it no longer fits.
Check your understanding
2 questions · correct answers earn XP once each
My notes
Saved in this browser. Highlight a line above and save it, or write it in your own words.
Nothing saved yet. Your highlights will live here.
References
Finished reading?
Ticking it here also ticks the chapter in the sidebar, the section count and your streak — it is all one number.
Related chapters
Spotted a mistake or want a topic covered? Report an issue