🏫 The School›🗺️ Vector Databases›✂️ धडा 06 — Chunking आणि metadata: कार्डे आणि रंगीत स्टिकर
🖼️ See the drawing + lab 🏠 Course home 🌿 Branch on GitHub ✏️ View source
🖼️ आकृती आणि labThe drawing + lab पूर्ण पानावर उघडा ↗Open full page ↗

✂️ धडा 06 — Chunking आणि metadata: कार्डे आणि रंगीत स्टिकर

📍 तुम्ही इथे आहात: 8 पैकी धडा 06 · मागे: lesson-05-build-a-vector-db · पुढे: lesson-07-rag-wiring


📦 या ब्रँचमध्ये काय आहे

धडे 01–05, आणि कोणत्याही index पेक्षा retrieval चा दर्जा जास्त बदलणारे साधेसुधे निर्णय: तुम्ही कसे कापता (chunking) आणि तुम्ही कसे label करता (metadata) — शिवाय hybrid search.

🧒 5 वर्षांच्या मुलाला समजावल्यासारखे

400 पानांची handbook तुम्ही एका खुर्चीत बसवू शकत नाही. 📚 Hall मध्ये नेण्याआधी ग्रंथपाल पुस्तकांचे index cards मध्ये तुकडे करते ✂️ — आणि कापणे ही एक कला आहे:

आणि प्रत्येक card ला रंगीत stickers 🏷️ (metadata) मिळतात: room: 3A, kind: rules, year: 2026, source: handbook-p12. Stickers मुळे दोन महाशक्ती मिळतात:

  1. Filtered search — "फक्त kind=rules मधली सर्वात जवळची cards" (आपले search(query, kind="rules") — तुम्ही demo मध्ये चालवले होते!). अर्थ उमेदवार शोधतो; stickers तथ्ये लागू करतात (tenant! permissions! — k8s namespaces ची सहजप्रवृत्ती).
  2. पावत्या — source stickers मुळेच RAG पाने सांगू (cite करू) शकते (धडा 07).

Hybrid search 🤝: अर्थ-search नेमक्या strings चुकवतो ("error E-4012", part numbers, नावे); keyword search समानार्थी शब्द चुकवतो. प्रौढ systems दोन्ही चालवतात आणि एकत्र करतात (RRF) — ग्रंथपाल card catalog पण तपासते आणि hall मध्ये फिरतेही.

🗺️ आकृती

flowchart LR
    book["📚 400-page handbook"]
    cut["✂️ chunking<br/>natural seams · self-contained ·<br/>size & overlap fit the docs (e.g. ~hundreds of tokens) 🔁"]
    cards["🗂️ index cards<br/>+ stickers 🏷️ room/kind/source"]
    hall["🗺️ the hall<br/>filtered search: stickers FIRST,<br/>then nearest-by-meaning"]
    hybrid["🤝 hybrid: + keyword catalog<br/>for exact strings (E-4012) — merge results"]
    book --> cut --> cards --> hall
    hall <--> hybrid

❓ काय

🤔 का

"RAG वाईट उत्तरे देतो" अशा दहापैकी नऊ tickets प्रत्यक्षात retrieval tickets असतात, आणि बहुतेक retrieval tickets म्हणजे chunking tickets: model ने वाईट card वरून बरोबर उत्तर दिलेले असते. खरा नॉब कुठे आहे ते आता तुम्हाला माहीत आहे — आणि तो कात्रीने धरला जातो, indexes ने नाही. ✂️

🧪 करून पाहा

python3 - <<'EOF'
import sys; sys.path.insert(0,'vectordb')
from vectordb import MiniVectorDB
db = MiniVectorDB()
# ONE fat card mixing topics vs two clean cards:
db.add("fat",  "Library books are due in two weeks. The picnic needs one pizza per three students.")
db.add("liba", "Library books are due in two weeks.", kind="rules")
db.add("picn", "The picnic needs one pizza per three students.", kind="rules")
for s, i, t, m in db.search("how many pizzas for the picnic?", k=3):
    print(f"{s:.2f}  [{i}] {t}")
EOF
# the clean picnic card outranks the fat mixed one — chunking IS ranking

⏭️ पुढे

हे सगळे प्रत्येक जण बनवतो त्या pipeline मध्ये जोडा: RAG, सुरुवातीपासून शेवटपर्यंत — आपल्याच database ला page-finder बनवून.

git checkout lesson-07-rag-wiring

✂️ Lesson 06 — Chunking & metadata: cards and colored stickers

📍 You are here: Lesson 06 of 8 · Previous: lesson-05-build-a-vector-db · Next: lesson-07-rag-wiring


📦 What's in this branch

Lessons 01–05, plus the unglamorous decisions that move retrieval quality more than any index: how you cut (chunking) and how you label (metadata) — plus hybrid search.

🧒 Explain like I'm 5

You can't seat a 400-page handbook in ONE chair. 📚 Before the hall, the librarian cuts books into index cards ✂️ — and the cutting is an art:

And every card gets colored stickers 🏷️ (metadata): room: 3A, kind: rules, year: 2026, source: handbook-p12. Stickers make two superpowers:

  1. Filtered search — "nearest cards among kind=rules only" (our search(query, kind="rules") — you ran it in the demo!). Meaning finds candidates; stickers enforce facts (tenant! permissions! — the k8s namespaces instinct).
  2. Receipts — source stickers are what lets RAG cite pages (lesson 07).

Hybrid search 🤝: meaning-search misses exact strings ("error E-4012", part numbers, names); keyword search misses synonyms. Grown-up systems run BOTH and merge (RRF) — the librarian checks the card catalog AND walks the hall.

🗺️ Diagram

flowchart LR
    book["📚 400-page handbook"]
    cut["✂️ chunking<br/>natural seams · self-contained ·<br/>size & overlap fit the docs (e.g. ~hundreds of tokens) 🔁"]
    cards["🗂️ index cards<br/>+ stickers 🏷️ room/kind/source"]
    hall["🗺️ the hall<br/>filtered search: stickers FIRST,<br/>then nearest-by-meaning"]
    hybrid["🤝 hybrid: + keyword catalog<br/>for exact strings (E-4012) — merge results"]
    book --> cut --> cards --> hall
    hall <--> hybrid

❓ What

🤔 Why

Nine out of ten "RAG gives bad answers" tickets are retrieval tickets, and most retrieval tickets are chunking tickets: the model answered correctly from a bad card. You now know where the real dial is — and that it's held with scissors, not indexes. ✂️

🧪 Try it

python3 - <<'EOF'
import sys; sys.path.insert(0,'vectordb')
from vectordb import MiniVectorDB
db = MiniVectorDB()
# ONE fat card mixing topics vs two clean cards:
db.add("fat",  "Library books are due in two weeks. The picnic needs one pizza per three students.")
db.add("liba", "Library books are due in two weeks.", kind="rules")
db.add("picn", "The picnic needs one pizza per three students.", kind="rules")
for s, i, t, m in db.search("how many pizzas for the picnic?", k=3):
    print(f"{s:.2f}  [{i}] {t}")
EOF
# the clean picnic card outranks the fat mixed one — chunking IS ranking

⏭️ Next

Wire it all into the pipeline everyone builds: RAG, end to end — with our own database as the page-finder.

git checkout lesson-07-rag-wiring
← Previousbuild a vector dbNext →rag wiring

This page is the lesson's README from the lesson-06-chunking-metadata branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.