🏫 The School›🧠 AI›⭐ धडा 07 — RLHF आणि LoRA: ग्रंथालय, शिष्टाचार वर्ग, सोनेरी तारे, चिकट चिठ्ठ्या
🖼️ See the drawing + lab 🏠 Course home 🌿 Branch on GitHub ✏️ View source
🖼️ आकृती आणि labThe drawing + lab पूर्ण पानावर उघडा ↗Open full page ↗

⭐ धडा 07 — RLHF आणि LoRA: ग्रंथालय, शिष्टाचार वर्ग, सोनेरी तारे, चिकट चिठ्ठ्या

📍 तुम्ही इथे आहात: 12 पैकी धडा 07 — भाग 1 संपला! · मागे: lesson-06-attention · पुढे: lesson-08-prompting-context


📦 या ब्रँचमध्ये काय आहे

धडे 01–06, आणि अंतिम संस्काराची शाळा: कच्चा text-predictor मदत करणारा assistant कसा बनतो — pretraining → fine-tuning → RLHF — आणि सगळे वापरतात ती बजेटची युक्ती: LoRA.

🧒 5 वर्षांच्या मुलाला समजावल्यासारखे

assistant घडवायला चार शाळा लागतात:

  1. 📚 संपूर्ण ग्रंथालय वाचा (pretraining): वर्षानुवर्षे ग्रंथालयात कोंडून सगळे वाचणे — हुशार, अब्जावधी गोष्टी माहीत, पण "bread कसा bake करायचा?" विचारले तर ते कदाचित उत्तर देईल "हा प्रश्न अनेक नवशिके विचारतात." ते text पूर्ण करते; ते तुम्हाला उत्तर देत नाही. (महिने, लाखो dollars, एकदाच केलेले.)
  2. 🎓 शिष्टाचार वर्ग (supervised fine-tuning): त्याला "प्रश्न → उपयुक्त उत्तर" ची हजारो उदाहरणे दाखवा, जोपर्यंत उत्तर देण्याची शैली मुरत नाही. तोच लाल पेन (धडा 02), छोटा खास अभ्यासक्रम. आता ते उत्तर देते!
  3. ⭐ सोनेरी तारे (RLHF — reinforcement learning from human feedback): ते दोन उत्तरे लिहिते; माणसे त्यातले चांगले निवडतात — हजारो वेळा. त्या निवडींमधून एक taste-judge model (reward model) train करा, मग विद्यार्थ्याला असे उत्तर लिहिण्याकडे ढकला ज्याला judge जास्त गुण देतो. इथेच helpful, honest, harmless हे जुळवले जाते — आणि स्वभाव इथूनच येतो. (हे जास्त केले की तुम्हाला sycophancy मिळते: शिक्षकांना जे ऐकायचे तेच बोलणारे मूल. 😬)
  4. 🗒️ चिकट चिठ्ठ्या (LoRA): assistant ने तुमच्या कंपनीच्या शैलीतही बोलावे असे हवे? अब्जावधी dials पुन्हा train करणे म्हणजे मुलाला नवा मेंदू विकत घेणे. LoRA त्याऐवजी मेंदू गोठवते आणि छोट्या जोड चिठ्ठ्या शिकते (model च्या parameters चा अगदी छोटा भाग) ज्यांच्या दुरुस्त्या वरून चालतात. संपूर्ण मेंदूपेक्षा train करायला खूप स्वस्त, बदलता येणाऱ्या (एक model, अनेक चिठ्ठी-संच), काढता येणाऱ्या.

🗺️ आकृती

flowchart LR
    p["📚 pretraining<br/>read everything<br/>months · $$$M · once"]
    sft["🎓 fine-tuning<br/>Q→A examples<br/>learn to ANSWER"]
    rlhf["⭐ RLHF<br/>humans pick better answer →<br/>taste-judge → nudge student"]
    chat["🤖 the assistant<br/>you actually talk to"]
    lora["🗒️ LoRA - sticky notes<br/>frozen brain + tiny add-on<br/>~0.1% params · swappable"]
    p -->|"1"| sft -->|"2"| rlhf -->|"3"| chat
    lora -.->|"4 personalize cheaply"| chat

❓ काय

🤔 का

कच्चा model आणि ChatGPT-सारखे assistants यांच्यात तुम्हाला जाणवणारा फरक हा धडा समजावतो — आणि "ते इतके होकारार्थी का आहे" हा स्वभाव नसून training चा परिणाम का आहे तेही. कामावरच्या #1 व्यावहारिक प्रश्नासाठीही तो तुम्हाला तयार करतो: "आपण fine-tune करावे का?" — tuning काय जोडू शकते आणि काय नाही (style हो, facts नाही), आणि चिकट चिठ्ठ्यांसह व त्यांशिवाय त्याचा खर्च किती, हे आता तुम्हाला माहीत आहे.

🧪 करून पाहा (कागदावरचा lab — मजेदार प्रकारचा)

शिष्टाचार शिक्षिका बना: "Why is the sky blue?" या प्रश्नासाठी लिहा (a) ग्रंथालयातले मूल देईल असे completion ("…is a common question in physics forums. Related: why are sunsets red?") आणि (b) तुम्हाला खरोखर हवे असलेले assistant चे उत्तर. आता तिसरे उत्तर तयार करा जे facts च्या दृष्टीने बरोबर पण उद्धट आहे — आणि लक्षात घ्या की RLHF चे काम नेमके बाकी दोन्हींपेक्षा (b) ला प्राधान्य देणे आहे: अचूकता AND वागणूक. तुम्ही आत्ताच सोनेरी-तारे pipeline हाताने चालवून पाहिली. ⭐

⏭️ पुढे — भाग 2 सुरू 🧰

model बांधून तयार झाला आहे. आता कोर्सचा दुसरा अर्धा भाग: तो नीट वापरणे — सुरुवात त्या कौशल्याने जे सगळे असल्याचा दावा करतात आणि फार थोडे सराव करतात: prompting, आणि ते सगळे ज्यावर मावते ते छोटे desk.

git checkout lesson-08-prompting-context

⭐ Lesson 07 — RLHF & LoRA: library, etiquette class, gold stars, sticky notes

📍 You are here: Lesson 07 of 12 — end of Part 1! · Previous: lesson-06-attention · Next: lesson-08-prompting-context


📦 What's in this branch

Lessons 01–06, plus the finishing school: how a raw text-predictor becomes a helpful assistant — pretraining → fine-tuning → RLHF — and the budget trick everyone uses: LoRA.

🧒 Explain like I'm 5

Raising an assistant takes four schools:

  1. 📚 Read the whole library (pretraining): years locked in the library reading everything — brilliant, knows a billion things, but ask "how do I bake bread?" and it might reply "is a question many beginners ask." It completes text; it doesn't answer you. (Months, millions of dollars, done once.)
  2. 🎓 Etiquette class (supervised fine-tuning): show it thousands of examples of "question → helpful answer" until the answering style sinks in. Same red pen (lesson 02), tiny special curriculum. Now it answers!
  3. ⭐ Gold stars (RLHF — reinforcement learning from human feedback): it writes two answers; humans pick the better one — thousands of times. From those picks, train a taste-judge model (the reward model), then nudge the student to write answers the judge scores high. This is where helpful, honest, harmless gets dialed in — and where the personality comes from. (Over-dial it and you get sycophancy: the kid who says what teachers want to hear. 😬)
  4. 🗒️ Sticky notes (LoRA): want the assistant to also speak YOUR company's tone? Retraining billions of dials is buying the kid a new brain. LoRA instead freezes the brain and learns tiny add-on note-sheets (a tiny fraction of the model's parameters) whose corrections ride on top. Far cheaper to train than the whole brain, swappable (one model, many note-sets), removable.

🗺️ Diagram

flowchart LR
    p["📚 pretraining<br/>read everything<br/>months · $$$M · once"]
    sft["🎓 fine-tuning<br/>Q→A examples<br/>learn to ANSWER"]
    rlhf["⭐ RLHF<br/>humans pick better answer →<br/>taste-judge → nudge student"]
    chat["🤖 the assistant<br/>you actually talk to"]
    lora["🗒️ LoRA - sticky notes<br/>frozen brain + tiny add-on<br/>~0.1% params · swappable"]
    p -->|"1"| sft -->|"2"| rlhf -->|"3"| chat
    lora -.->|"4 personalize cheaply"| chat

❓ What

🤔 Why

This lesson explains the difference you FEEL between a raw model and ChatGPT-style assistants — and why "it's so agreeable" is a training artifact, not a personality. It also arms you for the #1 practical question at work: "should we fine-tune?" — now you know what tuning can and cannot add (style yes, facts no), and what it costs with vs without sticky notes.

🧪 Try it (paper lab — the fun kind)

Play etiquette teacher: for the question "Why is the sky blue?" write (a) the completion a library kid might produce ("…is a common question in physics forums. Related: why are sunsets red?") and (b) the assistant answer you'd actually want. Now invent the third answer that's factually right but rude — and notice RLHF's job is exactly to prefer (b) over BOTH others: correctness AND manner. You've just hand-simulated the gold-star pipeline. ⭐

⏭️ Next — Part 2 begins 🧰

The model is built and finished. Now the OTHER half of the course: using it well — starting with the skill everyone claims and few practice: prompting, and the small desk it all fits on.

git checkout lesson-08-prompting-context
← PreviousattentionNext →prompting context

This page is the lesson's README from the lesson-07-rlhf-lora branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.