🏫 The School›🧠 AI›✏️ धडा 02 — Training: सराव परीक्षा आणि लाल पेन
🖼️ See the drawing + lab 🏠 Course home 🌿 Branch on GitHub ✏️ View source
🖼️ आकृती आणि labThe drawing + lab पूर्ण पानावर उघडा ↗Open full page ↗

✏️ धडा 02 — Training: सराव परीक्षा आणि लाल पेन

📍 तुम्ही इथे आहात: 12 पैकी धडा 02 · मागे: lesson-01-what-is-ai · पुढे: lesson-03-tokens


📦 या ब्रँचमध्ये काय आहे

धडा 01, आणि सगळ्या machine learning चे इंजिन: loss (मी किती चुकलो?) आणि gradient descent (कमी चूक कोणत्या दिशेला?).

🧒 5 वर्षांच्या मुलाला समजावल्यासारखे

मूल गणितात खरोखर सुधारते कसे? जादूने नाही — सराव-परीक्षेच्या loop ने ✏️:

  1. सराव परीक्षा द्या — model उदाहरणांवर predictions करतो.
  2. लाल पेन 🔴 — उत्तरे उत्तरपत्रिकेशी जुळवा आणि किती चुकलात ते मोजा. हा चुकीचा-गुण म्हणजे loss. (फक्त बरोबर/चूक नाही — किती चूक: 9×7=64 हे उत्तर 9×7=12 पेक्षा जवळ आहे.)
  3. थोडेसे बदला — ही आहे सुंदर युक्ती: model च्या लाखो dials (weights) पैकी प्रत्येकासाठी calculus विचारू शकते "मी हा dial थोडासा फिरवला, तर loss कमी होईल का?" मग प्रत्येक dial ला त्याच्या स्वतःच्या उताराच्या दिशेने छोटासा धक्का मिळतो. हेच gradient descent — "चुकीचा डोंगर एकेक छोटे पाऊल टाकत उतरणे." ⛰️
  4. पुन्हा करा. लाखो वेळा. loss हळूहळू कमी होतो, मूल अधिक तल्लख होते.

वर्गातील दोन अपयश जे तुम्हाला माहीत असलेच पाहिजेत:

🗺️ आकृती

flowchart LR
    d["📚 examples<br/>question + right answer"]
    m["🧠 model<br/>millions of dials"]
    p["✍️ prediction"]
    l["🔴 loss<br/>HOW wrong?"]
    g["⛰️ gradient descent<br/>nudge every dial<br/>slightly downhill"]
    d --> m --> p --> l -->|"1"| g -->|"2 repeat millions of times"| m
    v["🧪 validation set - unseen questions<br/>catches memorizers (overfitting)"]
    v -.->|"3 honest grade"| l

❓ काय

🤔 का

AI ची प्रत्येक क्षमता आणि प्रत्येक खर्चाची गोष्ट म्हणजे हाच loop. "एक model train करायला लाखो खर्च येतो" = मोठ्या प्रमाणावर step 3 साठी लागणारी वीज. "model biased आहे" = उत्तरपत्रिका (data) biased होती. "Fine-tuning" (धडा 07) = तोच लाल पेन, एका छोट्या, खास परीक्षेवर. सगळ्यांवर राज्य करणारा एकच loop.

🧪 करून पाहा

python3 demo/bigram_model.py

आपले हे खेळणे calculus वगळते — मोजण्यासाठी, मोजलेले आकडेच परिपूर्ण dials असतात (loss कमीत कमी करणारे, सिद्ध करता येते!). आता कल्पना करा की छापलेल्या तक्त्यात 175 billion dials आहेत जे मोजून सोडवता येत नाहीत — फक्त उताराकडे ढकलता येतात, उदाहरणामागून उदाहरण. तोच तक्ता, अवघड डोंगर. ⛰️

कागदावरचा सराव: तीन धक्क्यांमध्ये model 9×7 चे उत्तर 64, मग 61, मग 63 देतो. loss वाढतोय की कमी होतोय? (कमी — लाल पेन अंतर मोजतो, आणि म्हणूनच "किती चूक" हे "चूक" पेक्षा सरस आहे.)

⏭️ पुढे

LLM text वर सराव करण्याआधी, text ला calculator पचवू शकेल असे काहीतरी बनावे लागते: tokens — puzzle चे तुकडे.

git checkout lesson-03-tokens

✏️ Lesson 02 — Training: practice tests and the red pen

📍 You are here: Lesson 02 of 12 · Previous: lesson-01-what-is-ai · Next: lesson-03-tokens


📦 What's in this branch

Lesson 01, plus the engine of all machine learning: loss (how wrong am I?) and gradient descent (which way is less wrong?).

🧒 Explain like I'm 5

How does a kid actually get better at math? Not by magic — by the practice-test loop ✏️:

  1. Take a practice test — the model makes predictions on examples.
  2. The red pen 🔴 — compare answers to the answer key and count how wrong you were. That wrongness-score is the loss. (Not just right/wrong — HOW wrong: answering 9×7=64 is closer than 9×7=12.)
  3. Adjust, a tiny bit — here's the beautiful trick: for every one of the model's millions of dials (weights), calculus can ask "if I turned THIS dial slightly, would the loss go down?" Then every dial gets a tiny nudge in its own downhill direction. That's gradient descent — "descend the mountain of wrongness one small step at a time." ⛰️
  4. Repeat. Millions of times. Loss drifts down, the kid gets sharper.

Two classroom failure modes you must know:

🗺️ Diagram

flowchart LR
    d["📚 examples<br/>question + right answer"]
    m["🧠 model<br/>millions of dials"]
    p["✍️ prediction"]
    l["🔴 loss<br/>HOW wrong?"]
    g["⛰️ gradient descent<br/>nudge every dial<br/>slightly downhill"]
    d --> m --> p --> l -->|"1"| g -->|"2 repeat millions of times"| m
    v["🧪 validation set - unseen questions<br/>catches memorizers (overfitting)"]
    v -.->|"3 honest grade"| l

❓ What

🤔 Why

Every AI capability and every AI cost story is this loop. "Training a model costs millions" = electricity for step 3 at scale. "The model is biased" = the answer key (data) was. "Fine-tuning" (lesson 07) = the same red pen on a smaller, special test. One loop to rule them all.

🧪 Try it

python3 demo/bigram_model.py

Our toy skips calculus — for counting, counts ARE the perfect dials (loss-minimizing, provably!). Now imagine the printed table with 175 billion dials that can't be solved by counting — only nudged downhill, example by example. Same table, harder mountain. ⛰️

Paper exercise: the model answers 9×7 with 64 then 61 then 63 across three nudges. Is the loss going up or down? (Down — the red pen counts distance, and that's why "how wrong" beats "wrong".)

⏭️ Next

Before an LLM can practice on text, text must become something a calculator can chew: tokens — the puzzle pieces.

git checkout lesson-03-tokens
← Previouswhat is aiNext →tokens

This page is the lesson's README from the lesson-02-training branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.