🏫 The School›🛠️ SRE›✅ धडा 12 — Production readiness review आणि संपूर्ण चित्र
🖼️ See the drawing + lab 🏠 Course home 🌿 Branch on GitHub ✏️ View source
🖼️ आकृती आणि labThe drawing + lab पूर्ण पानावर उघडा ↗Open full page ↗

✅ धडा 12 — Production readiness review आणि संपूर्ण चित्र

📍 तुम्ही इथे आहात: 12 पैकी धडा 12 · मागे: lesson-11-chaos-engineering


📦 या ब्रँचमध्ये काय आहे

संपूर्ण course, अधिक production readiness review (PRR): पथक एखाद्या service चा pager सांभाळायला तयार होण्याआधी तिने पार करायची तपासणी. आधीचा प्रत्येक धडा एक मुद्दा बनतो, launch थांबवणाऱ्या blockers सह आणि नंतर करता येणाऱ्या मुद्द्यांसह. sre/design.py मधील prr आणि CHECKS आणि sre/demo.py मधील prr().

🧒 5 वर्षांच्या मुलाला समजावल्यासारखे

एक नवी इमारत तयार आहे: वेळापत्रकाचे नवे office. 🏢 शिक्षिकांना सोमवारी तिथे राहायला जायचे आहे.

पथक तिची देखभाल करायला तयार होण्याआधी, कतरिना checklist घेऊन ती फिरून पाहते. दुरुस्तीच्या बजेटवर सही झाली आहे का? दिवे पटकन पुन्हा पूर्वीसारखे करता येतात का? पुढच्या सत्रासाठी, एखादी बाजू बंद झाली तरी, खुर्च्यांच्या पुरेशा रांगा आहेत का? कोणी कधी वीज खंडित करून पाहिली आहे का?

बहुतेक चौकटींवर खूण झाली आहे. पण दोन लाल आहेत. बदल उलटवायला 15 मिनिटे लागतात — खूप हळू. आणि रांगेतील तीन दारे मिळून 99.84% होतात, वचन दिलेल्या 99.9% पेक्षा कमी.

"अजून तयार नाही," कतरिना म्हणते. "कधीच नाही असे नाही — फक्त अजून नाही." शिक्षिका rollbacks साठी झटपट switch आणि office चे दुसरे दार जोडतात. त्या वीज खंडित करण्याचा सराव करतात. कतरिना पुन्हा फिरून पाहते: सगळे हिरवे. आता पथक ड्युटीचा फोन हाती घेते.

🗺️ आकृती

flowchart LR
    s["🏢 timetable service<br/>asks for SRE support"] --> prr{"✅ PRR<br/>12 checks"}
    prr -->|"❌ rollback 15 min [blocker]<br/>❌ chain 99.84% < 99.9% [blocker]<br/>❌ no game day"| nr["NOT READY<br/>2 blockers, 1 other"]
    nr -->|"flag-off rollback 1 min<br/>second database copy<br/>a game day"| prr2{"✅ PRR again"}
    prr2 --> ok["READY<br/>the crew takes the pager 📟"]

🗺️ काढलेली आकृती + एक lab: https://school-edh.pages.dev/sre/lesson-diagrams.html#l12

❓ काय

🤔 का

कारण reliability दुरुस्त करण्याची सर्वात स्वस्त वेळ म्हणजे users service वर अवलंबून होण्याच्या आधी. Checklist मुळे पथकाचा अनुभव पुन्हा वापरता येतो: प्रत्येक जुने outage एक प्रश्न बनतो ज्याचे उत्तर पुढच्या service ला द्यावे लागते. आणि ते नाते न्याय्य ठेवते: तयार नसलेल्या services भिंतीवरून फेकून देण्याची जागा म्हणजे पथक नव्हे.

🔧 कसे (या repo मध्ये)

sre/design.py मधील CHECKS ही यादी आहे: (lesson, blocker, what, test), जिथे प्रत्येक test service बद्दलच्या facts ची dictionary वाचते. prr(facts) प्रत्येक row pass की fail ते परत देते, आणि निकाल: कोणताही blocker fail झाला तर NOT READY, नाहीतर follow-ups सह READY. मुद्दा L09 dependency साखळीवर धडा 09 मधील serial() वापरतो. sre/demo.py मधील prr() timetable service चा आढावा घेते, तीन facts दुरुस्त करते आणि पुन्हा आढावा घेते.

🧪 करून पाहा

python3 sre/demo.py prr
python3 - <<'EOF'
import sys; sys.path.insert(0, "sre"); from design import prr
facts = dict(slo=0.999, policy_signed=True, ops_share=0.55, toil_ledger=False, oncall_people=4, sites=1, runbooks=True,
             canary_gate=False, rollback_min=1, flags=True, forecast_q=2, zone_safe=True, ic_trained=True,
             chain=[0.9999, 0.9995], shedding=True, retry_budget=0.1, game_day_days=30, burn_alerts=True, restore_days=None)
rows, verdict = prr(facts)
print([f"{'L' + l if l.isdigit() else l}" for ok, l, b, w in rows if not ok])
print(verdict)
EOF
python3 sre/demo.py
python3 sre/test_sre.py

✅ तपासा — तुम्हाला काय दिसायला हवे

prr हे छापते:

   ❌ L06  rollback in 5 minutes or less; risky features behind flags  [blocker]
   ✅ L07  capacity: 2-quarter forecast, survives a zone loss
   ✅ L08  incident roles trained, handoff template ready
   ❌ L09  dependency chain meets the SLO  [blocker]
   ✅ L10  load shedding by priority + a client retry budget
   ❌ L11  a game day in the last 90 days, abort conditions written
   → NOT READY — 2 blocker(s), 1 other item(s) to fix
   after fixes (flag-off rollback 1 min, two database copies, a game day): → READY — the crew takes the pager

तुमचा snippet हे छापतो:

['L01', 'L04', 'L05', 'ops']
NOT READY — 3 blocker(s), 1 other item(s) to fix

संपूर्ण demo ✅ done — the crew fixed the gutter ने संपतो, आणि tests 12/12 passed ने.

🏁 तुम्ही आत्ताच काय सिद्ध केले

प्रत्येक धड्याला एकट्याने जे सापडले असते ते आढाव्याला सापडले — 15 मिनिटांचा rollback (06) आणि गुणाकाराने 99.84% होणारी साखळी (09) — आणि त्याने त्यांचे स्पष्ट "अजून नाही" मध्ये रूपांतर केले. तीन दुरुस्त्यांनंतर त्याच आढाव्याने READY म्हटले. तुमच्या snippet ची service वेगळ्या कारणांनी fail झाली: 4 on-call लोक, canary gate नाही, आणि कधीच restore न केलेला backup हे blockers आहेत; toil ची नोंदवही नसताना 55% ops हा follow-up आहे. PRR म्हणजे गुण नव्हेत — ती gate असलेली कामांची यादी आहे.

⚠️ नेहमीच्या चुका

🏭 प्रत्यक्ष वापरात

PRR service repo मध्ये एका file म्हणून ठेवा, प्रत्येक मुद्द्याच्या पुराव्याच्या link सह:

service: timetable
reviewed: 2026-09-27
reviewers: [katrina, dipika, aishwarya]
items:
  slo_and_policy:   {status: pass, evidence: "docs/slo.yaml, signed 2026-09-01"}
  rollback_minutes: {status: pass, evidence: "game day 2026-09-20: flag off in 1 min"}
  dependency_chain: {status: pass, evidence: "99.94% ceiling with two database copies"}
  game_day:         {status: pass, evidence: "docs/gamedays/2026-09-20.md"}
  restore_test:     {status: pass, evidence: "restore drill 2026-08-18, 42 min"}
verdict: ready
next_review: 2027-03-27

Service catalogues अशा checks अनेक services मध्ये scorecards म्हणून track करू शकतात (Backstage plugins, Cortex, OpsLevel आणि तशीच tools), म्हणजे उणिवा फक्त एका इमारतीत नाही तर संपूर्ण शाळेत दिसतात.

🏭 प्रत्यक्ष वापरात हे का महत्त्वाचे आहे: CHECKS मधील बारा checks copy करा, तुमच्या team ला बसतील असे बदला, आणि या महिन्यात तुमच्या सर्वात महत्त्वाच्या service चा त्यांच्याविरुद्ध आढावा घ्या. प्रत्येक postmortem नंतर एक नवा मुद्दा जोडा.

🎓 पथकाने पन्हाळ दुरुस्त केली

आठवड्याचा जास्तीत जास्त अर्धा वेळ बादल्या → दुरुस्तीचे बजेट गळतीच्या आधी सही होते → वही ठरवते कोणत्या बादल्या आधी automate करायच्या → ड्युटीचा फोन पुरेशा लोकांमध्ये, दिवसाउजेडी वाटला जातो → नव्या boiler ची तुलना ताज्या जुन्या boiler शी होते → बदल एकेक खोली करत, हातात switch ठेवून जातात → पुढच्या सत्रासाठी आणि बंद बाजूसाठी खुर्च्या मोजल्या जातात → incident ला एक commander, एक दुरुस्ती करणारी, एक आवाज आणि एक लेखनिक असतात → रांगेतील दारे गुणाकाराने कमी करतात, प्रती nines वाढवतात → उपाहारगृह परीक्षेच्या विद्यार्थिनींना आधी जेवण देते → ठरवून केलेली वीज खंडित दावा तपासते → आणि पथक फोन केव्हा घेते ते checklist ठरवते. तुम्ही फक्त SRE शिकला नाहीत — तुम्ही कोणत्याही service कडे पाहून पथकाचे प्रश्न विचारू शकता: आपल्या आठवड्याचा किती भाग बादल्या आहेत, budget संपल्यावर काय होते, आणि ज्या अपयशासाठी आपण design केले ते आपण कधी करून पाहिले आहे का? 🛠️📟🎓

⏭️ पुढे

इतर शाळा. Observability school हा course ज्यावर अवलंबून आहे ते signals, SLO चे गणित, alerting आणि postmortems शिकवते; Scaling school traffic हाताळते; Distributed Systems school अनेक machines अवघड का असतात ते समजावते; CI/CD school आणि Argo CD school release pipeline बांधतात; School portal मध्ये बाकी सगळे आहे.

git checkout main
python3 sre/demo.py     # one last run, for fun

✅ Lesson 12 — Production readiness review and the whole picture

📍 You are here: Lesson 12 of 12 · Previous: lesson-11-chaos-engineering


📦 What's in this branch

The whole course, plus the production readiness review (PRR): the check a service passes before the crew agrees to carry its pager. Every earlier lesson becomes one item, with blockers that stop a launch and items that can follow. prr and CHECKS in sre/design.py and prr() in sre/demo.py.

🧒 Explain like I'm 5

A new building is ready: the new timetable office. 🏢 The teachers want to move in on Monday.

Before the crew agrees to look after it, Katrina walks through it with a checklist. Is there a signed repair budget? Can the lights be switched back fast? Are there enough rows of chairs for next term, even if a wing closes? Has anyone ever tried a power cut?

Most boxes are ticked. But two are red. Undoing a change takes 15 minutes — too slow. And the three doors in a row add up to 99.84%, below the promised 99.9%.

"Not ready yet," says Katrina. "Not never — just not yet." The teachers add a quick switch for rollbacks and a second office door. They run a practice power cut. Katrina walks through again: all green. Now the crew takes the duty phone.

🗺️ Diagram

flowchart LR
    s["🏢 timetable service<br/>asks for SRE support"] --> prr{"✅ PRR<br/>12 checks"}
    prr -->|"❌ rollback 15 min [blocker]<br/>❌ chain 99.84% < 99.9% [blocker]<br/>❌ no game day"| nr["NOT READY<br/>2 blockers, 1 other"]
    nr -->|"flag-off rollback 1 min<br/>second database copy<br/>a game day"| prr2{"✅ PRR again"}
    prr2 --> ok["READY<br/>the crew takes the pager 📟"]

🗺️ Drawn version + a lab: https://school-edh.pages.dev/sre/lesson-diagrams.html#l12

❓ What

🤔 Why

Because the cheapest time to fix reliability is before users depend on the service. A checklist makes the crew's experience reusable: every past outage becomes a question that the next service must answer. And it keeps the relationship fair: the crew is not a place to throw unready services over the wall.

🔧 How (in this repo)

CHECKS in sre/design.py is the list: (lesson, blocker, what, test), where each test reads a dictionary of facts about the service. prr(facts) returns each row as passed or failed, and a verdict: NOT READY if any blocker fails, otherwise READY with any follow-ups. Item L09 uses serial() from lesson 09 on the dependency chain. prr() in sre/demo.py reviews the timetable service, fixes three facts and reviews it again.

🧪 Try it

python3 sre/demo.py prr
python3 - <<'EOF'
import sys; sys.path.insert(0, "sre"); from design import prr
facts = dict(slo=0.999, policy_signed=True, ops_share=0.55, toil_ledger=False, oncall_people=4, sites=1, runbooks=True,
             canary_gate=False, rollback_min=1, flags=True, forecast_q=2, zone_safe=True, ic_trained=True,
             chain=[0.9999, 0.9995], shedding=True, retry_budget=0.1, game_day_days=30, burn_alerts=True, restore_days=None)
rows, verdict = prr(facts)
print([f"{'L' + l if l.isdigit() else l}" for ok, l, b, w in rows if not ok])
print(verdict)
EOF
python3 sre/demo.py
python3 sre/test_sre.py

✅ Verify — what you should see

prr prints:

   ❌ L06  rollback in 5 minutes or less; risky features behind flags  [blocker]
   ✅ L07  capacity: 2-quarter forecast, survives a zone loss
   ✅ L08  incident roles trained, handoff template ready
   ❌ L09  dependency chain meets the SLO  [blocker]
   ✅ L10  load shedding by priority + a client retry budget
   ❌ L11  a game day in the last 90 days, abort conditions written
   → NOT READY — 2 blocker(s), 1 other item(s) to fix
   after fixes (flag-off rollback 1 min, two database copies, a game day): → READY — the crew takes the pager

Your snippet prints:

['L01', 'L04', 'L05', 'ops']
NOT READY — 3 blocker(s), 1 other item(s) to fix

The full demo ends with ✅ done — the crew fixed the gutter, and the tests with 12/12 passed.

🏁 What you just proved

The review found what each lesson would have found alone — a 15-minute rollback (06) and a chain that multiplies to 99.84% (09) — and turned them into a clear "not yet". Three fixes later, the same review said READY. Your snippet's service failed differently: 4 on-call people, no canary gate, and a backup never restored are blockers; ops at 55% with no toil ledger is a follow-up. A PRR is not a grade — it is a to-do list with a gate.

⚠️ Common mistakes

🏭 In production

Keep the PRR as a file in the service repo, with a link to the evidence for each item:

service: timetable
reviewed: 2026-09-27
reviewers: [katrina, dipika, aishwarya]
items:
  slo_and_policy:   {status: pass, evidence: "docs/slo.yaml, signed 2026-09-01"}
  rollback_minutes: {status: pass, evidence: "game day 2026-09-20: flag off in 1 min"}
  dependency_chain: {status: pass, evidence: "99.94% ceiling with two database copies"}
  game_day:         {status: pass, evidence: "docs/gamedays/2026-09-20.md"}
  restore_test:     {status: pass, evidence: "restore drill 2026-08-18, 42 min"}
verdict: ready
next_review: 2027-03-27

Service catalogues can track such checks as scorecards across many services (Backstage plugins, Cortex, OpsLevel and similar tools), so the gaps are visible across the whole school, not just one building.

🏭 Why this matters in production: copy the twelve checks from CHECKS, change them to fit your team, and review your most important service against them this month. Add a new item after every postmortem.

🎓 The crew fixed the gutter

At most half the week is buckets → the repair budget is signed before the leaks → the notebook ranks which buckets to automate → the duty phone is shared by enough people, in daylight → a new boiler is judged against a fresh old one → changes go room by room with a switch → chairs are counted for next term and for a closed wing → an incident has a commander, a fixer, a voice and a scribe → doors in a row multiply down, copies add nines → the canteen feeds the exam pupils first → a planned power cut tests the claim → and a checklist decides when the crew takes the phone. You didn't just learn SRE — you can look at any service and ask the crew's questions: how much of our week is buckets, what happens when the budget is spent, and have we ever tried the failure we designed for? 🛠️📟🎓

⏭️ Next

The other schools. The Observability school teaches the signals, SLO arithmetic, alerting and postmortems this course leans on; the Scaling school handles the traffic; the Distributed Systems school explains why many machines are hard; the CI/CD school and the Argo CD school build the release pipeline; the School portal has the rest.

git checkout main
python3 sre/demo.py     # one last run, for fun
← Previouschaos engineeringFinished! Take the quiz →check what stuck

This page is the lesson's README from the lesson-12-production-readiness branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.