ЁЯПл The SchoolтА║ЁЯй║ ObservabilityтА║ЁЯУЭ рдзрдбрд╛ 12 тАФ Postmortems рдЖрдгрд┐ рдЪрдХреНрд░: рд╢рд╛рд│рд╛ рдЕрдзрд┐рдХ рд╕реБрд░рдХреНрд╖рд┐рдд рдХрд░рдгрд╛рд░рд╛ рдЕрд╣рд╡рд╛рд▓
ЁЯЦ╝я╕П See the drawing + lab ЁЯПа Course home ЁЯМ┐ Branch on GitHub тЬПя╕П View source
ЁЯЦ╝я╕П рдЖрдХреГрддреА рдЖрдгрд┐ labThe drawing + lab рдкреВрд░реНрдг рдкрд╛рдирд╛рд╡рд░ рдЙрдШрдбрд╛ тЖЧOpen full page тЖЧ

ЁЯУЭ рдзрдбрд╛ 12 тАФ Postmortems рдЖрдгрд┐ рдЪрдХреНрд░: рд╢рд╛рд│рд╛ рдЕрдзрд┐рдХ рд╕реБрд░рдХреНрд╖рд┐рдд рдХрд░рдгрд╛рд░рд╛ рдЕрд╣рд╡рд╛рд▓

ЁЯУН рддреБрдореНрд╣реА рдЗрдереЗ рдЖрд╣рд╛рдд: 12 рдкреИрдХреА рдзрдбрд╛ 12 ┬╖ рдорд╛рдЧреАрд▓: lesson-11-root-cause


ЁЯУж рдпрд╛ рдмреНрд░рдБрдЪрдордзреНрдпреЗ рдХрд╛рдп рдЖрд╣реЗ

рдзрдбреЗ 01тАУ11, рдЖрдгрд┐ рдЪрдХреНрд░ рдкреВрд░реНрдг рдХрд░рдгрд╛рд░реА рдкрд╛рдпрд░реА: blameless postmortem тАФ impact, timeline, root cause рдЖрдгрд┐ contributing factors, рдЖрдгрд┐ owner рдЖрдгрд┐ date рдЕрд╕рд▓реЗрд▓реЗ action items тАФ рдЖрдгрд┐ signals рдкрд╛рд╕реВрди рдкрд░рдд рдЕрдзрд┐рдХ рдЪрд╛рдВрдЧрд▓реНрдпрд╛ signals рдкрд░реНрдпрдВрддрдЪреЗ рдЪрдХреНрд░. postmortem() obs/reliability.py рдордзреНрдпреЗ рдЖрд╣реЗ; obs/demo.py рдордзрд▓реЗ postmortem_demo() (section postmortem) рдирд┐рдХрд╛рд▓рд╛рдЪреНрдпрд╛ рджрд┐рд╡рд╕рд╛рдЪрд╛ рдЕрд╣рд╡рд╛рд▓ рд▓рд┐рд╣рд┐рддреЗ.

ЁЯзТ 5 рд╡рд░реНрд╖рд╛рдВрдЪреНрдпрд╛ рдореБрд▓рд╛рд▓рд╛ рд╕рдордЬрд╛рд╡рд▓реНрдпрд╛рд╕рд╛рд░рдЦреЗ

рд╕рд░рд╛рд╡рд╛рдирдВрддрд░ рдПрдХрд╛ рдЖрдард╡рдбреНрдпрд╛рдиреЗ, рджреАрдкрд┐рдХрд╛ рдмреИрдардХ рдмреЛрд▓рд╛рд╡рддреЗ рдЖрдгрд┐ рдПрдХ рдЕрд╣рд╡рд╛рд▓ ЁЯУЭ рд▓рд┐рд╣рд┐рддреЗ. рддреНрдпрд╛рдЪреЗ рдЪрд╛рд░ рднрд╛рдЧ рдЖрд╣реЗрдд:

рдЕрд╣рд╡рд╛рд▓рд╛рдЪреНрдпрд╛ рд╡рд░ рдореЛрдареНрдпрд╛ рдЕрдХреНрд╖рд░рд╛рдВрдд рдПрдХ рдирд┐рдпрдо рдЖрд╣реЗ: "рдЖрдореНрд╣реА рд╢рд╛рд│реЗрдХрдбреЗ рдкрд╛рд╣рддреЛ, рдПрдЦрд╛рджреНрдпрд╛ рд╡реНрдпрдХреНрддреАрдХрдбреЗ рдирд╛рд╣реА." рдХрд╛рд░рдг рдЦрд░реЗ рд╕рд╛рдВрдЧрд┐рддрд▓реНрдпрд╛рдмрджреНрджрд▓ рд▓реЛрдХ рдЕрдбрдЪрдгреАрдд рдЖрд▓реЗ, рддрд░ рдкреБрдврдЪреНрдпрд╛ рд╡реЗрд│реА рдХреЛрдгреАрдЪ рдЦрд░реЗ рд╕рд╛рдВрдЧрдгрд╛рд░ рдирд╛рд╣реА тАФ рдЖрдгрд┐ toaster рдкреБрдиреНрд╣рд╛ рдЬрд│реЗрд▓.

рдордЧ рдЕрд╣рд╡рд╛рд▓ рд╕рдЧрд│реНрдпрд╛рдВрдирд╛ рд╡рд╛рдЪрддрд╛ рдпрд╛рд╡рд╛ рдореНрд╣рдгреВрди staff-room рдЪреНрдпрд╛ рднрд┐рдВрддреАрд╡рд░ рд▓рд╛рд╡рд▓рд╛ рдЬрд╛рддреЛ. рдЖрдгрд┐ рджрд░ рд╢реБрдХреНрд░рд╡рд╛рд░реА, рджреАрдкрд┐рдХрд╛ рдпрд╛рджреА рддрдкрд╛рд╕рддреЗ: рдЭрд╛рд▓реЗ, рдХреА рдЕрдЬреВрди рдирд╛рд╣реА? рдмрджрд▓рд╛рдВрд╡рд░ рдХреБрдгрд╛рдЪреЗрдЪ рдирд╛рд╡ рдирд╕рд▓реЗрд▓рд╛ рдЕрд╣рд╡рд╛рд▓ рдореНрд╣рдгрдЬреЗ рдлрдХреНрдд рдПрдХ рдЧреЛрд╖реНрдЯ.

ЁЯЧ║я╕П рдЖрдХреГрддреА

flowchart LR
    sig["ЁЯУКЁЯУУЁЯЧ║я╕П signals<br/>L02тАУL07"] --> al["ЁЯЪи alert on the SLO<br/>L08тАУL09"]
    al --> inc["ЁЯзп incident<br/>L10"]
    inc --> rc["ЁЯФН root cause<br/>L11"]
    rc --> pm["ЁЯУЭ blameless postmortem<br/>impact ┬╖ timeline ┬╖ cause"]
    pm --> ai["тЬЕ action items<br/>owner + date"]
    ai -->|"new tests, guards, alerts"| sig

ЁЯЧ║я╕П рдХрд╛рдврд▓реЗрд▓реА рдЖрдХреГрддреА + рдПрдХ lab: https://school-edh.pages.dev/observability/lesson-diagrams.html#l12

тЭУ рдХрд╛рдп

ЁЯдФ рдХрд╛

рдХрд╛рд░рдг incident рдорд╣рд╛рдЧ рдЕрд╕рддреЗ, рдЖрдгрд┐ postmortem рд╣рд╛ рддреНрдпрд╛ рдкреИрд╢рд╛рдЪреНрдпрд╛ рдмрджрд▓реНрдпрд╛рдд рдХрд╛рд╣реАрддрд░реА рдорд┐рд│рд╡рдгреНрдпрд╛рдЪрд╛ рдорд╛рд░реНрдЧ рдЖрд╣реЗ. рддреНрдпрд╛рд╢рд┐рд╡рд╛рдп, рдзрдбреЗ рддреАрди рд▓реЛрдХрд╛рдВрдЪреНрдпрд╛ рдбреЛрдХреНрдпрд╛рдд рд░рд╛рд╣рддрд╛рдд рдЖрдгрд┐ рддреЗ рдЧреЗрд▓реЗ рдХреА рддреЗрд╣реА рдЬрд╛рддрд╛рдд. Owners рдЖрдгрд┐ dates рдирд╕рддреАрд▓, рддрд░ action items рдХреЛрдгреАрдЪ рди рд╡рд╛рдЪрдгрд╛рд░реА рдпрд╛рджреА рдмрдирддрд╛рдд, рдЖрдгрд┐ рддреЗрдЪ outage рддреАрди рдорд╣рд┐рдиреНрдпрд╛рдВрдд рдкрд░рдд рдпреЗрддреЗ. "Blameless" рдирд╕реЗрд▓, рддрд░ рд▓реЛрдХ рдереЛрдбрдХреНрдпрд╛рдд рдЯрд│рд▓реЗрд▓реНрдпрд╛ рдШрдЯрдирд╛ рд╕рд╛рдВрдЧрдгреЗ рдмрдВрдж рдХрд░рддрд╛рдд тАФ рдЬреЗ рд╕рдЧрд│реНрдпрд╛рдд рд╕реНрд╡рд╕реНрдд рдзрдбреЗ рдЕрд╕рддрд╛рдд.

ЁЯФз рдХрд╕реЗ (рдпрд╛ repo рдордзреНрдпреЗ)

obs/reliability.py рдордзрд▓реЗ postmortem(title, impact, timeline, root_cause, actions) рдПрдХ Markdown document рджреЗрддреЗ: title, blameless рдУрд│, Impact, Timeline, Root cause рдЖрдгрд┐ Action items, рдкреНрд░рддреНрдпреЗрдХ action - [owner] what (by due) рдЕрд╕реЗ рд▓рд┐рд╣рд┐рд▓реЗрд▓реЗ. obs/demo.py рдордзрд▓реЗ postmortem_demo() bad-deploy рдЪрд╛ рдЕрд╣рд╡рд╛рд▓ рд▓рд┐рд╣рд┐рддреЗ: 32 рдорд┐рдирд┐рдЯреЗ, рд╕реБрдорд╛рд░реЗ 1,900 fail рдЭрд╛рд▓реЗрд▓реЗ views, рдХрддрд░рд┐рдирд╛, рдРрд╢реНрд╡рд░реНрдпрд╛ рдЖрдгрд┐ рджреАрдкрд┐рдХрд╛ рдпрд╛рдВрдЪреНрдпрд╛рд╕рд╛рдареА рддреАрди action items. рд╢реЗрд╡рдЯреА рддреЗ рдЪрдХреНрд░ print рдХрд░рддреЗ.

ЁЯзк рдХрд░реВрди рдкрд╛рд╣рд╛

рдЬреНрдпрд╛ incident рд╕рд╛рдареА рдХреБрдгрд╛рд▓рд╛рдЪ page рдЭрд╛рд▓реЗ рдирд╛рд╣реА рддреНрдпрд╛рдЪрд╛ postmortem рд▓рд┐рд╣рд╛ тАФ рдзрдбрд╛ 11 рдиреЗ cache TTL рдмрджрд▓рд╛рд╢реА рдЬреЛрдбрд▓реЗрд▓рд╛ slow leak тАФ рдЦрд▒реНрдпрд╛ dates рд╕рд╣:

python3 obs/demo.py postmortem
python3 - <<'EOF'
import sys; sys.path.insert(0, "obs"); from reliability import postmortem
print(postmortem("slow results page from a flaky grade dependency",
                 "0.9% of results views failed for 280 minutes (minutes 100тАУ379); no page, one ticket at minute 266",
                 ["090 config: cache TTL 60 s тЖТ 5 s", "100 errors rise to 0.9%", "266 burn-rate ticket opened",
                  "380 dependency recovered; errors back to 0.1%"],
                 "a 5 s cache TTL sent 12├Ч more calls to a grade service that could not take them",
                 [("alert on the grade service's own SLO", "Aishwarya", "2026-10-02"),
                  ("config changes go through the same canary as deploys", "Dipika", "2026-10-09")]))
EOF
python3 obs/test_obs.py

тЬЕ рддрдкрд╛рд╕рд╛ тАФ рддреБрдореНрд╣рд╛рд▓рд╛ рдХрд╛рдп рджрд┐рд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ

postmortem рд╣реЗ print рдХрд░рддреЗ:

   # Postmortem: results page 504s on results day

   Blameless: we look at the system, not at a person.

   ## Impact
   6% of results views failed for 32 minutes (minutes 400тАУ431); ~1,900 parents saw an error

   ## Timeline
   - 398 deploy v42
   - 409 burn-rate page fired
   - 411 Dipika acknowledged, took command
   - 432 rollback to v41 тАФ errors back to normal
   - 470 resolved; grade cache restored

   ## Root cause
   v42 removed the grade cache; the release load test did not match results-day traffic

   ## Action items
   - [Katrina] restore the grade cache with a test that fails without it (by 2 days)
   - [Aishwarya] results-day load test in the release checklist (by 1 week)
   - [Dipika] canary 5% for 15 min with burn-rate auto-rollback (by 2 weeks)
тФАтФА the loop: signals тЖТ alerts on the SLO тЖТ incident тЖТ root cause тЖТ postmortem тЖТ action items тЖТ better signals

рддреБрдордЪрд╛ snippet рд╣реЗ print рдХрд░рддреЛ:

# Postmortem: slow results page from a flaky grade dependency

Blameless: we look at the system, not at a person.

## Impact
0.9% of results views failed for 280 minutes (minutes 100тАУ379); no page, one ticket at minute 266

## Timeline
- 090 config: cache TTL 60 s тЖТ 5 s
- 100 errors rise to 0.9%
- 266 burn-rate ticket opened
- 380 dependency recovered; errors back to 0.1%

## Root cause
a 5 s cache TTL sent 12├Ч more calls to a grade service that could not take them

## Action items
- [Aishwarya] alert on the grade service's own SLO (by 2026-10-02)
- [Dipika] config changes go through the same canary as deploys (by 2026-10-09)

Tests тЬЕ L12 the postmortem is blameless and has action items with owners рдЖрдгрд┐ 12/12 passed рдиреЗ рд╕рдВрдкрддрд╛рдд.

ЁЯПБ рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдХрд╛рдп рд╕рд┐рджреНрдз рдХреЗрд▓реЗ

рдкреНрд░рддреНрдпреЗрдХ action item рдордзреНрдпреЗ рдПрдХрд╛ рд╡реНрдпрдХреНрддреАрдЪреЗ рдирд╛рд╡ рдЖрдгрд┐ рдПрдХ deadline рдЖрд╣реЗ, рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХ item system рдмрджрд▓рддреЛ: рдПрдХ test, checklist рдордзреНрдпреЗ load test, canary, alert, config рдмрджрд▓рд╛рдВрд╕рд╛рдареА рдПрдХ рдирд┐рдпрдо. рдХреЛрдгрддрд╛рдЪ "рдХрд╛рд│рдЬреА рдШреНрдпрд╛" рдореНрд╣рдгрдд рдирд╛рд╣реА. рддреБрдордЪрд╛ рджреБрд╕рд░рд╛ рдЕрд╣рд╡рд╛рд▓ рдЪрдХреНрд░ рдЪрд╛рд▓рдд рдЕрд╕рд▓реНрдпрд╛рдЪреЗ рджрд╛рдЦрд╡рддреЛ: slow leak рдиреЗ рдХрдзреАрдЪ рдХреБрдгрд╛рд▓рд╛ рдЙрдард╡рд▓реЗ рдирд╛рд╣реА (рдзрдбрд╛ 08 рдиреЗ рддреНрдпрд╛рдЪреЗ рдХрд╛рдо рдХреЗрд▓реЗ), рддрд░реАрд╣реА рддреНрдпрд╛рдЪреА рдХрд┐рдВрдордд 2,520 fail рдЭрд╛рд▓реЗрд▓реЗ views рд╣реЛрддреА тАФ bad deploy рдЪреНрдпрд╛ ~1,920 рдкреЗрдХреНрд╖рд╛ рдЬрд╛рд╕реНрдд тАФ рдореНрд╣рдгреВрди рддреНрдпрд╛рдЪрд╛рд╣реА postmortem рд╡реНрд╣рд╛рдпрд▓рд╛ рд╣рд╡рд╛, рдЖрдгрд┐ рддреНрдпрд╛рдЪреЗ fixes рдЕрдзрд┐рдХ рдЪрд╛рдВрдЧрд▓реЗ signals рдмрдирддрд╛рдд: рдПрдХ рдирд╡рд╛ SLO рдЖрдгрд┐ config рд╕рд╛рдареА canary. 2026-10-02 рд╕рд╛рд░рдЦреА рддрд╛рд░реАрдЦ рддрдкрд╛рд╕рддрд╛ рдпреЗрддреЗ; demo рдЪреНрдпрд╛ рдЕрд╣рд╡рд╛рд▓рд╛рддрд▓реЗ "2 days" рдлрдХреНрдд рддреЛ рд▓рд┐рд╣рд┐рд▓рд╛ рддреНрдпрд╛ рджрд┐рд╡рд╢реАрдЪ рд╕реНрдкрд╖реНрдЯ рдЕрд╕рддреЗ.

тЪая╕П рдиреЗрд╣рдореАрдЪреНрдпрд╛ рдЪреБрдХрд╛

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд

рдЦрд▒реНрдпрд╛ account рд╡рд░ тАФ рдПрдХ postmortem template (Markdown, wiki рдХрд┐рдВрд╡рд╛ repo рд╕рд╛рдареА):

# Postmortem: <title>  ┬╖  SEV<1-3>  ┬╖  <date>  ┬╖  status: draft | reviewed | closed
Blameless: we look at the system, not at a person.

## Summary
<2тАУ3 sentences: what users saw, for how long, what fixed it>

## Impact
- Users affected: <number / share>   ┬╖ Duration: <start тЖТ mitigated>
- Error budget used: <% of the 30-day budget>   ┬╖ SLA breached: yes / no

## Timeline (all times UTC)
| time | event |
|---|---|
| 06:08 | deploy results-api v42 |
| 06:10 | errors start (SLI) |
| 06:19 | burn-rate page fires тАФ detected |
| 06:21 | on-call acknowledges, becomes IC тАФ acknowledged |
| 06:42 | rollback to v41 тАФ mitigated |
| 07:20 | cache restored, SLI normal тАФ resolved |

## Root cause and contributing factors
<five whys from lesson 11; list every factor that lined up>

## What went well ┬╖ what went badly ┬╖ where we got lucky

## Action items
| type | action | owner | due | ticket |
|---|---|---|---|---|
| prevent | test that fails when the grade cache is removed | Katrina | 2026-09-29 | RES-101 |
| prevent | results-day load test in the release checklist | Aishwarya | 2026-10-04 | RES-102 |
| mitigate | canary 5% for 15 min with burn-rate auto-rollback | Dipika | 2026-10-11 | RES-103 |

Team рдЬрд┐рдереЗ рдХрд╛рдо track рдХрд░рддреЗ рддрд┐рдереЗ action items рддрдпрд╛рд░ рдХрд░рд╛, рдореНрд╣рдгрдЬреЗ рддреЗ planning рдордзреНрдпреЗ рджрд┐рд╕рддреАрд▓ (GitHub CLI):

gh issue create --title "Test that fails when the grade cache is removed" \
    --body "Postmortem: results page 504s on results day (SEV1). Due 2026-09-29." \
    --assignee katrina --label postmortem,prevent
gh issue list --label postmortem --state open        # review this list every week

PagerDuty, incident.io, FireHydrant, Rootly рдЖрдгрд┐ Jira Service Management рд╕рд╛рд░рдЦреА tools incident channel рдордзреВрди timeline рддрдпрд╛рд░ рдХрд░рддрд╛рдд рдЖрдгрд┐ follow-ups track рдХрд░рддрд╛рдд.

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд рд╣реЗ рдХрд╛ рдорд╣рддреНрддреНрд╡рд╛рдЪреЗ: рдлрдХреНрдд outages рдирд╡реНрд╣реЗ, рдЪрдХреНрд░ рдореЛрдЬрд╛: рдХрд┐рддреА action items рдЙрдШрдбреЗ рдЖрд╣реЗрдд, рдХрд┐рддреА рдЙрд╢реАрд░ рдЭрд╛рд▓реЗрд▓реЗ рдЖрд╣реЗрдд, рдЖрдгрд┐ рддреЗрдЪ cause рдХрд┐рддреА рд╡реЗрд│рд╛ рдкрд░рдд рдпреЗрддреЗ. рдЖрдкрд▓реЗ items рдмрдВрдж рдХрд░рдгрд╛рд▒реНрдпрд╛ team рдЪреНрдпрд╛ рд░рд╛рддреНрд░реА рджрд░ рддрд┐рдорд╛рд╣реАрдд рдЕрдзрд┐рдХ рд╢рд╛рдВрдд рд╣реЛрддрд╛рдд.

ЁЯОУ рдЖрд░реЛрдЧреНрдп рдХрдХреНрд╖ рдЖрддрд╛ рддреБрдордЪрд╛ рдЖрд╣реЗ

рдЖрд░реЛрдЧреНрдп рдХрдХреНрд╖ рдХрд╛ тЖТ рджреИрдирдВрджрд┐рдиреА тЖТ рддрдкрд╛рд╕рдгреА рддрдХреНрддрд╛ тЖТ рд╣рд│реВ рдореНрд╣рдгрдЬреЗ рдХрд┐рддреА рд╣рд│реВ тЖТ рдПрдХрд╛ рд╡рд┐рджреНрдпрд╛рд░реНрдереНрдпрд╛рдЪреЗ рдорд╛рд░реНрдЧ рдХрд╛рд░реНрдб тЖТ рдкреНрд░рддреНрдпреЗрдХ рдЦреЛрд▓реАрд╕рд╛рдареА рдПрдХрдЪ form тЖТ scrape рдХрд░рд╛, рд╡рд┐рдЪрд╛рд░рд╛, рдХрд╛рдврд╛ тЖТ рд╡рд┐рджреНрдпрд╛рд░реНрдереНрдпрд╛рдВрдирд╛ рддреНрд░рд╛рд╕ рд╣реЛрдд рдЕрд╕реЗрд▓ рддреЗрд╡реНрд╣рд╛рдЪ рдЙрдард╡рдгрд╛рд░реА рдзреЛрдХреНрдпрд╛рдЪреА рдШрдВрдЯрд╛ тЖТ рдЖрдЬрд╛рд░реА-рд░рдЬреЗрдЪреА рдореБрднрд╛ тЖТ рдЕрдЧреНрдирд┐рд╢рдорди рд╕рд░рд╛рд╡ тЖТ рдХрд╛, рдкрд╛рдЪ рд╡реЗрд│рд╛ тЖТ рдирд╛рд╡реЗ рдЖрдгрд┐ рддрд╛рд░рдЦрд╛рдВрд╕рд╣ рдЕрд╣рд╡рд╛рд▓. рддреБрдореНрд╣реА рдлрдХреНрдд observability рд╢рд┐рдХрд▓рд╛ рдирд╛рд╣реАрдд тАФ system рдХрд╛рдп рдХрд░рдд рдЖрд╣реЗ рдЖрдгрд┐ рдХрд╛, рд╣реЗ рддреБрдореНрд╣реА рд╕рд╛рдВрдЧреВ рд╢рдХрддрд╛, рдЖрдгрд┐ рдкреБрдврдЪреЗ incident рд▓рд╣рд╛рди рдХрд░реВ рд╢рдХрддрд╛. ЁЯй║ЁЯУКЁЯОУ

тПня╕П рдкреБрдвреЗ

рдЗрддрд░ рд╢рд╛рд│рд╛. System Design рд╢рд╛рд│рд╛ рдкреНрд░рддреНрдпреЗрдХ box рдордзреНрдпреЗ observability рдЖрдзреАрдЪ рдард░рд╡рддреЗ; Scaling рд╢рд╛рд│рд╛ load рдЦрд╛рд▓реА рдХрд╛рдп рдкрд╛рд╣рд╛рдпрдЪреЗ рддреЗ рджрд╛рдЦрд╡рддреЗ; Kubernetes рд╢рд╛рд│рд╛ рддреЗ pods рдЪрд╛рд▓рд╡рддреЗ рдЬреНрдпрд╛рдВрд╡рд░ рдЖрддрд╛ рддреБрдореНрд╣реА рд▓рдХреНрд╖ рдареЗрд╡реВ рд╢рдХрддрд╛; School portal рдордзреНрдпреЗ рдмрд╛рдХреА рд╕рдЧрд│реЗ рдЖрд╣реЗ.

git checkout main
python3 obs/demo.py     # one last run, for fun

ЁЯУЭ Lesson 12 тАФ Postmortems & the loop: the report that makes the school safer

ЁЯУН You are here: Lesson 12 of 12 ┬╖ Previous: lesson-11-root-cause


ЁЯУж What's in this branch

Lessons 01тАУ11, plus the step that closes the loop: the blameless postmortem тАФ impact, timeline, root cause and contributing factors, and action items with an owner and a date тАФ and the loop from signals back to better signals. postmortem() lives in obs/reliability.py; postmortem_demo() (section postmortem) in obs/demo.py writes the results-day report.

ЁЯзТ Explain like I'm 5

A week after the drill, Dipika calls a meeting and writes a report ЁЯУЭ. It has four parts:

There is one rule at the top of the report, in big letters: "We look at the school, not at a person." Because if people get into trouble for telling the truth, next time nobody tells the truth тАФ and the toaster burns again.

Then the report goes on the staff-room wall for everyone to read. And every Friday, Dipika checks the list: done, or not yet? A report with nobody's name on the changes is just a story.

ЁЯЧ║я╕П Diagram

flowchart LR
    sig["ЁЯУКЁЯУУЁЯЧ║я╕П signals<br/>L02тАУL07"] --> al["ЁЯЪи alert on the SLO<br/>L08тАУL09"]
    al --> inc["ЁЯзп incident<br/>L10"]
    inc --> rc["ЁЯФН root cause<br/>L11"]
    rc --> pm["ЁЯУЭ blameless postmortem<br/>impact ┬╖ timeline ┬╖ cause"]
    pm --> ai["тЬЕ action items<br/>owner + date"]
    ai -->|"new tests, guards, alerts"| sig

ЁЯЧ║я╕П Drawn version + a lab: https://school-edh.pages.dev/observability/lesson-diagrams.html#l12

тЭУ What

ЁЯдФ Why

Because an incident is expensive, and the postmortem is how you get something for the money. Without it, the lessons live in three people's heads and leave when they do. Without owners and dates, action items become a list nobody reads, and the same outage returns in three months. Without "blameless", people stop reporting near misses тАФ the cheapest lessons of all.

ЁЯФз How (in this repo)

postmortem(title, impact, timeline, root_cause, actions) in obs/reliability.py returns a Markdown document: a title, the blameless line, Impact, Timeline, Root cause and Action items, with each action written as - [owner] what (by due). postmortem_demo() in obs/demo.py writes the bad-deploy report: 32 minutes, about 1,900 failed views, three action items for Katrina, Aishwarya and Dipika. It ends by printing the loop.

ЁЯзк Try it

Write the postmortem for the incident nobody was paged for тАФ the slow leak that lesson 11 tied to the cache TTL change тАФ with real dates:

python3 obs/demo.py postmortem
python3 - <<'EOF'
import sys; sys.path.insert(0, "obs"); from reliability import postmortem
print(postmortem("slow results page from a flaky grade dependency",
                 "0.9% of results views failed for 280 minutes (minutes 100тАУ379); no page, one ticket at minute 266",
                 ["090 config: cache TTL 60 s тЖТ 5 s", "100 errors rise to 0.9%", "266 burn-rate ticket opened",
                  "380 dependency recovered; errors back to 0.1%"],
                 "a 5 s cache TTL sent 12├Ч more calls to a grade service that could not take them",
                 [("alert on the grade service's own SLO", "Aishwarya", "2026-10-02"),
                  ("config changes go through the same canary as deploys", "Dipika", "2026-10-09")]))
EOF
python3 obs/test_obs.py

тЬЕ Verify тАФ what you should see

postmortem prints:

   # Postmortem: results page 504s on results day

   Blameless: we look at the system, not at a person.

   ## Impact
   6% of results views failed for 32 minutes (minutes 400тАУ431); ~1,900 parents saw an error

   ## Timeline
   - 398 deploy v42
   - 409 burn-rate page fired
   - 411 Dipika acknowledged, took command
   - 432 rollback to v41 тАФ errors back to normal
   - 470 resolved; grade cache restored

   ## Root cause
   v42 removed the grade cache; the release load test did not match results-day traffic

   ## Action items
   - [Katrina] restore the grade cache with a test that fails without it (by 2 days)
   - [Aishwarya] results-day load test in the release checklist (by 1 week)
   - [Dipika] canary 5% for 15 min with burn-rate auto-rollback (by 2 weeks)
тФАтФА the loop: signals тЖТ alerts on the SLO тЖТ incident тЖТ root cause тЖТ postmortem тЖТ action items тЖТ better signals

Your snippet prints:

# Postmortem: slow results page from a flaky grade dependency

Blameless: we look at the system, not at a person.

## Impact
0.9% of results views failed for 280 minutes (minutes 100тАУ379); no page, one ticket at minute 266

## Timeline
- 090 config: cache TTL 60 s тЖТ 5 s
- 100 errors rise to 0.9%
- 266 burn-rate ticket opened
- 380 dependency recovered; errors back to 0.1%

## Root cause
a 5 s cache TTL sent 12├Ч more calls to a grade service that could not take them

## Action items
- [Aishwarya] alert on the grade service's own SLO (by 2026-10-02)
- [Dipika] config changes go through the same canary as deploys (by 2026-10-09)

The tests end with тЬЕ L12 the postmortem is blameless and has action items with owners and 12/12 passed.

ЁЯПБ What you just proved

Every action item names one person and a deadline, and each one changes the system: a test, a load test in the checklist, a canary, an alert, a rule for config changes. None says "be careful". Your second report shows the loop working: the slow leak never woke anyone (lesson 08 did its job), yet it still cost 2,520 failed views тАФ more than the bad deploy's ~1,920 тАФ so it earns a postmortem too, and its fixes become better signals: a new SLO and a canary for config. A date like 2026-10-02 can be checked; "2 days" in the demo's report is only clear on the day it was written.

тЪая╕П Common mistakes

ЁЯПн In production

On a real account тАФ a postmortem template (Markdown, for a wiki or a repo):

# Postmortem: <title>  ┬╖  SEV<1-3>  ┬╖  <date>  ┬╖  status: draft | reviewed | closed
Blameless: we look at the system, not at a person.

## Summary
<2тАУ3 sentences: what users saw, for how long, what fixed it>

## Impact
- Users affected: <number / share>   ┬╖ Duration: <start тЖТ mitigated>
- Error budget used: <% of the 30-day budget>   ┬╖ SLA breached: yes / no

## Timeline (all times UTC)
| time | event |
|---|---|
| 06:08 | deploy results-api v42 |
| 06:10 | errors start (SLI) |
| 06:19 | burn-rate page fires тАФ detected |
| 06:21 | on-call acknowledges, becomes IC тАФ acknowledged |
| 06:42 | rollback to v41 тАФ mitigated |
| 07:20 | cache restored, SLI normal тАФ resolved |

## Root cause and contributing factors
<five whys from lesson 11; list every factor that lined up>

## What went well ┬╖ what went badly ┬╖ where we got lucky

## Action items
| type | action | owner | due | ticket |
|---|---|---|---|---|
| prevent | test that fails when the grade cache is removed | Katrina | 2026-09-29 | RES-101 |
| prevent | results-day load test in the release checklist | Aishwarya | 2026-10-04 | RES-102 |
| mitigate | canary 5% for 15 min with burn-rate auto-rollback | Dipika | 2026-10-11 | RES-103 |

Create the action items where the team tracks work, so they show up in planning (GitHub CLI):

gh issue create --title "Test that fails when the grade cache is removed" \
    --body "Postmortem: results page 504s on results day (SEV1). Due 2026-09-29." \
    --assignee katrina --label postmortem,prevent
gh issue list --label postmortem --state open        # review this list every week

Tools such as PagerDuty, incident.io, FireHydrant, Rootly and Jira Service Management build the timeline from the incident channel and track the follow-ups.

ЁЯПн Why this matters in production: measure the loop, not only the outages: how many action items are open, how many are overdue, and how often the same cause comes back. A team that closes its items gets quieter nights every quarter.

ЁЯОУ The health room is yours

Why a health room тЖТ the diary тЖТ the vital-signs chart тЖТ how slow is slow тЖТ one pupil's route card тЖТ one form for every room тЖТ scrape, ask, draw тЖТ an alarm that wakes you only when pupils are hurting тЖТ a sick-day allowance тЖТ the fire drill тЖТ why, five times тЖТ the report, with names and dates. You didn't just learn observability тАФ you can tell what the system is doing, why, and make the next incident shorter. ЁЯй║ЁЯУКЁЯОУ

тПня╕П Next

The other schools. The System Design school plans observability into every box; the Scaling school shows what to watch under load; the Kubernetes school runs the pods you now know how to watch; the School portal has the rest.

git checkout main
python3 obs/demo.py     # one last run, for fun
тЖР Previousroot causeFinished! Take the quiz тЖТcheck what stuck

This page is the lesson's README from the lesson-12-postmortems branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.