ЁЯПл The SchoolтА║ЁЯй║ ObservabilityтА║ЁЯФН рдзрдбрд╛ 11 тАФ Root-cause analysis: рд╣реЗ рдЦрд░реЛрдЦрд░ рдХрд╛ рдШрдбрд▓реЗ
ЁЯЦ╝я╕П See the drawing + lab ЁЯПа Course home ЁЯМ┐ Branch on GitHub тЬПя╕П View source
ЁЯЦ╝я╕П рдЖрдХреГрддреА рдЖрдгрд┐ labThe drawing + lab рдкреВрд░реНрдг рдкрд╛рдирд╛рд╡рд░ рдЙрдШрдбрд╛ тЖЧOpen full page тЖЧ

ЁЯФН рдзрдбрд╛ 11 тАФ Root-cause analysis: рд╣реЗ рдЦрд░реЛрдЦрд░ рдХрд╛ рдШрдбрд▓реЗ

ЁЯУН рддреБрдореНрд╣реА рдЗрдереЗ рдЖрд╣рд╛рдд: 12 рдкреИрдХреА рдзрдбрд╛ 11 ┬╖ рдорд╛рдЧреАрд▓: lesson-10-incident-response ┬╖ рдкреБрдвреАрд▓: lesson-12-postmortems


ЁЯУж рдпрд╛ рдмреНрд░рдБрдЪрдордзреНрдпреЗ рдХрд╛рдп рдЖрд╣реЗ

рдзрдбреЗ 01тАУ10, рдЖрдгрд┐ рдЦрд░реЗ рдХрд╛рд░рдг рд╢реЛрдзрдгреЗ: change correlation (рдЕрдЧрджреА рдЖрдзреА рдХрд╛рдп рдмрджрд▓рд▓реЗ?), рдкрд╛рдЪ рдХрд╛, рдЖрдгрд┐ рдЙрддреНрддрд░ рдЬрд╡рд│рдЬрд╡рд│ рдиреЗрд╣рдореАрдЪ рд╡реНрдпрдХреНрддреА рдирд╡реНрд╣реЗ рддрд░ process рдордзрд▓реЗ рдирд╕рд▓реЗрд▓реЗ рд╕рдВрд░рдХреНрд╖рдг рдХрд╛ рдЕрд╕рддреЗ. first_bad_minute(), suspect_changes() рдЖрдгрд┐ five_whys() obs/reliability.py рдордзреНрдпреЗ рдЖрд╣реЗрдд; obs/demo.py рдордзрд▓реЗ rootcause() рддреНрдпрд╛рдВрдирд╛ рдирд┐рдХрд╛рд▓рд╛рдЪреНрдпрд╛ рджрд┐рд╡рд╕рд╛рд╡рд░ рдЪрд╛рд▓рд╡рддреЗ.

ЁЯзТ 5 рд╡рд░реНрд╖рд╛рдВрдЪреНрдпрд╛ рдореБрд▓рд╛рд▓рд╛ рд╕рдордЬрд╛рд╡рд▓реНрдпрд╛рд╕рд╛рд░рдЦреЗ

рд╕рд░рд╛рд╡рд╛рдирдВрддрд░ рд╕рдЧрд│реЗ рд╕реБрд░рдХреНрд╖рд┐рдд рдЖрд╣реЗрдд. рдЖрддрд╛ рдХрддрд░рд┐рдирд╛ рд╡рд┐рдЪрд╛рд░рддреЗ: рд╣реЗ рдХрд╛ рдШрдбрд▓реЗ?

рдЖрдзреА рддреА рдШрдбреНрдпрд╛рд│ рдЖрдгрд┐ рднреЗрдЯ рдиреЛрдВрджрд╡рд╣реА рдкрд╛рд╣рддреЗ. рдзреВрд░ 11:40 рд▓рд╛ рд╕реБрд░реВ рдЭрд╛рд▓рд╛. 11:40 рдЪреНрдпрд╛ рдЕрдЧрджреА рдЖрдзреА рдХреЛрдг рдЖрдд рдЖрд▓реЗ рдХрд┐рдВрд╡рд╛ рдХреЛрдгреА рдХрд╛рд╣реАрддрд░реА рдмрджрд▓рд▓реЗ? рдиреЛрдВрджрд╡рд╣реА рд╕рд╛рдВрдЧрддреЗ: 11:38 рд▓рд╛, staff room рдордзреНрдпреЗ рдПрдХ рдирд╡рд╛ toaster рд▓рд╛рд╡рд▓рд╛ рдЧреЗрд▓рд╛. рддреЛ рдкрд╣рд┐рд▓рд╛ рд╕рдВрд╢рдпрд┐рдд.

рдордЧ рддреА рд╡рд┐рдЪрд╛рд░рддреЗ "рдХрд╛?" тАФ рдЖрдгрд┐ рдкреБрдиреНрд╣рд╛ рд╡рд┐рдЪрд╛рд░рддреЗ, рдкреБрдиреНрд╣рд╛ рдкреБрдиреНрд╣рд╛:

  1. рдзреВрд░ рдХрд╛ рдЖрд▓рд╛? Toaster рдиреЗ рдкрд╛рд╡ рдЬрд╛рд│рд▓рд╛.
  2. рддреЛ рдХрд╛ рдЬрд│рд▓рд╛? Toaster 20 рдорд┐рдирд┐рдЯреЗ рдЪрд╛рд▓реВ рд░рд╛рд╣рд┐рд▓рд╛.
  3. рддреЛ рдЪрд╛рд▓реВ рдХрд╛ рд░рд╛рд╣рд┐рд▓рд╛? рддреНрдпрд╛рд▓рд╛ рдмрдВрдж рдХрд░рдгрд╛рд░рд╛ timer рдирд╛рд╣реА.
  4. Timer рдирд╕рд▓реЗрд▓реНрдпрд╛ toaster рд▓рд╛ рдкрд░рд╡рд╛рдирдЧреА рдХрд╛ рдорд┐рд│рд╛рд▓реА? рд╕реНрд╡рдпрдВрдкрд╛рдХрдШрд░рд╛рдЪреНрдпрд╛ рдирд╡реНрдпрд╛ рд╡рд╕реНрддреВ рдХреЛрдгреАрдЪ рддрдкрд╛рд╕рдд рдирд╛рд╣реА.
  5. рдХреЛрдгреАрдЪ рдХрд╛ рддрдкрд╛рд╕рдд рдирд╛рд╣реА? рд╕реНрд╡рдпрдВрдкрд╛рдХрдШрд░рд╛рдЪреНрдпрд╛ рдирд╡реНрдпрд╛ рд╡рд╕реНрддреВрдВрд╕рд╛рдареА рдХреЛрдгрддреАрд╣реА checklist рдирд╛рд╣реА.

рдХрддрд░рд┐рдирд╛ "рдХреБрдгреАрддрд░реА рддреЛ рдЪрд╛рд▓реВ рдареЗрд╡рд▓рд╛" рдЗрдереЗ рдерд╛рдВрдмрдд рдирд╛рд╣реА. рдорд╛рдгрд╕реЗ рдЧреЛрд╖реНрдЯреА рд╡рд┐рд╕рд░рддрд╛рдд; рддреЗ рдЙрджреНрдпрд╛рд╣реА рдкреБрдиреНрд╣рд╛ рдШрдбреЗрд▓. Timer рдЖрдгрд┐ checklist рд╣реЗ рд╕рдЧрд│реНрдпрд╛рдВрд╕рд╛рдареА рджреБрд░реБрд╕реНрдд рдХрд░рддрд╛рдд.

ЁЯЧ║я╕П рдЖрдХреГрддреА

flowchart TB
    spike["ЁЯУИ error ratio first above 5% at minute 400"]
    ch["ЁЯЧВя╕П changes in the 30 minutes before<br/>398 deploy results-api v42 тЬЕ suspect<br/>(90 config ┬╖ 250 v41 ┬╖ 430 rollback тАФ outside)"]
    spike --> ch
    ch --> w1["why 1? parents saw 504 on /results"]
    w1 --> w2["why 2? the grade service call took 1 s and timed out"]
    w2 --> w3["why 3? v42 removed the grade cache"]
    w3 --> w4["why 4? the load test used 3 classes, not 900"]
    w4 --> w5["why 5? no production-like load test in the release checklist"]
    w5 --> fix["ЁЯФз fix the PROCESS, not the person"]

ЁЯЧ║я╕П рдХрд╛рдврд▓реЗрд▓реА рдЖрдХреГрддреА + рдПрдХ lab: https://school-edh.pages.dev/observability/lesson-diagrams.html#l11

тЭУ рдХрд╛рдп

ЁЯдФ рдХрд╛

рдХрд╛рд░рдг рддреБрдореНрд╣реА рдирд┐рд╡рдбрд▓реЗрд▓рд╛ fix рддреБрдореНрд╣реА рдХрд┐рддреА рдЦреЛрд▓ рдЬрд╛рддрд╛ рдпрд╛рд╡рд░ рдЕрд╡рд▓рдВрдмреВрди рдЕрд╕рддреЛ. "v42 рд╡рд╛рдИрдЯ рд╣реЛрддреЗ" рдЗрдереЗ рдерд╛рдВрдмрд▓рд╛рдд рддрд░ рддреБрдореНрд╣реА roll back рдХрд░рддрд╛ рдЖрдгрд┐ v43 рдиреЗ рддреЗрдЪ рдХрд░рдгреНрдпрд╛рдЪреА рд╡рд╛рдЯ рдкрд╛рд╣рддрд╛. "engineer cache рд╡рд┐рд╕рд░рд▓рд╛" рдЗрдереЗ рдерд╛рдВрдмрд▓рд╛рдд рддрд░ рддреБрдореНрд╣реА рдПрдХрд╛ рд╡реНрдпрдХреНрддреАрд▓рд╛ рджреЛрд╖ рджреЗрддрд╛, рд▓реЛрдХ рдкреБрдврдЪреА рдЪреВрдХ рд▓рдкрд╡рддрд╛рдд, рдЖрдгрд┐ рдХрд╛рд╣реАрдЪ рдмрджрд▓рдд рдирд╛рд╣реА. "release checklist рдордзреНрдпреЗ production рд╕рд╛рд░рдЦреА load test рдирд╛рд╣реА" рдЗрдердкрд░реНрдпрдВрдд рдкреЛрд╣реЛрдЪрд╛ рдЖрдгрд┐ рддреБрдореНрд╣реА рдЕрд╢реА рдкреЛрдХрд│реА рднрд░рддрд╛ рдЬрд┐рдиреЗ рд╣рд╛ deploy рдкрдХрдбрд▓рд╛ рдЕрд╕рддрд╛ тАФ рдЖрдгрд┐ рдкреБрдврдЪреЗ рдкрд╛рдЪрд╣реА.

ЁЯФз рдХрд╕реЗ (рдпрд╛ repo рдордзреНрдпреЗ)

obs/reliability.py рдордзрд▓реЗ first_bad_minute(errors, requests, ratio=0.05) error ratio ratio рдкреЗрдХреНрд╖рд╛ рдЬрд╛рд╕реНрдд рдЕрд╕рд▓реЗрд▓рд╛ рдкрд╣рд┐рд▓рд╛ minute рджреЗрддреЗ тАФ рд╣рд╛рдиреАрдЪреА рд╕реБрд░реБрд╡рд╛рдд. suspect_changes(changes, spike_minute, lookback=30) рддреНрдпрд╛ minute рдкрд░реНрдпрдВрдд рдЖрдгрд┐ рддреЛ рдзрд░реВрди, lookback рдорд┐рдирд┐рдЯрд╛рдВрддреАрд▓ рдкреНрд░рддреНрдпреЗрдХ (minute, what) рдмрджрд▓ рджреЗрддреЗ (рджрд┐рд▓реЗрд▓реНрдпрд╛ рдХреНрд░рдорд╛рдиреЗ). five_whys(chain) рдЙрддреНрддрд░рд╛рдВрдЪреНрдпрд╛ рд╕рд╛рдЦрд│реАрд▓рд╛ рдХреНрд░рдорд╛рдВрдХ рджреЗрддреЗ. obs/demo.py рдордзрд▓реЗ CHANGES рджрд┐рд╡рд╕рд╛рддрд▓реЗ рдЪрд╛рд░ рдмрджрд▓ рдиреЛрдВрджрд╡рддреЗ: 90 рд▓рд╛ рдПрдХ cache TTL config, 250 рдЖрдгрд┐ 398 рд▓рд╛ deploys, рдЖрдгрд┐ rollback.

ЁЯзк рдХрд░реВрди рдкрд╛рд╣рд╛

Minute 385 рд▓рд╛ рдПрдХ feature flag рдЬреЛрдбрд╛ (рджреБрд╕рд░рд╛ рд╕рдВрд╢рдпрд┐рдд), рдЖрдгрдЦреА рдорд╛рдЧреЗ рдкрд╛рд╣рд╛ тАФ рдордЧ рддреАрдЪ tools slow leak рдХрдбреЗ рд╡рд│рд╡рд╛ (0.5% рдЪреА рд░реЗрд╖рд╛), рдЖрдгрд┐ рдЦреВрдк рд▓рд╡рдХрд░ рдерд╛рдВрдмрдгрд╛рд░реА рд╕рд╛рдЦрд│реА рд▓рд┐рд╣рд╛:

python3 obs/demo.py rootcause
python3 - <<'EOF'
import sys; sys.path.insert(0, "obs"); from demo import ERR, REQ, CHANGES; from reliability import first_bad_minute, suspect_changes, five_whys
changes = sorted(CHANGES + [(385, "feature flag: new PDF certificates ON")])
for ratio in (0.05, 0.005):
    spike = first_bad_minute(ERR, REQ, ratio)
    print(f"first minute above {ratio:.1%}: {spike}")
    for lookback in (30, 180):
        print(f"   changes in the {lookback} min before тЖТ {suspect_changes(changes, spike, lookback)}")
for w in five_whys(["the results page returned 504", "the grade service was slow", "Katrina pressed deploy"]): print(w)
EOF
python3 obs/test_obs.py

тЬЕ рддрдкрд╛рд╕рд╛ тАФ рддреБрдореНрд╣рд╛рд▓рд╛ рдХрд╛рдп рджрд┐рд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ

rootcause рд╣реЗ print рдХрд░рддреЗ:

тФАтФА the error ratio first crosses 5% at minute 400
   changes in the 30 minutes before: [(398, 'deploy results-api v42')]
   why 1? parents saw 504 on /results
   why 2? the grade service call took 1 s and timed out
   why 3? v42 removed the grade cache, so every view hit the grade service
   why 4? the load test used 3 classes, not 900
   why 5? there is no production-like load test in the release checklist
   the root cause is usually a missing guard in the PROCESS, not a person

рддреБрдордЪрд╛ snippet рд╣реЗ print рдХрд░рддреЛ:

first minute above 5.0%: 400
   changes in the 30 min before тЖТ [(385, 'feature flag: new PDF certificates ON'), (398, 'deploy results-api v42')]
   changes in the 180 min before тЖТ [(250, 'deploy results-api v41'), (385, 'feature flag: new PDF certificates ON'), (398, 'deploy results-api v42')]
first minute above 0.5%: 100
   changes in the 30 min before тЖТ [(90, 'config: cache TTL 60 s тЖТ 5 s')]
   changes in the 180 min before тЖТ [(90, 'config: cache TTL 60 s тЖТ 5 s')]
why 1? the results page returned 504
why 2? the grade service was slow
why 3? Katrina pressed deploy

Tests рдордзреНрдпреЗ тЬЕ L10/L11 incident maths and the suspect change рдЕрд╕рддреЛ.

ЁЯПБ рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдХрд╛рдп рд╕рд┐рджреНрдз рдХреЗрд▓реЗ

Window рдордзреНрдпреЗ рдПрдХрдЪ рдмрджрд▓ рдЕрд╕реЗрд▓ рддрд░ рд╕рдВрд╢рдпрд┐рдд рд╕рд╣рдЬ рдорд┐рд│рддреЛ. рдПрдХ flag рдЬреЛрдбрд╛ рдЖрдгрд┐ рджреЛрди рд╣реЛрддрд╛рдд, рдЖрдгрд┐ 180-minute lookback рддрд┐рд╕рд░рд╛ рдЬреЛрдбрддреЛ тАФ correlation рдпрд╛рджреА рд▓рд╣рд╛рди рдХрд░рддреЗ; рддреЗ рдирд┐рд╡рдб рдХрд░рдд рдирд╛рд╣реА. v42 рдЪреНрдпрд╛ rollback рдиреЗ errors рд╕рдВрдкрд▓реЗ, рдпрд╛рдиреЗрдЪ рдЦрд╛рддреНрд░реА рд╣реЛрддреЗ. Slow leak рд╡рд░ рддреАрдЪ tools рдХрд╛рд╣реАрддрд░реА рдирд╡реЗ рджрд╛рдЦрд╡рддрд╛рдд: minute 90 рдЪрд╛ cache TTL рдмрджрд▓, 0.9% errors рд╕реБрд░реВ рд╣реЛрдгреНрдпрд╛рдЪреНрдпрд╛ рджрд╣рд╛ рдорд┐рдирд┐рдЯреЗ рдЖрдзреА тАФ incident рдиреЗ рдХрдзреАрдЪ рди рдкрд╛рд╣рд┐рд▓реЗрд▓рд╛ рдзрд╛рдЧрд╛, рдХрд╛рд░рдг leak рдиреЗ рдХрдзреАрдЪ page рдХреЗрд▓реЗ рдирд╛рд╣реА. рдЖрдгрд┐ рддреАрди-рдХрд╛ рдЪреА рд╕рд╛рдЦрд│реА рдПрдХрд╛ рд╡реНрдпрдХреНрддреАрд╡рд░ рдерд╛рдВрдмрд▓реА. рддреНрдпрд╛рддрд▓реЗ рдХрд╛рд╣реАрдЪ рдХрддрд░рд┐рдирд╛рд▓рд╛ рджреЛрд╖ рджрд┐рд▓реНрдпрд╛рд╢рд┐рд╡рд╛рдп рджреБрд░реБрд╕реНрдд рдХрд░рддрд╛ рдпреЗрдд рдирд╛рд╣реА. рд╡рд┐рдЪрд╛рд░рдд рд░рд╛рд╣рд╛: рдХреЛрдгрддреАрд╣реА test fail рди рд╣реЛрддрд╛ рдПрдХрд╛ deploy рдиреЗ cache рдХрд╕реЗ рдХрд╛рдвреВрди рдЯрд╛рдХрд▓реЗ?

тЪая╕П рдиреЗрд╣рдореАрдЪреНрдпрд╛ рдЪреБрдХрд╛

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд

рдЦрд▒реНрдпрд╛ account рд╡рд░ тАФ рдкреНрд░рддреНрдпреЗрдХ рдмрджрд▓ рдкреНрд░рддреНрдпреЗрдХ graph рд╡рд░ рдЯрд╛рдХрд╛. Grafana annotation рдЬреЛрдбрдгрд╛рд░реА deploy pipeline step (рд╡реЗрд│ epoch milliseconds рдордзреНрдпреЗ):

curl -s -X POST https://grafana.example.school/api/annotations \
  -H "Authorization: Bearer $GRAFANA_TOKEN" -H 'Content-Type: application/json' \
  -d "{\"time\": $(date +%s000), \"tags\": [\"deploy\", \"results-api\"], \"text\": \"deploy results-api v42\"}"

рд╣рд╛рдиреАрдЪреНрдпрд╛ рд╕реБрд░реБрд╡рд╛рддреАрдЪреНрдпрд╛ рдЖрд╕рдкрд╛рд╕ рдХрд╛рдп рдмрджрд▓рд▓реЗ рддреЗ рд╢реЛрдзрд╛:

kubectl rollout history deployment/results-api
git log --since="2026-09-27 11:00" --until="2026-09-27 11:45" --oneline -- services/results-api
aws cloudtrail lookup-events --start-time 2026-09-27T05:30:00Z --end-time 2026-09-27T06:15:00Z \
    --lookup-attributes AttributeKey=ReadOnly,AttributeValue=false \
    --query 'Events[].[EventTime,EventName,Username]' --output table

Error ratio version рдиреБрд╕рд╛рд░ рддреБрд▓рдирд╛ рдХрд░рд╛ тАФ рдлрдХреНрдд v42 fail рд╣реЛрдд рдЕрд╕реЗрд▓, рддрд░ рд╕рдВрд╢рдпрд┐рддрд╛рдЪреА рдЦрд╛рддреНрд░реА рд╣реЛрддреЗ (рдпрд╛рд╕рд╛рдареА metric рд╡рд░ version label рд▓рд╛рдЧрддреЛ, рдЬреЛ low-cardinality рдЖрд╣реЗ рдЖрдгрд┐ рдареЗрд╡рдгреНрдпрд╛рд╕рд╛рд░рдЦрд╛ рдЖрд╣реЗ):

sum by (version) (rate(http_requests_total{job="results-api", status=~"5.."}[5m]))
/
sum by (version) (rate(http_requests_total{job="results-api"}[5m]))

Datadog graphs рд╡рд░ deploys events рдореНрд╣рдгреВрди рджрд╛рдЦрд╡рддреЛ рдЖрдгрд┐ APM рдЪреНрдпрд╛ Deployment Tracking рдордзреНрдпреЗ versions рдЪреА рддреБрд▓рдирд╛ рдХрд░рддреЛ; CloudWatch Amazon DevOps Guru рдХрд┐рдВрд╡рд╛ change events рдЕрд╕рд▓реЗрд▓реНрдпрд╛ dashboard рдордзреВрди CloudTrail рдмрджрд▓ рд╡рд░ рдорд╛рдВрдбреВ рд╢рдХрддреЛ. Feature flag tools (LaunchDarkly, Unleash, OpenFeature providers) audit log рдареЗрд╡рддрд╛рдд тАФ рддреЛрд╣реА рд╢реЛрдзрд╛рдд рд╕рдорд╛рд╡рд┐рд╖реНрдЯ рдХрд░рд╛.

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд рд╣реЗ рдХрд╛ рдорд╣рддреНрддреНрд╡рд╛рдЪреЗ: рдПрдЦрд╛рджреНрдпрд╛ рдмрджрд▓рд╛рдЪреА рд╡реЗрд│реЗрд╕рд╣ рдиреЛрдВрдж рдирд╕реЗрд▓, рддрд░ рддреНрдпрд╛рдЪреА correlation рдХрд░рддрд╛ рдпреЗрдд рдирд╛рд╣реА. рдкреНрд░рддреНрдпреЗрдХ deploy, config рдмрджрд▓ рдЖрдгрд┐ flag flip рдиреЗ рдЕрд╕рд╛ timestamped event рд╕реЛрдбрд▓рд╛ рдкрд╛рд╣рд┐рдЬреЗ рдЬреЛ graphs рджрд╛рдЦрд╡реВ рд╢рдХрддреАрд▓.

тПня╕П рдкреБрдвреЗ

Cause рд╕рд╛рдкрдбрд▓реЗ. рдЖрддрд╛ рддреЗ рд▓рд┐рд╣реВрди рдареЗрд╡рд╛ тАФ рджреЛрд╖ рди рджреЗрддрд╛ тАФ рдЕрд╢рд╛ fixes рд╕рд╣ рдЬреНрдпрд╛рдВрдирд╛ owner рдЖрдгрд┐ date рдЖрд╣реЗ, рдореНрд╣рдгрдЬреЗ рддреАрдЪ рдЖрдЧ рдкреБрдиреНрд╣рд╛ рдХрдзреАрдЪ рд▓рд╛рдЧрдгрд╛рд░ рдирд╛рд╣реА.

git checkout lesson-12-postmortems

ЁЯФН Lesson 11 тАФ Root-cause analysis: why did it really happen

ЁЯУН You are here: Lesson 11 of 12 ┬╖ Previous: lesson-10-incident-response ┬╖ Next: lesson-12-postmortems


ЁЯУж What's in this branch

Lessons 01тАУ10, plus finding the real cause: change correlation (what changed just before?), the five whys, and why the answer is almost always a missing guard in the process, not a person. first_bad_minute(), suspect_changes() and five_whys() live in obs/reliability.py; rootcause() in obs/demo.py runs them on results day.

ЁЯзТ Explain like I'm 5

After the drill, everyone is safe. Now Katrina asks: why did it happen?

First she looks at the clock and the visitors' book. The smoke started at 11:40. Who came in or changed something just before 11:40? The book says: 11:38, a new toaster was plugged in the staff room. That is the first suspect.

Then she asks "why?" тАФ and asks again, and again:

  1. Why was there smoke? The toaster burned the bread.
  2. Why did it burn? It was left on for 20 minutes.
  3. Why was it left on? It has no timer that switches it off.
  4. Why was a toaster with no timer allowed? Nobody checks new kitchen things.
  5. Why does nobody check? There is no checklist for new kitchen things.

Katrina does not stop at "Mr. Somebody left it on". People forget things; that will happen again tomorrow. A timer and a checklist fix it for everyone.

ЁЯЧ║я╕П Diagram

flowchart TB
    spike["ЁЯУИ error ratio first above 5% at minute 400"]
    ch["ЁЯЧВя╕П changes in the 30 minutes before<br/>398 deploy results-api v42 тЬЕ suspect<br/>(90 config ┬╖ 250 v41 ┬╖ 430 rollback тАФ outside)"]
    spike --> ch
    ch --> w1["why 1? parents saw 504 on /results"]
    w1 --> w2["why 2? the grade service call took 1 s and timed out"]
    w2 --> w3["why 3? v42 removed the grade cache"]
    w3 --> w4["why 4? the load test used 3 classes, not 900"]
    w4 --> w5["why 5? no production-like load test in the release checklist"]
    w5 --> fix["ЁЯФз fix the PROCESS, not the person"]

ЁЯЧ║я╕П Drawn version + a lab: https://school-edh.pages.dev/observability/lesson-diagrams.html#l11

тЭУ What

ЁЯдФ Why

Because the fix you choose depends on how deep you go. Stop at "v42 was bad" and you roll back and wait for v43 to do the same. Stop at "the engineer forgot the cache" and you blame a person, people hide the next mistake, and nothing changes. Go down to "the release checklist has no production-like load test" and you fix a gap that would have caught this deploy тАФ and the next five.

ЁЯФз How (in this repo)

first_bad_minute(errors, requests, ratio=0.05) in obs/reliability.py returns the first minute where the error ratio is above ratio тАФ the start of the harm. suspect_changes(changes, spike_minute, lookback=30) returns every (minute, what) change in the lookback minutes up to and including that minute (in the order given). five_whys(chain) numbers a chain of answers. CHANGES in obs/demo.py records the day's four changes: a cache TTL config at 90, deploys at 250 and 398, and the rollback.

ЁЯзк Try it

Add a feature flag at minute 385 (a second suspect), look back further тАФ then point the same tools at the slow leak (a 0.5% line), and write a chain that stops too early:

python3 obs/demo.py rootcause
python3 - <<'EOF'
import sys; sys.path.insert(0, "obs"); from demo import ERR, REQ, CHANGES; from reliability import first_bad_minute, suspect_changes, five_whys
changes = sorted(CHANGES + [(385, "feature flag: new PDF certificates ON")])
for ratio in (0.05, 0.005):
    spike = first_bad_minute(ERR, REQ, ratio)
    print(f"first minute above {ratio:.1%}: {spike}")
    for lookback in (30, 180):
        print(f"   changes in the {lookback} min before тЖТ {suspect_changes(changes, spike, lookback)}")
for w in five_whys(["the results page returned 504", "the grade service was slow", "Katrina pressed deploy"]): print(w)
EOF
python3 obs/test_obs.py

тЬЕ Verify тАФ what you should see

rootcause prints:

тФАтФА the error ratio first crosses 5% at minute 400
   changes in the 30 minutes before: [(398, 'deploy results-api v42')]
   why 1? parents saw 504 on /results
   why 2? the grade service call took 1 s and timed out
   why 3? v42 removed the grade cache, so every view hit the grade service
   why 4? the load test used 3 classes, not 900
   why 5? there is no production-like load test in the release checklist
   the root cause is usually a missing guard in the PROCESS, not a person

Your snippet prints:

first minute above 5.0%: 400
   changes in the 30 min before тЖТ [(385, 'feature flag: new PDF certificates ON'), (398, 'deploy results-api v42')]
   changes in the 180 min before тЖТ [(250, 'deploy results-api v41'), (385, 'feature flag: new PDF certificates ON'), (398, 'deploy results-api v42')]
first minute above 0.5%: 100
   changes in the 30 min before тЖТ [(90, 'config: cache TTL 60 s тЖТ 5 s')]
   changes in the 180 min before тЖТ [(90, 'config: cache TTL 60 s тЖТ 5 s')]
why 1? the results page returned 504
why 2? the grade service was slow
why 3? Katrina pressed deploy

The tests include тЬЕ L10/L11 incident maths and the suspect change.

ЁЯПБ What you just proved

One change in the window makes an easy suspect. Add a flag and there are two, and a 180-minute lookback adds a third тАФ correlation narrows the list; it does not choose. The rollback of v42 ending the errors is what confirms it. The same tools on the slow leak point at something new: the cache TTL change at minute 90, ten minutes before the 0.9% errors began тАФ a lead the incident never looked at, because the leak never paged. And the three-why chain stopped at a person. Nothing in it can be fixed except by blaming Katrina. Keep asking: why could one deploy remove a cache with no test failing?

тЪая╕П Common mistakes

ЁЯПн In production

On a real account тАФ put every change on every graph. A deploy pipeline step that adds a Grafana annotation (time in epoch milliseconds):

curl -s -X POST https://grafana.example.school/api/annotations \
  -H "Authorization: Bearer $GRAFANA_TOKEN" -H 'Content-Type: application/json' \
  -d "{\"time\": $(date +%s000), \"tags\": [\"deploy\", \"results-api\"], \"text\": \"deploy results-api v42\"}"

Find what changed around the start of the harm:

kubectl rollout history deployment/results-api
git log --since="2026-09-27 11:00" --until="2026-09-27 11:45" --oneline -- services/results-api
aws cloudtrail lookup-events --start-time 2026-09-27T05:30:00Z --end-time 2026-09-27T06:15:00Z \
    --lookup-attributes AttributeKey=ReadOnly,AttributeValue=false \
    --query 'Events[].[EventTime,EventName,Username]' --output table

Compare the error ratio by version тАФ if only v42 fails, the suspect is confirmed (this needs a version label on the metric, which is low-cardinality and worth having):

sum by (version) (rate(http_requests_total{job="results-api", status=~"5.."}[5m]))
/
sum by (version) (rate(http_requests_total{job="results-api"}[5m]))

Datadog shows deploys as events on graphs and compares versions in APM's Deployment Tracking; CloudWatch can overlay CloudTrail changes through Amazon DevOps Guru or a dashboard with the change events. Feature flag tools (LaunchDarkly, Unleash, OpenFeature providers) keep an audit log тАФ include it in the search.

ЁЯПн Why this matters in production: if a change is not recorded with a time, it cannot be correlated. Make every deploy, config change and flag flip leave a timestamped event where the graphs can show it.

тПня╕П Next

The cause is found. Now write it down тАФ without blame тАФ with fixes that have an owner and a date, so the same fire never starts twice.

git checkout lesson-12-postmortems
тЖР Previousincident responseNext тЖТpostmortems

This page is the lesson's README from the lesson-11-root-cause branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.