ЁЯПл The SchoolтА║ЁЯУИ ScalingтА║ЁЯЫЯ рдзрдбрд╛ 13 тАФ рдмрд┐рдШрд╛рдбрд╛рддреВрди рдЯрд┐рдХрдгреЗ: рдирд┐рдХрд╛рд▓рд╛рдЪреНрдпрд╛ рджрд┐рд╡рд╢реА рдПрдЦрд╛рджрд╛ stall рдмрд┐рдШрдбрддреЛ рддреЗрд╡реНрд╣рд╛
ЁЯЦ╝я╕П See the drawing + lab ЁЯПа Course home ЁЯМ┐ Branch on GitHub тЬПя╕П View source
ЁЯЦ╝я╕П рдЖрдХреГрддреА рдЖрдгрд┐ labThe drawing + lab рдкреВрд░реНрдг рдкрд╛рдирд╛рд╡рд░ рдЙрдШрдбрд╛ тЖЧOpen full page тЖЧ

ЁЯЫЯ рдзрдбрд╛ 13 тАФ рдмрд┐рдШрд╛рдбрд╛рддреВрди рдЯрд┐рдХрдгреЗ: рдирд┐рдХрд╛рд▓рд╛рдЪреНрдпрд╛ рджрд┐рд╡рд╢реА рдПрдЦрд╛рджрд╛ stall рдмрд┐рдШрдбрддреЛ рддреЗрд╡реНрд╣рд╛

ЁЯУН рддреБрдореНрд╣реА рдЗрдереЗ рдЖрд╣рд╛рдд: 13 рдкреИрдХреА рдзрдбрд╛ 13 ┬╖ рдорд╛рдЧреАрд▓: lesson-12-queues-and-blueprint


ЁЯУж рдпрд╛ рдмреНрд░рдБрдЪрдордзреНрдпреЗ рдХрд╛рдп рдЖрд╣реЗ

рдзрдбреЗ 01тАУ12, рдЖрдгрд┐ scaling рдЪрд╛ рджреБрд╕рд░рд╛ рдЕрд░реНрдзрд╛ рднрд╛рдЧ: рдПрдЦрд╛рджрд╛ рднрд╛рдЧ рдмрд┐рдШрдбрд▓рд╛ рддрд░реА рдЬрддреНрд░рд╛ рдЪрд╛рд▓реВ рдареЗрд╡рдгреЗ. Servers рдЬреНрдпрд╛ service рд▓рд╛ call рдХрд░рддрд╛рдд рддреА рдЕрдбрдХрд▓реА рдЕрд╕реЗрд▓, рддрд░ рдЬрд╛рд╕реНрдд servers рдиреЗ рдлрд╛рдпрджрд╛ рд╣реЛрдд рдирд╛рд╣реА. рддреБрдореНрд╣реА timeouts, exponential backoff + jitter рд╕рд╣ retries, circuit breaker, health checks, graceful degradation, idempotency, dead-letter queue, multi-AZ, рдЖрдгрд┐ dependencies рдЪреНрдпрд╛ рд╕рд╛рдЦрд│реАрдд availability рдХрд╛ рдЧреБрдгрд▓реА рдЬрд╛рддреЗ рд╣реЗ рд╢рд┐рдХрддрд╛. scale/demo.py рдордзрд▓реЗ resilience() рдпрд╛рддрд▓реЗ рдкреНрд░рддреНрдпреЗрдХ рджрд╛рдЦрд╡рддреЗ, рдЖрдгрд┐ рддреНрдпрд╛рдЪреЗ models scale/sim.py рдЪреНрдпрд╛ "surviving failure" рднрд╛рдЧрд╛рдд рдЖрд╣реЗрдд.

ЁЯзТ 5 рд╡рд░реНрд╖рд╛рдВрдЪреНрдпрд╛ рдореБрд▓рд╛рд▓рд╛ рд╕рдордЬрд╛рд╡рд▓реНрдпрд╛рд╕рд╛рд░рдЦреЗ

рдирд┐рдХрд╛рд▓рд╛рдЪрд╛ рджрд┐рд╡рд╕, рд╕рдХрд╛рд│рдЪреЗ 10 рд╡рд╛рдЬрд▓реЗрдд. рдЬрддреНрд░реЗрдЪреНрдпрд╛ рдорд╛рдЧрдЪреНрдпрд╛ рдмрд╛рдЬреВрдЪрд╛ grades stall рдмрд┐рдШрдбрддреЛ. рддреНрдпрд╛рдЪрд╛ рдХрд╛рд░рдХреВрди рдЕрдбрдХрд▓рд╛ рдЖрд╣реЗ рдЖрдгрд┐ рдЙрддреНрддрд░ рджреЗрдд рдирд╛рд╣реА.

рдкреНрд░рддреНрдпреЗрдХ рдЦрд┐рдбрдХреА рдкрд╛рд▓рдХрд╛рдВрдирд╛ grades stall рдХрдбреЗ рдкрд╛рдард╡рддреЗ рдЖрдгрд┐ рд╡рд╛рдЯ рдкрд╛рд╣рддреЗ. рдЖрдгрд┐ рд╡рд╛рдЯ рдкрд╛рд╣рддреЗ. рд▓рд╡рдХрд░рдЪ рдкреНрд░рддреНрдпреЗрдХ рдХрд╛рд░рдХреВрди рд╡рд╛рдЯ рдкрд╛рд╣рдд рдЕрд╕рддреЛ, рдЖрдгрд┐ рд╕рдВрдкреВрд░реНрдг рдЬрддреНрд░рд╛ рдардкреНрдк рд╣реЛрддреЗ тАФ рдПрдХрд╛ stall рдореБрд│реЗ. рдореНрд╣рдгреВрди рджреАрдкрд┐рдХрд╛ рдкреНрд░рддреНрдпреЗрдХ рдЦрд┐рдбрдХреАрд▓рд╛ рдХрд╛рд╣реА рд╕рд╛рдзрдиреЗ рджреЗрддреЗ:

рдЖрдгрд┐ рдЖрдгрдЦреА рдПрдХ рдирд┐рдпрдо: рдкрд╛рд▓рдХрд╛рдЪреА рднреЗрдЯ gate, рдПрдХ рдЦрд┐рдбрдХреА, рд╕реВрдЪрдирд╛ рдлрд▓рдХ рдЖрдгрд┐ рдиреЛрдВрджрд╡рд╣реА рдпрд╛рдВрдордзреВрди рдЬрд╛рддреЗ. рддреНрдпрд╛рддрд▓реЗ рдкреНрд░рддреНрдпреЗрдХ 1,000 рдкреИрдХреА 999 рд╡реЗрд│рд╛ рдЪрд╛рд▓рдд рдЕрд╕реЗрд▓, рддрд░ рд╕рдВрдкреВрд░реНрдг рднреЗрдЯ рддреНрдпрд╛рддрд▓реНрдпрд╛ рдХреЛрдгрддреНрдпрд╛рд╣реА рдПрдХрд╛рдкреЗрдХреНрд╖рд╛ рдереЛрдбреА рдХрдореА рд╡реЗрд│рд╛ рдпрд╢рд╕реНрд╡реА рд╣реЛрддреЗ.

ЁЯЧ║я╕П рдЖрдХреГрддреА

flowchart LR
    u["ЁЯСк parent"] --> api["ЁЯН│ API"]
    api -->|"тП▒я╕П timeout 1 s"| cb{"ЁЯФМ circuit breaker<br/>closed ┬╖ open ┬╖ half-open"}
    cb -->|"closed"| dep["ЁЯУТ grades service"]
    cb -->|"open: fail fast"| fb["ЁЯкз fallback<br/>cached page, '5 minutes old'"]
    dep -.->|"error"| rt["ЁЯЩП retry<br/>backoff + jitter<br/>500 at once тЖТ 81 per 10 ms"]
    rt -.-> cb
    q["ЁЯУм SQS"] -->|"at least once"| wk["ЁЯЦия╕П worker<br/>ЁЯФЦ idempotency key"]
    q -.->|"after maxReceiveCount"| dlq["ЁЯЧВя╕П DLQ"]
    az["ЁЯПвЁЯПвЁЯПв 3 AZs<br/>one AZ 99.5% тЖТ three 99.999988%"]

ЁЯЧ║я╕П рд░реЗрдЦрд╛рдЯрд▓реЗрд▓реА рдЖрд╡реГрддреНрддреА + рдПрдХ lab: https://school-edh.pages.dev/scaling/lesson-diagrams.html#l13

тЭУ рдХрд╛рдп

ЁЯдФ рдХрд╛

рдХрд╛рд░рдг scaling рднрд╛рдЧ рд╡рд╛рдврд╡рддреЗ, рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХ рднрд╛рдЧ рдмрд┐рдШрдбреВ рд╢рдХрддреЛ. рдзрдбреЗ 01тАУ12 рдиреЗ рдЬрддреНрд░рд╛ рдореЛрдареА рдХреЗрд▓реА: рдЬрд╛рд╕реНрдд servers, pods, Lambdas, caches, replicas, partitions рдЖрдгрд┐ queues. рдЬрд╛рд╕реНрдд рднрд╛рдЧ рдЕрд╕рд▓реЗ рдХреА рдХреЛрдгрддреНрдпрд╛рд╣реА рджрд┐рд╡рд╢реА рдХреБрдареЗрддрд░реА рдмрд┐рдШрд╛рдб рд╣реЛрдгреНрдпрд╛рдЪреА рд╢рдХреНрдпрддрд╛ рдЬрд╛рд╕реНрдд. Timeouts рдирд╕рддреАрд▓ рддрд░ рдПрдХ рдЕрдбрдХрд▓реЗрд▓реА service рдкреНрд░рддреНрдпреЗрдХ thread рднрд░реВрди рдЯрд╛рдХрддреЗ. Jitter рдирд╕реЗрд▓ рддрд░ retries рдПрдХрд╛ рдЫреЛрдЯреНрдпрд╛ рдЕрдбрдерд│реНрдпрд╛рдЪреЗ рд╡рд╛рджрд│ рдХрд░рддрд╛рдд. Idempotency рдирд╕реЗрд▓ рддрд░ retry рдкрд╛рд▓рдХрд╛рдХрдбреВрди рджреЛрдирджрд╛ рд╢реБрд▓реНрдХ рдШреЗрддреЛ. рдпрд╛ рдзрдбреНрдпрд╛рддрд▓реА рд╕рд╛рдзрдиреЗ рдореЛрдареНрдпрд╛ system рд▓рд╛ рд╕реНрдерд┐рд░рд╣реА рдмрдирд╡рддрд╛рдд.

ЁЯФз рдХрд╕реЗ (рдпрд╛ repo рдордзреНрдпреЗ)

scale/sim.py рдЪреНрдпрд╛ "surviving failure" рднрд╛рдЧрд╛рдд рдЫреЛрдЯреЗ models рдЖрд╣реЗрдд:

scale/demo.py рдордзрд▓реЗ resilience() рдкреНрд░рддреНрдпреЗрдХрд╛рдЪреЗ рдПрдХ рдЙрджрд╛рд╣рд░рдг рдЪрд╛рд▓рд╡рддреЗ. Breaker рдЙрджрд╛рд╣рд░рдгрд╛рддрд▓реА payments service рд╕реЗрдХрдВрдж 0 рдкрд╛рд╕реВрди рд╕реЗрдХрдВрдж 29 рдкрд░реНрдпрдВрдд рдмрдВрдж рдЕрд╕рддреЗ, рдЖрдгрд┐ demo рддрд┐рд▓рд╛ 40 рд╕реЗрдХрдВрдж рджрд░ рд╕реЗрдХрдВрджрд╛рд▓рд╛ рдПрдХрджрд╛ call рдХрд░рддреЗ.

ЁЯзк рдХрд░реВрди рдкрд╛рд╣рд╛

python3 scale/demo.py resilience
python3 - <<'EOF'
import sys; sys.path.insert(0, "scale"); from sim import call_with_timeout, littles_law, backoff_ms, retry_storm
for lat in (120, 900, 1500, 30000):
    print(f"latency {lat:>5} ms, timeout 1,000 ms тЖТ", call_with_timeout(lat, 1000))
for wait in (30, 1, 0.2):
    print(f"2,000 req/s, each caller waits up to {wait:>4} s тЖТ {littles_law(2000, wait):>7,.0f} requests stuck at worst")
print("backoff (no jitter):", [backoff_ms(a) for a in range(7)], "ms")
for clients in (100, 500, 2000):
    print(f"{clients:>4} clients, 4 retries тЖТ busiest 10 ms: {retry_storm(clients, 4, False):>4} without jitter, {retry_storm(clients, 4, True):>3} with full jitter")
EOF
python3 - <<'EOF'
import sys; sys.path.insert(0, "scale"); from sim import CircuitBreaker, deliver_at_least_once, process, redrive, parallel, serial
for threshold, open_s in ((5, 10), (3, 10), (5, 2), (10, 30)):
    b = CircuitBreaker(threshold, open_s); log = [b.call(t, t >= 30) for t in range(40)]
    print(f"breaker {threshold:>2} failures, open {open_s:>2} s тЖТ reached the sick service {log.count('failed'):>2} ┬╖ failed fast {log.count('fast-fail'):>2} ┬╖ first ok at {log.index('ok')} s")
d = deliver_at_least_once(["pay-1", "pay-2", "pay-3", "pay-4"], ["pay-2", "pay-4", "pay-4"])
print(f"{len(d)} deliveries of 4 payments тЖТ plain {process(d, False)} charges ┬╖ idempotent {process(d, True)}")
for mr in (1, 3, 5):
    done, dlq, rec = redrive([f"pdf-{i}" for i in range(10)], {"pdf-3", "pdf-7"}, max_receives=mr)
    print(f"maxReceiveCount {mr} тЖТ made {len(done)}, DLQ {dlq}, receives {rec}")
for a in (0.99, 0.995):
    print(f"one AZ {a:.1%} тЖТ two {parallel(a, 2):.4%} ┬╖ three {parallel(a, 3):.6%}")
print(f"chain of 3 at 99.9% тЖТ {serial(0.999, 0.999, 0.999):.2%} ┬╖ chain of 6 тЖТ {serial(*[0.999] * 6):.2%} ┬╖ 6 links each on 3 AZs тЖТ {serial(*[parallel(0.999, 3)] * 6):.6%}")
EOF
python3 scale/test_scale.py

тЬЕ рддрдкрд╛рд╕рд╛ тАФ рддреБрдореНрд╣рд╛рд▓рд╛ рдХрд╛рдп рджрд┐рд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ

resilience рд╣реЗ рдЫрд╛рдкрддреЗ:

тФАтФА timeouts: the results API usually answers in 120 ms; today it is stuck at 30 s
   latency    120 ms, timeout 1,000 ms тЖТ ('ok', 120)
   latency  30000 ms, timeout 1,000 ms тЖТ ('timeout', 1000)
   without a timeout every thread waits 30 s тАФ 2,000 req/s ├Ч 30 s = 60,000 stuck requests (Little's law again)
тФАтФА retries: 500 clients fail at once and retry 4 times (backoff 100 тЖТ 800 ms)
   no jitter   тЖТ busiest 10 ms:  500 retries hit the database together
   full jitter тЖТ busiest 10 ms:   81 retries hit the database together
тФАтФА circuit breaker (5 failures тЖТ open for 10 s тЖТ one trial call); the payments service is down for 30 s
   calls that reached the sick service: 7 ┬╖ failed fast without waiting: 27 ┬╖ state at 40 s: closed
тФАтФА graceful degradation: the grade service is down тЖТ show the results page from Redis, marked '5 minutes old'
   a slower, older page beats an error page; non-essential parts (photos, recommendations) can switch off
тФАтФА at-least-once: 4 deliveries for 3 payments тЖТ plain worker charges 4, idempotent worker charges 3
тФАтФА a poison message: 9 PDFs made, ['pdf-7'] moved to the DLQ after 3 receives (12 receives in total)
тФАтФА multi-AZ: one AZ at 99.5% тЖТ two AZs 99.9975% ┬╖ three AZs 99.999988%
   a chain API 99.9% ├Ч Redis 99.9% ├Ч database 99.95% = 99.75% тАФ every dependency lowers the total
   scaling adds capacity; resilience keeps the fair open when one part of it breaks

рддреБрдордЪрд╛ рдкрд╣рд┐рд▓рд╛ snippet рд╣реЗ рдЫрд╛рдкрддреЛ:

latency   120 ms, timeout 1,000 ms тЖТ ('ok', 120)
latency   900 ms, timeout 1,000 ms тЖТ ('ok', 900)
latency  1500 ms, timeout 1,000 ms тЖТ ('timeout', 1000)
latency 30000 ms, timeout 1,000 ms тЖТ ('timeout', 1000)
2,000 req/s, each caller waits up to   30 s тЖТ  60,000 requests stuck at worst
2,000 req/s, each caller waits up to    1 s тЖТ   2,000 requests stuck at worst
2,000 req/s, each caller waits up to  0.2 s тЖТ     400 requests stuck at worst
backoff (no jitter): [100, 200, 400, 800, 1600, 3200, 3200] ms
 100 clients, 4 retries тЖТ busiest 10 ms:  100 without jitter,  21 with full jitter
 500 clients, 4 retries тЖТ busiest 10 ms:  500 without jitter,  81 with full jitter
2000 clients, 4 retries тЖТ busiest 10 ms: 2000 without jitter, 292 with full jitter

рддреБрдордЪрд╛ рджреБрд╕рд░рд╛ snippet рд╣реЗ рдЫрд╛рдкрддреЛ:

breaker  5 failures, open 10 s тЖТ reached the sick service  7 ┬╖ failed fast 27 ┬╖ first ok at 34 s
breaker  3 failures, open 10 s тЖТ reached the sick service  5 ┬╖ failed fast 27 ┬╖ first ok at 32 s
breaker  5 failures, open  2 s тЖТ reached the sick service 17 ┬╖ failed fast 13 ┬╖ first ok at 30 s
breaker 10 failures, open 30 s тЖТ reached the sick service 10 ┬╖ failed fast 29 ┬╖ first ok at 39 s
7 deliveries of 4 payments тЖТ plain 7 charges ┬╖ idempotent 4
maxReceiveCount 1 тЖТ made 8, DLQ ['pdf-3', 'pdf-7'], receives 10
maxReceiveCount 3 тЖТ made 8, DLQ ['pdf-3', 'pdf-7'], receives 14
maxReceiveCount 5 тЖТ made 8, DLQ ['pdf-3', 'pdf-7'], receives 18
one AZ 99.0% тЖТ two 99.9900% ┬╖ three 99.999900%
one AZ 99.5% тЖТ two 99.9975% ┬╖ three 99.999988%
chain of 3 at 99.9% тЖТ 99.70% ┬╖ chain of 6 тЖТ 99.40% ┬╖ 6 links each on 3 AZs тЖТ 99.999999%

Tests тЬЕ L13 timeouts, jittered retries, a circuit breaker and idempotent workers рдЖрдгрд┐ 13/13 passed рдиреЗ рд╕рдВрдкрддрд╛рдд.

ЁЯПБ рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдХрд╛рдп рд╕рд┐рджреНрдз рдХреЗрд▓реЗ

Timeout "60,000 рдЕрдбрдХрд▓реЗрд▓реНрдпрд╛ requests" рдЪреЗ рдЬрд╛рд╕реНрддреАрдд рдЬрд╛рд╕реНрдд 2,000 (1 s) рдХрд┐рдВрд╡рд╛ 400 (0.2 s) рдХрд░рддреЛ. Jitter рдирд╕реЗрд▓ рддрд░ рдкреНрд░рддреНрдпреЗрдХ client рдЪрд╛ retry рддреНрдпрд╛рдЪ 10 ms рдордзреНрдпреЗ рдпреЗрддреЛ тАФ рдПрдХрджрдо 2,000; full jitter рддреНрдпрд╛рдВрдирд╛ 292 рдкрд░реНрдпрдВрдд рдкрд╕рд░рд╡рддреЛ. Breaker рдЖрдЬрд╛рд░реА service рдЪреНрдпрд╛ 30 рд╕реЗрдХрдВрджрд╛рдВрдЪреНрдпрд╛ рдмрдВрджреАрдд рдлрдХреНрдд 7 calls рддрд┐рдЪреНрдпрд╛рдкрд░реНрдпрдВрдд рдкреЛрд╣реЛрдЪреВ рджреЗрддреЛ; рдмрд╛рдХреАрдЪреЗ 23 рд▓рдЧреЗрдЪ рдЕрдкрдпрд╢реА рд╣реЛрддрд╛рдд. рддреНрдпрд╛рдЪреА рдПрдХ рдХрд┐рдВрдордд рдЖрд╣реЗ: service рд╕реЗрдХрдВрдж 30 рд▓рд╛ рдкрд░рдд рдЖрд▓реА, рдЖрдгрд┐ 10 рд╕реЗрдХрдВрджрд╛рдВрдЪреНрдпрд╛ open рд╡реЗрд│реЗрдиреЗ рддреЗ рдлрдХреНрдд 34 рд▓рд╛ рдУрд│рдЦрд▓реЗ. рдХрдореА open рд╡реЗрд│ (2 s) рд▓рдЧреЗрдЪ рд╕рд╛рд╡рд░рддреЗ рдкрдг рдЖрдЬрд╛рд░реА service рдХрдбреЗ 17 calls рдкрд╛рдард╡рддреЗ. рдпрд╛ рджреЛрдШрд╛рдВрдордзреВрди рддреБрдореНрд╣реА рдирд┐рд╡рдб рдХрд░рддрд╛. рд╢рд┐рдХреНрдХрд╛ (idempotency) 7 deliveries рдЕрд╕рддрд╛рдирд╛рд╣реА 4 payments рд╕рд╛рдареА 4 рд╢реБрд▓реНрдХрдЪ рдШреЗрддреЛ. maxReceiveCount рдХрд╛рд╣реАрд╣реА рдЕрд╕реЛ, DLQ 2 рдЦрд░рд╛рдм messages рдмрд╛рдЬреВрд▓рд╛ рдареЗрд╡рддреЗ; рдЬрд╛рд╕реНрдд count рдлрдХреНрдд рддреНрдпрд╛рдВрдЪреНрдпрд╛рд╡рд░ рдЬрд╛рд╕реНрдд receives рдЦрд░реНрдЪ рдХрд░рддреЛ тАФ рдкрдг 1 рдЪрд╛ count рдПрдХрд╛ рдЫреЛрдЯреНрдпрд╛ рдЕрдбрдерд│реНрдпрд╛рдирдВрддрд░ рдЪрд╛рдВрдЧрд▓реНрдпрд╛ message рд▓рд╛рд╣реА рдмрд╛рдЬреВрд▓рд╛ рдареЗрд╡реЗрд▓. рдЖрдгрд┐ рд╕рд╛рдЦрд│реА рдкреНрд░рддреНрдпреЗрдХ рджреБрд╡реНрдпрд╛рд╕реЛрдмрдд рдереЛрдбреЗ рдЧрдорд╛рд╡рддреЗ (6 рджреБрд╡реЗ: 99.40%), рддрд░ 3 AZs рдордзрд▓реНрдпрд╛ рдкреНрд░рддреА рддреЛ рдЖрдХрдбрд╛ рдкреБрдиреНрд╣рд╛ рд╡рд╛рдврд╡рддрд╛рдд тАФ рдЬреЛрдкрд░реНрдпрдВрдд AZs рд╕реНрд╡рддрдВрддреНрд░рдкрдгреЗ рдмрд┐рдШрдбрддрд╛рдд рддреЛрдкрд░реНрдпрдВрдд.

тЪая╕П рдиреЗрд╣рдореАрдЪреНрдпрд╛ рдЪреБрдХрд╛

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд

рдЦрд▒реНрдпрд╛ account рд╡рд░ тАФ рдкреНрд░рддреНрдпреЗрдХ client рд╡рд░ timeouts рдЖрдгрд┐ retries. Connect рдЖрдгрд┐ read timeout, full-jitter backoff рдЖрдгрд┐ рдПрдХреВрдг рд╡реЗрд│реЗрдЪреА рдорд░реНрдпрд╛рджрд╛ рдЕрд╕рд▓реЗрд▓реЗ Python requests:

import random, time, requests

def get_results(cls, attempts=4, base=0.1, cap=3.2, budget=5.0):
    start = time.monotonic()
    for a in range(attempts):
        try:
            r = requests.get(f"https://grades.internal/results/{cls}", timeout=(0.5, 1.0))  # connect, read (s)
            if r.status_code < 500 and r.status_code != 429:
                return r
        except requests.RequestException:
            pass
        wait = random.uniform(0, min(cap, base * 2 ** a))          # full jitter
        if time.monotonic() - start + wait > budget:
            break
        time.sleep(wait)
    return None            # the caller shows the cached page, marked "5 minutes old"

(requests рдордзреНрдпреЗ read timeout рдореНрд╣рдгрдЬреЗ bytes рдордзрд▓реЗ рд╕рд░реНрд╡рд╛рдд рдореЛрдареЗ рдЕрдВрддрд░, рд╕рдВрдкреВрд░реНрдг рдЙрддреНрддрд░рд╛рд╡рд░рдЪреА рдорд░реНрдпрд╛рджрд╛ рдирд╛рд╣реА тАФ loop рдордзрд▓реЗ budget рд╣реА рдПрдХреВрдг рдорд░реНрдпрд╛рджрд╛ рдЖрд╣реЗ.)

AWS SDKs рддреБрдордЪреНрдпрд╛рд╕рд╛рдареА retry рдХрд░рддрд╛рдд. Retry mode рдирд┐рд╡рдбрд╛: standard (jitter рд╕рд╣ exponential backoff, default 3 рдкреНрд░рдпрддреНрди) рдХрд┐рдВрд╡рд╛ adaptive (standard рдЕрдзрд┐рдХ client-side rate limiter, рдЬреЛ AWS рддреБрдореНрд╣рд╛рд▓рд╛ throttle рдХрд░рддреЗ рддреЗрд╡реНрд╣рд╛ рд╡реЗрдЧ рдХрдореА рдХрд░рддреЛ). boto3 рд╕рд╛рдареА:

import boto3
from botocore.config import Config
cfg = Config(connect_timeout=1, read_timeout=2, retries={"mode": "standard", "max_attempts": 3})
ddb = boto3.client("dynamodb", config=cfg)

рдХрд┐рдВрд╡рд╛ рдХреЛрдгрддреНрдпрд╛рд╣реА SDK рдЖрдгрд┐ AWS CLI рд╕рд╛рдареА: export AWS_RETRY_MODE=standard AWS_MAX_ATTEMPTS=3.

Circuit breakers тАФ code рдордзрд▓реА library (Java рд╕рд╛рдареА Resilience4j, .NET рд╕рд╛рдареА Polly, Node.js рд╕рд╛рдареА opossum, Python рд╕рд╛рдареА pybreaker), рдХрд┐рдВрд╡рд╛ service mesh рд╣реЗ рдкреНрд░рддреНрдпреЗрдХ service рд╕рд╛рдареА рдХрд░рддреЗ. Istio: route рд╡рд░ timeout рдЖрдгрд┐ retries, рдЖрдгрд┐ рдЖрдЬрд╛рд░реА pod рд▓рд╛ рдХрд╛рд╣реА рд╡реЗрд│ pool рдордзреВрди рдХрд╛рдврдгрд╛рд░реЗ outlier detection:

apiVersion: networking.istio.io/v1
kind: VirtualService
metadata: { name: grades }
spec:
  hosts: [grades]
  http:
    - route: [{ destination: { host: grades } }]
      timeout: 2s
      retries: { attempts: 2, perTryTimeout: 1s, retryOn: "5xx,reset,connect-failure" }
---
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata: { name: grades }
spec:
  host: grades
  trafficPolicy:
    outlierDetection:
      consecutive5xxErrors: 5
      interval: 10s
      baseEjectionTime: 30s
      maxEjectionPercent: 50

Health checks тАФ ALB target group /health рддрдкрд╛рд╕рддреЛ (рдзрдбрд╛ 05). Route 53 health checks рд╕рдВрдкреВрд░реНрдг endpoint рд╡рд░ рд▓рдХреНрд╖ рдареЗрд╡рддрд╛рдд, рдЖрдгрд┐ рдкрд╣рд┐рд▓рд╛ рдЕрдкрдпрд╢реА рдард░рд▓рд╛ рдХреА failover record users рдирд╛ рджреБрд╕рд▒реНрдпрд╛ endpoint рдХрдбреЗ рдкрд╛рдард╡рддреЛ:

aws route53 create-health-check --caller-reference results-api-2026-05 \
    --health-check-config '{"Type":"HTTPS","FullyQualifiedDomainName":"api.school.example",
      "ResourcePath":"/health","RequestInterval":30,"FailureThreshold":3}'

SQS DLQ redrive policy + maxReceiveCount (рдзрдбрд╛ 12), рдЖрдгрд┐ рдордЧ рджреБрд░реБрд╕реНрддреАрдирдВрддрд░ рдмрд╛рдЬреВрд▓рд╛ рдареЗрд╡рд▓реЗрд▓реЗ messages рдкрд░рдд рд╣рд▓рд╡рд╛:

aws sqs set-queue-attributes --queue-url https://sqs.ap-south-1.amazonaws.com/111122223333/certificates \
    --attributes '{"RedrivePolicy":"{\"deadLetterTargetArn\":\"arn:aws:sqs:ap-south-1:111122223333:certificates-dlq\",\"maxReceiveCount\":\"5\"}"}'
aws sqs start-message-move-task \
    --source-arn arn:aws:sqs:ap-south-1:111122223333:certificates-dlq \
    --destination-arn arn:aws:sqs:ap-south-1:111122223333:certificates

Idempotency keys тАФ Powertools for AWS Lambda (Python) рдкреНрд░рддреНрдпреЗрдХ key DynamoDB рдордзреНрдпреЗ рд╕рд╛рдард╡рддреЗ рдЖрдгрд┐ рдкреБрдиреНрд╣рд╛ рдЖрд▓реЗрд▓реНрдпрд╛ request рд▓рд╛ рд╕рд╛рдард╡рд▓реЗрд▓реЗ рдЙрддреНрддрд░ рдкрд░рдд рджреЗрддреЗ (table рд▓рд╛ string partition key id, рдЖрдгрд┐ expiration attribute рд╡рд░ TTL рд▓рд╛рдЧрддреЛ):

from aws_lambda_powertools.utilities.idempotency import (
    DynamoDBPersistenceLayer, IdempotencyConfig, idempotent)

persistence = DynamoDBPersistenceLayer(table_name="school-idempotency")
config = IdempotencyConfig(event_key_jmespath="powertools_json(body).payment_id",
                           expires_after_seconds=3600)

@idempotent(config=config, persistence_store=persistence)
def handler(event, context):
    charge_fee(event)                     # runs once per payment_id, even if the event comes twice
    return {"statusCode": 200, "body": "paid"}

Database рдЖрдгрд┐ cache рд╕рд╛рдареА Multi-AZ тАФ рдЖрдгрд┐ рдирд┐рдХрд╛рд▓рд╛рдЪреНрдпрд╛ рджрд┐рд╡рд╕рд╛рдЖрдзреА failover рдЪреА рдЪрд╛рдЪрдгреА рдШреНрдпрд╛:

aws rds modify-db-instance --db-instance-identifier school-db --multi-az --apply-immediately
aws rds reboot-db-instance --db-instance-identifier school-db --force-failover
aws elasticache modify-replication-group --replication-group-id school-sessions \
    --multi-az-enabled --automatic-failover-enabled --apply-immediately

AWS Fault Injection Service (FIS) рд╕рд╣ chaos testing тАФ рдПрдХрд╛ AZ рдордзрд▓реЗ API instances 10 рдорд┐рдирд┐рдЯрд╛рдВрд╕рд╛рдареА рдмрдВрдж рдХрд░рд╛, рдЖрдгрд┐ p99 alarm рд╡рд╛рдЬрд▓рд╛ рддрд░ experiment рд▓рдЧреЗрдЪ рдерд╛рдВрдмрд╡рд╛:

{
  "description": "results day drill: lose AZ a",
  "roleArn": "arn:aws:iam::111122223333:role/fis-school",
  "stopConditions": [{ "source": "aws:cloudwatch:alarm",
    "value": "arn:aws:cloudwatch:ap-south-1:111122223333:alarm:results-p99-high" }],
  "targets": { "api-in-az-a": {
    "resourceType": "aws:ec2:instance",
    "resourceTags": { "app": "school-api" },
    "filters": [{ "path": "Placement.AvailabilityZone", "values": ["ap-south-1a"] }],
    "selectionMode": "ALL" } },
  "actions": { "stop-az-a": {
    "actionId": "aws:ec2:stop-instances",
    "parameters": { "startInstancesAfterDuration": "PT10M" },
    "targets": { "Instances": "api-in-az-a" } } }
}
aws fis create-experiment-template --cli-input-json file://lose-az-a.json
aws fis start-experiment --experiment-template-id EXT1a2b3c4d5e6f7

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд рд╣реЗ рдХрд╛ рдорд╣рддреНрддреНрд╡рд╛рдЪреЗ: рдирд┐рдХрд╛рд▓рд╛рдЪреНрдпрд╛ рдкрд╛рдирд╛рдЪреА рдкреНрд░рддреНрдпреЗрдХ dependency рд▓рд┐рд╣реВрди рдХрд╛рдврд╛, рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХрд╛рд╕рд╛рдареА: рддрд┐рдЪрд╛ timeout, рддрд┐рдЪрд╛ retry рдирд┐рдпрдо, рддрд┐рдЪрд╛ breaker, рдЖрдгрд┐ рддрд┐рдЪрд╛ fallback. рдордЧ рдореБрджреНрджрд╛рдо рдПрдХ рддреЛрдбрд╛ (FIS, рдХрд┐рдВрд╡рд╛ staging рдордзреНрдпреЗ рдПрдЦрд╛рджрд╛ container рдерд╛рдВрдмрд╡рд╛) рдЖрдгрд┐ p99 рдЖрдгрд┐ error rate рдкрд╛рд╣рд╛. рд╕рд░рд╛рд╡ рдХреЗрд▓реЗрд▓рд╛ рдмрд┐рдШрд╛рдб рд╣реА рдЫреЛрдЯреА рдШрдЯрдирд╛ рдЕрд╕рддреЗ; рд╕рд░рд╛рд╡ рди рдХреЗрд▓реЗрд▓рд╛ рдмрд┐рдШрд╛рдб рдореНрд╣рдгрдЬреЗ outage.

ЁЯОУ рдЬрддреНрд░рд╛ рдЖрддрд╛ рддреБрдордЪреА рдЖрд╣реЗ

Up рдХрд┐рдВрд╡рд╛ out тЖТ рдмрд╛рдВрдзрдгреНрдпрд╛рдЖрдзреА рдореЛрдЬрд╛ тЖТ рдкреНрд░рддреНрдпреЗрдХ gate рд╡рд░ рдЭреЗрд░реЙрдХреНрд╕ рдкреНрд░рддреА тЖТ live pages рдХрд╛рд╣реА рд╕реЗрдХрдВрджрд╛рдВрд╕рд╛рдареА cache тЖТ рдХреЛрдгрддреАрд╣реА рдЦрд┐рдбрдХреА рдХреЛрдгрддреНрдпрд╛рд╣реА рдкрд╛рд▓рдХрд╛рд▓рд╛ рд╕реЗрд╡рд╛ рджреЗрддреЗ тЖТ рд░рд╛рдВрдЧ рд╡рд╛рдврд▓реА рдХреА рдЬрд╛рд╕реНрдд рдЦрд┐рдбрдХреНрдпрд╛ тЖТ pods рдЖрдгрд┐ рддреЗ рдмрд╕рддрд╛рдд рддреА рдЯреЗрдмрд▓реЗ тЖТ рдкреНрд░рддреНрдпреЗрдХ рдкрд╛рд▓рдХрд╛рд╕рд╛рдареА рдПрдХ рдЦрд┐рдбрдХреА тЖТ рд╕реВрдЪрдирд╛ рдлрд▓рдХ тЖТ рдиреЛрдВрджрд╡рд╣реАрдЪреНрдпрд╛ рдкреНрд░рддреА тЖТ рдкрд╕рд░рдгрд╛рд▒реНрдпрд╛ key рдиреЗ рд╡рд┐рднрд╛рдЧрд▓реЗрд▓реЗ рдХрдкреНрдкреЗ тЖТ рдкреЗрдЯреАрддрд▓реЗ tokens тЖТ рдПрдХ stopwatch, рд╕рднреНрдп retries, рдПрдХ breaker switch рдЖрдгрд┐ рддреАрди рдЗрдорд╛рд░рддреА. рддреБрдореНрд╣реА рдлрдХреНрдд scaling рд╢рд┐рдХрд▓рд╛ рдирд╛рд╣реАрдд тАФ рддреБрдореНрд╣реА рдирд┐рдХрд╛рд▓рд╛рдЪреНрдпрд╛ рджрд┐рд╡рд╕рд╛рдЪреЗ рдирд┐рдпреЛрдЬрди рдХрд░реВ рд╢рдХрддрд╛, рдЖрдгрд┐ рдПрдЦрд╛рджрд╛ stall рдмрд┐рдШрдбрд▓рд╛ рддрд░реА рдЬрддреНрд░рд╛ рдЪрд╛рд▓реВ рдареЗрд╡реВ рд╢рдХрддрд╛. ЁЯУИЁЯЫЯЁЯОУ

тПня╕П рдкреБрдвреЗ

рдЗрддрд░ рд╢рд╛рд│рд╛. Kubernetes рд╢рд╛рд│рд╛ рдЦрд┐рдбрдХреНрдпрд╛ рдЪрд╛рд▓рд╡рддреЗ; Database рд╢рд╛рд│рд╛ рдиреЛрдВрджрд╡рд╣реАрдд рдЖрдгрдЦреА рдЦреЛрд▓рд╡рд░ рдЬрд╛рддреЗ; API Gateway рд╢рд╛рд│рд╛ API рд╕рдореЛрд░рдЪреЗ рдкреБрдврдЪреЗ рдХрд╛рд░реНрдпрд╛рд▓рдп рдмрд╛рдВрдзрддреЗ; School portal рдордзреНрдпреЗ рдмрд╛рдХреА рд╕рдЧрд│реЗ рдЖрд╣реЗ.

git checkout main
python3 scale/demo.py     # one last run, for fun

ЁЯЫЯ Lesson 13 тАФ Surviving failure: when a stall breaks on results day

ЁЯУН You are here: Lesson 13 of 13 ┬╖ Previous: lesson-12-queues-and-blueprint


ЁЯУж What's in this branch

Lessons 01тАУ12, plus the other half of scaling: keeping the fair open when one part breaks. More servers do not help when a service they call is stuck. You learn timeouts, retries with exponential backoff + jitter, the circuit breaker, health checks, graceful degradation, idempotency, the dead-letter queue, multi-AZ, and why availability multiplies along a chain of dependencies. resilience() in scale/demo.py shows each one, with the models in the "surviving failure" part of scale/sim.py.

ЁЯзТ Explain like I'm 5

Results day, 10 o'clock. The grades stall at the back of the fair breaks. Its clerk is stuck and does not answer.

Every counter sends parents to the grades stall and waits. And waits. Soon every clerk is waiting, and the whole fair stands still тАФ because of one stall. So Dipika gives each counter a few tools:

And one more rule: a parent's visit passes the gate, a counter, the notice board and the register. If each one works 999 times in 1,000, the whole visit works a little less often than any one of them.

ЁЯЧ║я╕П Diagram

flowchart LR
    u["ЁЯСк parent"] --> api["ЁЯН│ API"]
    api -->|"тП▒я╕П timeout 1 s"| cb{"ЁЯФМ circuit breaker<br/>closed ┬╖ open ┬╖ half-open"}
    cb -->|"closed"| dep["ЁЯУТ grades service"]
    cb -->|"open: fail fast"| fb["ЁЯкз fallback<br/>cached page, '5 minutes old'"]
    dep -.->|"error"| rt["ЁЯЩП retry<br/>backoff + jitter<br/>500 at once тЖТ 81 per 10 ms"]
    rt -.-> cb
    q["ЁЯУм SQS"] -->|"at least once"| wk["ЁЯЦия╕П worker<br/>ЁЯФЦ idempotency key"]
    q -.->|"after maxReceiveCount"| dlq["ЁЯЧВя╕П DLQ"]
    az["ЁЯПвЁЯПвЁЯПв 3 AZs<br/>one AZ 99.5% тЖТ three 99.999988%"]

ЁЯЧ║я╕П Drawn version + a lab: https://school-edh.pages.dev/scaling/lesson-diagrams.html#l13

тЭУ What

ЁЯдФ Why

Because scaling adds parts, and every part can fail. Lessons 01тАУ12 made the fair bigger: more servers, pods, Lambdas, caches, replicas, partitions and queues. With more parts, a failure somewhere is more likely on any given day. Without timeouts, one stuck service fills every thread. Without jitter, retries turn a blip into a storm. Without idempotency, a retry charges a parent twice. The tools in this lesson make a large system also a steady one.

ЁЯФз How (in this repo)

The "surviving failure" part of scale/sim.py holds small models:

resilience() in scale/demo.py runs one example of each. The payments service in the breaker example is down from second 0 to second 29, and the demo calls it once a second for 40 seconds.

ЁЯзк Try it

python3 scale/demo.py resilience
python3 - <<'EOF'
import sys; sys.path.insert(0, "scale"); from sim import call_with_timeout, littles_law, backoff_ms, retry_storm
for lat in (120, 900, 1500, 30000):
    print(f"latency {lat:>5} ms, timeout 1,000 ms тЖТ", call_with_timeout(lat, 1000))
for wait in (30, 1, 0.2):
    print(f"2,000 req/s, each caller waits up to {wait:>4} s тЖТ {littles_law(2000, wait):>7,.0f} requests stuck at worst")
print("backoff (no jitter):", [backoff_ms(a) for a in range(7)], "ms")
for clients in (100, 500, 2000):
    print(f"{clients:>4} clients, 4 retries тЖТ busiest 10 ms: {retry_storm(clients, 4, False):>4} without jitter, {retry_storm(clients, 4, True):>3} with full jitter")
EOF
python3 - <<'EOF'
import sys; sys.path.insert(0, "scale"); from sim import CircuitBreaker, deliver_at_least_once, process, redrive, parallel, serial
for threshold, open_s in ((5, 10), (3, 10), (5, 2), (10, 30)):
    b = CircuitBreaker(threshold, open_s); log = [b.call(t, t >= 30) for t in range(40)]
    print(f"breaker {threshold:>2} failures, open {open_s:>2} s тЖТ reached the sick service {log.count('failed'):>2} ┬╖ failed fast {log.count('fast-fail'):>2} ┬╖ first ok at {log.index('ok')} s")
d = deliver_at_least_once(["pay-1", "pay-2", "pay-3", "pay-4"], ["pay-2", "pay-4", "pay-4"])
print(f"{len(d)} deliveries of 4 payments тЖТ plain {process(d, False)} charges ┬╖ idempotent {process(d, True)}")
for mr in (1, 3, 5):
    done, dlq, rec = redrive([f"pdf-{i}" for i in range(10)], {"pdf-3", "pdf-7"}, max_receives=mr)
    print(f"maxReceiveCount {mr} тЖТ made {len(done)}, DLQ {dlq}, receives {rec}")
for a in (0.99, 0.995):
    print(f"one AZ {a:.1%} тЖТ two {parallel(a, 2):.4%} ┬╖ three {parallel(a, 3):.6%}")
print(f"chain of 3 at 99.9% тЖТ {serial(0.999, 0.999, 0.999):.2%} ┬╖ chain of 6 тЖТ {serial(*[0.999] * 6):.2%} ┬╖ 6 links each on 3 AZs тЖТ {serial(*[parallel(0.999, 3)] * 6):.6%}")
EOF
python3 scale/test_scale.py

тЬЕ Verify тАФ what you should see

resilience prints:

тФАтФА timeouts: the results API usually answers in 120 ms; today it is stuck at 30 s
   latency    120 ms, timeout 1,000 ms тЖТ ('ok', 120)
   latency  30000 ms, timeout 1,000 ms тЖТ ('timeout', 1000)
   without a timeout every thread waits 30 s тАФ 2,000 req/s ├Ч 30 s = 60,000 stuck requests (Little's law again)
тФАтФА retries: 500 clients fail at once and retry 4 times (backoff 100 тЖТ 800 ms)
   no jitter   тЖТ busiest 10 ms:  500 retries hit the database together
   full jitter тЖТ busiest 10 ms:   81 retries hit the database together
тФАтФА circuit breaker (5 failures тЖТ open for 10 s тЖТ one trial call); the payments service is down for 30 s
   calls that reached the sick service: 7 ┬╖ failed fast without waiting: 27 ┬╖ state at 40 s: closed
тФАтФА graceful degradation: the grade service is down тЖТ show the results page from Redis, marked '5 minutes old'
   a slower, older page beats an error page; non-essential parts (photos, recommendations) can switch off
тФАтФА at-least-once: 4 deliveries for 3 payments тЖТ plain worker charges 4, idempotent worker charges 3
тФАтФА a poison message: 9 PDFs made, ['pdf-7'] moved to the DLQ after 3 receives (12 receives in total)
тФАтФА multi-AZ: one AZ at 99.5% тЖТ two AZs 99.9975% ┬╖ three AZs 99.999988%
   a chain API 99.9% ├Ч Redis 99.9% ├Ч database 99.95% = 99.75% тАФ every dependency lowers the total
   scaling adds capacity; resilience keeps the fair open when one part of it breaks

Your first snippet prints:

latency   120 ms, timeout 1,000 ms тЖТ ('ok', 120)
latency   900 ms, timeout 1,000 ms тЖТ ('ok', 900)
latency  1500 ms, timeout 1,000 ms тЖТ ('timeout', 1000)
latency 30000 ms, timeout 1,000 ms тЖТ ('timeout', 1000)
2,000 req/s, each caller waits up to   30 s тЖТ  60,000 requests stuck at worst
2,000 req/s, each caller waits up to    1 s тЖТ   2,000 requests stuck at worst
2,000 req/s, each caller waits up to  0.2 s тЖТ     400 requests stuck at worst
backoff (no jitter): [100, 200, 400, 800, 1600, 3200, 3200] ms
 100 clients, 4 retries тЖТ busiest 10 ms:  100 without jitter,  21 with full jitter
 500 clients, 4 retries тЖТ busiest 10 ms:  500 without jitter,  81 with full jitter
2000 clients, 4 retries тЖТ busiest 10 ms: 2000 without jitter, 292 with full jitter

Your second snippet prints:

breaker  5 failures, open 10 s тЖТ reached the sick service  7 ┬╖ failed fast 27 ┬╖ first ok at 34 s
breaker  3 failures, open 10 s тЖТ reached the sick service  5 ┬╖ failed fast 27 ┬╖ first ok at 32 s
breaker  5 failures, open  2 s тЖТ reached the sick service 17 ┬╖ failed fast 13 ┬╖ first ok at 30 s
breaker 10 failures, open 30 s тЖТ reached the sick service 10 ┬╖ failed fast 29 ┬╖ first ok at 39 s
7 deliveries of 4 payments тЖТ plain 7 charges ┬╖ idempotent 4
maxReceiveCount 1 тЖТ made 8, DLQ ['pdf-3', 'pdf-7'], receives 10
maxReceiveCount 3 тЖТ made 8, DLQ ['pdf-3', 'pdf-7'], receives 14
maxReceiveCount 5 тЖТ made 8, DLQ ['pdf-3', 'pdf-7'], receives 18
one AZ 99.0% тЖТ two 99.9900% ┬╖ three 99.999900%
one AZ 99.5% тЖТ two 99.9975% ┬╖ three 99.999988%
chain of 3 at 99.9% тЖТ 99.70% ┬╖ chain of 6 тЖТ 99.40% ┬╖ 6 links each on 3 AZs тЖТ 99.999999%

The tests end with тЬЕ L13 timeouts, jittered retries, a circuit breaker and idempotent workers and 13/13 passed.

ЁЯПБ What you just proved

A timeout turns "60,000 stuck requests" into at most 2,000 (1 s) or 400 (0.2 s). Without jitter, every client's retry lands in the same 10 ms тАФ 2,000 at once; full jitter spreads them to 292. The breaker lets only 7 calls reach the sick service during its 30-second outage; the other 23 fail fast. It has a price: the service was back at second 30, and a 10-second open time noticed it only at 34. A short open time (2 s) recovers at once but sends 17 calls to the sick service. You choose between the two. The stamp (idempotency) charges 4 fees for 4 payments, even with 7 deliveries. The DLQ parks the 2 bad messages whatever maxReceiveCount is; a higher count only spends more receives on them тАФ but a count of 1 would also park a good message after a single blip. And the chain loses a little with every link (6 links: 99.40%), while copies in 3 AZs raise it again тАФ as long as the AZs fail independently.

тЪая╕П Common mistakes

ЁЯПн In production

On a real account тАФ timeouts and retries on every client. Python requests with a connect and a read timeout, full-jitter backoff, and a total time budget:

import random, time, requests

def get_results(cls, attempts=4, base=0.1, cap=3.2, budget=5.0):
    start = time.monotonic()
    for a in range(attempts):
        try:
            r = requests.get(f"https://grades.internal/results/{cls}", timeout=(0.5, 1.0))  # connect, read (s)
            if r.status_code < 500 and r.status_code != 429:
                return r
        except requests.RequestException:
            pass
        wait = random.uniform(0, min(cap, base * 2 ** a))          # full jitter
        if time.monotonic() - start + wait > budget:
            break
        time.sleep(wait)
    return None            # the caller shows the cached page, marked "5 minutes old"

(In requests, the read timeout is the longest gap between bytes, not a limit on the whole answer тАФ the loop's budget is the total limit.)

The AWS SDKs retry for you. Choose the retry mode: standard (exponential backoff with jitter, 3 attempts by default) or adaptive (standard plus a client-side rate limiter that slows down when AWS throttles you). For boto3:

import boto3
from botocore.config import Config
cfg = Config(connect_timeout=1, read_timeout=2, retries={"mode": "standard", "max_attempts": 3})
ddb = boto3.client("dynamodb", config=cfg)

or for any SDK and the AWS CLI: export AWS_RETRY_MODE=standard AWS_MAX_ATTEMPTS=3.

Circuit breakers тАФ a library in the code (Resilience4j for Java, Polly for .NET, opossum for Node.js, pybreaker for Python), or the service mesh does it for every service. Istio: a timeout and retries on the route, and outlier detection that removes a sick pod from the pool for a while:

apiVersion: networking.istio.io/v1
kind: VirtualService
metadata: { name: grades }
spec:
  hosts: [grades]
  http:
    - route: [{ destination: { host: grades } }]
      timeout: 2s
      retries: { attempts: 2, perTryTimeout: 1s, retryOn: "5xx,reset,connect-failure" }
---
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata: { name: grades }
spec:
  host: grades
  trafficPolicy:
    outlierDetection:
      consecutive5xxErrors: 5
      interval: 10s
      baseEjectionTime: 30s
      maxEjectionPercent: 50

Health checks тАФ the ALB target group checks /health (lesson 05). Route 53 health checks watch a whole endpoint, and a failover record sends users to a second one when the first fails:

aws route53 create-health-check --caller-reference results-api-2026-05 \
    --health-check-config '{"Type":"HTTPS","FullyQualifiedDomainName":"api.school.example",
      "ResourcePath":"/health","RequestInterval":30,"FailureThreshold":3}'

SQS DLQ redrive policy + maxReceiveCount (lesson 12), then move parked messages back after the fix:

aws sqs set-queue-attributes --queue-url https://sqs.ap-south-1.amazonaws.com/111122223333/certificates \
    --attributes '{"RedrivePolicy":"{\"deadLetterTargetArn\":\"arn:aws:sqs:ap-south-1:111122223333:certificates-dlq\",\"maxReceiveCount\":\"5\"}"}'
aws sqs start-message-move-task \
    --source-arn arn:aws:sqs:ap-south-1:111122223333:certificates-dlq \
    --destination-arn arn:aws:sqs:ap-south-1:111122223333:certificates

Idempotency keys тАФ Powertools for AWS Lambda (Python) stores each key in DynamoDB and returns the saved answer for a repeat (the table needs a string partition key id, and TTL on the expiration attribute):

from aws_lambda_powertools.utilities.idempotency import (
    DynamoDBPersistenceLayer, IdempotencyConfig, idempotent)

persistence = DynamoDBPersistenceLayer(table_name="school-idempotency")
config = IdempotencyConfig(event_key_jmespath="powertools_json(body).payment_id",
                           expires_after_seconds=3600)

@idempotent(config=config, persistence_store=persistence)
def handler(event, context):
    charge_fee(event)                     # runs once per payment_id, even if the event comes twice
    return {"statusCode": 200, "body": "paid"}

Multi-AZ for the database and the cache тАФ and test a failover before results day:

aws rds modify-db-instance --db-instance-identifier school-db --multi-az --apply-immediately
aws rds reboot-db-instance --db-instance-identifier school-db --force-failover
aws elasticache modify-replication-group --replication-group-id school-sessions \
    --multi-az-enabled --automatic-failover-enabled --apply-immediately

Chaos testing with AWS Fault Injection Service (FIS) тАФ stop the API instances in one AZ for 10 minutes, and stop the experiment at once if the p99 alarm fires:

{
  "description": "results day drill: lose AZ a",
  "roleArn": "arn:aws:iam::111122223333:role/fis-school",
  "stopConditions": [{ "source": "aws:cloudwatch:alarm",
    "value": "arn:aws:cloudwatch:ap-south-1:111122223333:alarm:results-p99-high" }],
  "targets": { "api-in-az-a": {
    "resourceType": "aws:ec2:instance",
    "resourceTags": { "app": "school-api" },
    "filters": [{ "path": "Placement.AvailabilityZone", "values": ["ap-south-1a"] }],
    "selectionMode": "ALL" } },
  "actions": { "stop-az-a": {
    "actionId": "aws:ec2:stop-instances",
    "parameters": { "startInstancesAfterDuration": "PT10M" },
    "targets": { "Instances": "api-in-az-a" } } }
}
aws fis create-experiment-template --cli-input-json file://lose-az-a.json
aws fis start-experiment --experiment-template-id EXT1a2b3c4d5e6f7

ЁЯПн Why this matters in production: write down every dependency of the results page, and for each one: its timeout, its retry rule, its breaker, and its fallback. Then break one on purpose (FIS, or stop a container in staging) and watch p99 and the error rate. A failure you have practised is a small event; one you have not is an outage.

ЁЯОУ The fair is yours

Up or out тЖТ count before you build тЖТ photocopies at every gate тЖТ live pages cached for seconds тЖТ any counter serves any parent тЖТ more counters when the queue grows тЖТ pods and the tables they sit at тЖТ a counter for every parent тЖТ the notice board тЖТ copies of the register тЖТ drawers split by a key that spreads тЖТ tokens in a box тЖТ a stopwatch, polite retries, a breaker switch and three buildings. You didn't just learn scaling тАФ you can plan results day, and keep the fair open when a stall breaks. ЁЯУИЁЯЫЯЁЯОУ

тПня╕П Next

The other schools. The Kubernetes school runs the counters; the Database school goes deeper on the register; the API Gateway school builds the front office in front of the API; the School portal has the rest.

git checkout main
python3 scale/demo.py     # one last run, for fun
тЖР Previousqueues and blueprintFinished! Take the quiz тЖТcheck what stuck

This page is the lesson's README from the lesson-13-surviving-failure branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.