ЁЯПл The SchoolтА║ЁЯУР System DesignтА║ЁЯУМ рдзрдбрд╛ 05 тАФ Caching: рд╡рд╛рдЪрдгрд╛рд▒реНрдпрд╛рдЬрд╡рд│ рдкреНрд░рддреА
ЁЯЦ╝я╕П See the drawing + lab ЁЯПа Course home ЁЯМ┐ Branch on GitHub тЬПя╕П View source
ЁЯЦ╝я╕П рдЖрдХреГрддреА рдЖрдгрд┐ labThe drawing + lab рдкреВрд░реНрдг рдкрд╛рдирд╛рд╡рд░ рдЙрдШрдбрд╛ тЖЧOpen full page тЖЧ

ЁЯУМ рдзрдбрд╛ 05 тАФ Caching: рд╡рд╛рдЪрдгрд╛рд▒реНрдпрд╛рдЬрд╡рд│ рдкреНрд░рддреА

ЁЯУН рддреБрдореНрд╣реА рдЗрдереЗ рдЖрд╣рд╛рдд: 18 рдкреИрдХреА рдзрдбрд╛ 05 ┬╖ рдорд╛рдЧреЗ: lesson-04-data-model ┬╖ рдкреБрдвреЗ: lesson-06-load-balancing


ЁЯУж рдпрд╛ рдмреНрд░рдБрдЪрдордзреНрдпреЗ рдХрд╛рдп рдЖрд╣реЗ

рдзрдбреЗ 01тАУ04, рдЖрдгрд┐ рднрд╛рдЧ 2 рдЪрд╛ рдкрд╣рд┐рд▓рд╛ рдмрд╛рдВрдзрдгреАрдЪрд╛ рдареЛрдХрд│рд╛: cache. рдкреНрд░рддреА рдХреБрдареЗ рд░рд╛рд╣реВ рд╢рдХрддрд╛рдд (browser, CDN, app, database), cache рдХрд┐рддреА рдореЛрдард╛ рд╣рд╡рд╛ (hit-ratio рд╡рдХреНрд░), рдЖрдгрд┐ рдкреНрд░рддреА рдкреБрд░реЗрд╢рд╛ рддрд╛рдЬреНрдпрд╛ рдХрд╢рд╛ рд░рд╛рд╣рддрд╛рдд (cache-aside, write-through, TTL, invalidation). design/blocks.py рдордзреАрд▓ LRUCache рдЖрдгрд┐ zipf_stream() рд╡реЗрдЧрд╡реЗрдЧрд│реНрдпрд╛ рдЖрдХрд╛рд░рд╛рдВрдЪреНрдпрд╛ caches рдордзреВрди 20,000 feed reads рдЪрд╛рд▓рд╡рддрд╛рдд; design/demo.py рдордзреАрд▓ cache() рд╡рдХреНрд░ рдЫрд╛рдкрддреЗ.

ЁЯзТ 5 рд╡рд░реНрд╖рд╛рдВрдЪреНрдпрд╛ рдореБрд▓рд╛рд▓рд╛ рд╕рдордЬрд╛рд╡рд▓реНрдпрд╛рд╕рд╛рд░рдЦреЗ

рджрд░рд░реЛрдЬ рд╕рдХрд╛рд│реА рд╢реЗрдХрдбреЛ рдкрд╛рд▓рдХ рдХрд╛рд░реНрдпрд╛рд▓рдпрд╛рдд рдпреЗрддрд╛рдд рдЖрдгрд┐ рд╡рд┐рдЪрд╛рд░рддрд╛рдд: "рд╡рд░реНрдЧ 3A рдЪреНрдпрд╛ рд╕реВрдЪрдирд╛ рдХрд╛рдп рдЖрд╣реЗрдд?" рдХрд╛рд░рдХреВрди рдорд╛рдЧрдЪреНрдпрд╛ рдореЛрдареНрдпрд╛ рдиреЛрдВрджрд╡рд╣реАрдХрдбреЗ рдЬрд╛рддреЗ, 3A рдЪреА рдкрдЯреНрдЯреА рд╢реЛрдзрддреЗ, рдЖрдгрд┐ рддреНрдпрд╛ рд╡рд╛рдЪреВрди рджрд╛рдЦрд╡рддреЗ. рдкреБрдиреНрд╣рд╛. рдкреБрдиреНрд╣рд╛. рддреЗрдЪ рдЙрддреНрддрд░, рд╢реЗрдХрдбреЛ рд╡реЗрд│рд╛.

рдореНрд╣рдгреВрди рджреАрдкрд┐рдХрд╛ рдХрд╛рд░реНрдпрд╛рд▓рдпрд╛рдЪреНрдпрд╛ рджрд╛рд░рд╛рдЬрд╡рд│ рдПрдХ рдЫреЛрдЯрд╛ рд╕реВрдЪрдирд╛ рдлрд▓рдХ ЁЯУМ рд▓рд╛рд╡рддреЗ. рдХрд╛рд░рдХреВрди рд╕рд░реНрд╡рд╛рдд рдЬрд╛рд╕реНрдд рд╡рд┐рдЪрд╛рд░рд▓реНрдпрд╛ рдЬрд╛рдгрд╛рд▒реНрдпрд╛ рд╡рд░реНрдЧрд╛рдВрдЪреА рдкреНрд░рдд рддрд┐рдереЗ рд▓рд╛рд╡рддреЗ. рдЖрддрд╛ рдмрд╣реБрддреЗрдХ рдкрд╛рд▓рдХ рдлрд▓рдХ рд╡рд╛рдЪрддрд╛рдд рдЖрдгрд┐ рдирд┐рдШреВрди рдЬрд╛рддрд╛рдд. рдПрдЦрд╛рджрд╛ рд╡рд░реНрдЧ рдлрд▓рдХрд╛рд╡рд░ рдирд╕реЗрд▓ рддреЗрд╡реНрд╣рд╛рдЪ рдХрд╛рд░рдХреВрди рдиреЛрдВрджрд╡рд╣реАрдХрдбреЗ рдЬрд╛рддреЗ тАФ рдЖрдгрд┐ рдордЧ рддреЛрд╣реА рдлрд▓рдХрд╛рд╡рд░ рд▓рд╛рд╡рддреЗ.

рдлрд▓рдХ рд▓рд╣рд╛рди рдЖрд╣реЗ. рддреЛ рднрд░рд▓рд╛ рдХреА, рдЬрд╛рдЧрд╛ рдХрд░рдгреНрдпрд╛рд╕рд╛рдареА рдХрд╛рд░рдХреВрди рд╕рд░реНрд╡рд╛рдд рдЬрд╛рд╕реНрдд рдХрд╛рд│ рдХреЛрдгреАрд╣реА рди рдкрд╛рд╣рд┐рд▓реЗрд▓реА рдкреНрд░рдд (least recently used) рдХрд╛рдврддреЗ.

рджреАрдкрд┐рдХрд╛рд▓рд╛ рджреЛрди рдЕрдбрдЪрдгреАрдВрдЪреА рддрдпрд╛рд░реА рдХрд░рд╛рд╡реА рд▓рд╛рдЧрддреЗ:

ЁЯЧ║я╕П рдЖрдХреГрддреА

flowchart LR
    ph["ЁЯУ▒ app<br/>browser / phone cache"] --> cdn["ЁЯМН CDN<br/>edge copies"]
    cdn --> api["ЁЯФМ API servers"]
    api -->|"1 get"| rc[("ЁЯУМ Redis<br/>cache-aside")]
    rc -.->|"hit: answer"| api
    api -->|"2 miss: read"| db[("ЁЯРШ database<br/>+ its buffer cache")]
    api -->|"3 put, TTL 60 s"| rc
    post["тЬНя╕П new notice"] -->|"delete key 'feed:3A'"| rc
    curve["ЁЯУИ hit ratio<br/>1% of classes тЖТ 33.8%<br/>10% тЖТ 64.0% ┬╖ 50% тЖТ 83.1%"]

ЁЯЧ║я╕П рдХрд╛рдврд▓реЗрд▓реА рдЖрд╡реГрддреНрддреА + рдПрдХ lab: https://school-edh.pages.dev/system-design/lesson-diagrams.html#l05

тЭУ рдХрд╛рдп

ЁЯдФ рдХрд╛

рдХрд╛рд░рдг рдзрдбрд╛ 02 рдореНрд╣рдгрд╛рд▓рд╛ рдкреНрд░рддреНрдпреЗрдХ write рдорд╛рдЧреЗ 100 reads рдЖрдгрд┐ 3,507 req/s рдЪрд╛ peak. Cache рдЙрддреНрддрд░ рджреЗрддреЛ рддреЛ рдкреНрд░рддреНрдпреЗрдХ read рдореНрд╣рдгрдЬреЗ database рд▓рд╛ рди рдХрд░рд╛рд╡рд╛ рд▓рд╛рдЧрдгрд╛рд░рд╛ read, рдЖрдгрд┐ рддреНрдпрд╛рдЪреЗ рдЙрддреНрддрд░ рджрд╣рд╛рдРрд╡рдЬреА рд╕реБрдорд╛рд░реЗ рдПрдХрд╛ millisecond рдордзреНрдпреЗ рдорд┐рд│рддреЗ. Read-heavy system рд╕рд╛рдареА "read p99 < 200 ms" рдЧрд╛рдардгреНрдпрд╛рдЪрд╛ cache рд╣рд╛ рд╕рд░реНрд╡рд╛рдд рд╕реНрд╡рд╕реНрдд рдорд╛рд░реНрдЧ рдЖрд╣реЗ. рдкрдг рддреЛ рдПрдХ рджреБрд╕рд░реА рдкреНрд░рдд рдЬреЛрдбрддреЛ тАФ рдЖрдгрд┐ рджреЛрди рдкреНрд░рддреА рдПрдХрдореЗрдХрд╛рдВрд╢реА рдЬреБрд│рдд рдирд╛рд╣реАрдд рдЕрд╕реЗ рд╣реЛрдК рд╢рдХрддреЗ. рдореНрд╣рдгреВрди рд░рдЪрдиреЗрддрд▓реНрдпрд╛ рдкреНрд░рддреНрдпреЗрдХ cache рд╕рд╛рдареА рджреЛрди рд▓рд┐рд╣рд┐рд▓реЗрд▓реА рдЙрддреНрддрд░реЗ рд╣рд╡реАрдд: рддреА рдХрд┐рддреА рдЬреБрдиреА рдЕрд╕реВ рд╢рдХрддреЗ? рдЖрдгрд┐ рддреА рдХрд╢рд╛рдиреЗ delete рд╣реЛрддреЗ?

ЁЯФз рдХрд╕реЗ (рдпрд╛ repo рдордзреНрдпреЗ)

design/blocks.py рдордзреАрд▓ zipf_stream(n, items) reads рдЪреА рдПрдХ рдард░рд▓реЗрд▓реА рдпрд╛рджреА рдмрдирд╡рддреЗ, рдЬрд┐рдереЗ item i рд╕рд░реНрд╡рд╛рдд рд▓реЛрдХрдкреНрд░рд┐рдп item рдЪреНрдпрд╛ рд╕реБрдорд╛рд░реЗ 1/(i+1) рдЗрддрдХреНрдпрд╛ рд╡реЗрд│рд╛ рд╡рд┐рдЪрд╛рд░рд▓рд╛ рдЬрд╛рддреЛ тАФ рдХрд╛рд╣реА рд╡рд░реНрдЧ рдЦреВрдк рд▓реЛрдХрдкреНрд░рд┐рдп, рдмрд╣реБрддреЗрдХ рд╢рд╛рдВрдд. LRUCache(size) hits рдЖрдгрд┐ misses рдореЛрдЬрддреЛ рдЖрдгрд┐ рднрд░рд▓реНрдпрд╛рд╡рд░ least recently used рдиреЛрдВрдж рд╡рд┐рд╕рд░рддреЛ. design/demo.py рдордзреАрд▓ cache() 5,000 рд╡рд░реНрдЧрд╛рдВрд╡рд░рдЪреЗ 20,000 reads рддреНрдпрд╛рдВрдкреИрдХреА 1%, 10% рдЖрдгрд┐ 50% рдареЗрд╡рдгрд╛рд▒реНрдпрд╛ caches рдордзреВрди рдЪрд╛рд▓рд╡рддреЗ. Snippet рд╡рдХреНрд░рд╛рдЪреЗ рдЖрдгрдЦреА рдмрд┐рдВрджреВ рдХрд╛рдврддреЛ.

ЁЯзк рдХрд░реВрди рдкрд╛рд╣рд╛

python3 design/demo.py cache
python3 - <<'EOF'
import sys; sys.path.insert(0, "design"); from blocks import LRUCache, zipf_stream
stream = zipf_stream(20_000, 5_000)
for size in (10, 50, 250, 500, 1_000, 2_500, 5_000):
    c = LRUCache(size)
    for k in stream: c.get(k)
    print(f"cache {size:>5} classes тЖТ hit ratio {c.hit_ratio():6.1%} ┬╖ database reads {c.misses:>6,}")
top = sum(1 for k in stream if k < 50)
print(f"the 50 most popular classes get {top:,} of 20,000 reads ({top / 200:.1f}%)")
EOF

тЬЕ рддрдкрд╛рд╕рд╛ тАФ рддреБрдореНрд╣рд╛рд▓рд╛ рдХрд╛рдп рджрд┐рд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ

cache рдЫрд╛рдкрддреЗ:

тФАтФА 20,000 feed reads over 5,000 classes (a few classes are very popular) through an LRU cache
   cache holds    50 classes (1% of them) тЖТ hit ratio  33.8%
   cache holds   500 classes (10% of them) тЖТ hit ratio  64.0%
   cache holds  2500 classes (50% of them) тЖТ hit ratio  83.1%
   a cache holding 1% of the classes already catches a third of the reads тАФ sizing is a curve, measure it
   where: browser ┬╖ CDN ┬╖ API (Redis) ┬╖ database buffer ┬╖ invalidate on write or use a short TTL

рддреБрдордЪрд╛ snippet рдЫрд╛рдкрддреЛ:

cache    10 classes тЖТ hit ratio  14.9% ┬╖ database reads 17,018
cache    50 classes тЖТ hit ratio  33.8% ┬╖ database reads 13,240
cache   250 classes тЖТ hit ratio  55.0% ┬╖ database reads  9,008
cache   500 classes тЖТ hit ratio  64.0% ┬╖ database reads  7,203
cache  1000 classes тЖТ hit ratio  73.1% ┬╖ database reads  5,382
cache  2500 classes тЖТ hit ratio  83.1% ┬╖ database reads  3,375
cache  5000 classes тЖТ hit ratio  83.9% ┬╖ database reads  3,212
the 50 most popular classes get 9,822 of 20,000 reads (49.1%)

ЁЯПБ рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдХрд╛рдп рд╕рд┐рджреНрдз рдХреЗрд▓реЗ

рд╡рдХреНрд░ рд╡рд╛рдХрддреЛ: 10 рд╡рд░реВрди 500 рд╡рд░реНрдЧрд╛рдВрд╡рд░ рдЧреЗрд▓реНрдпрд╛рдиреЗ database reads 17,018 рд╡рд░реВрди 7,203 рд╡рд░ рдпреЗрддрд╛рдд; 2,500 рд╡рд░реВрди рд╕рдЧрд│реНрдпрд╛ 5,000 рд╡рд░ рдЧреЗрд▓реНрдпрд╛рдиреЗ рдлрдХреНрдд рдЖрдгрдЦреА 163 рд╡рд╛рдЪрддрд╛рдд. рд╕рдЧрд│реЗ рдХрд╛рд╣реА рдареЗрд╡рдгрд╛рд░рд╛ cache рд╕реБрджреНрдзрд╛ 83.9% рд╡рд░ рдерд╛рдВрдмрддреЛ тАФ рдкреНрд░рддреНрдпреЗрдХ рд╡рд░реНрдЧрд╛рдЪрд╛ рдкрд╣рд┐рд▓рд╛ read рдиреЗрд╣рдореА miss рдЕрд╕рддреЛ. рдЖрдгрд┐ 50 рд╕рд░реНрд╡рд╛рдд рд▓реЛрдХрдкреНрд░рд┐рдп рд╡рд░реНрдЧрд╛рдВрдирд╛ 49.1% reads рдорд┐рд│рддрд╛рдд, рдкрдг 50 рдЬрд╛рдЧрд╛рдВрдЪрд╛ LRU рдлрдХреНрдд 33.8% рдкрдХрдбрддреЛ: рд╢рд╛рдВрдд рд╡рд░реНрдЧ рд▓реЛрдХрдкреНрд░рд┐рдп рд╡рд░реНрдЧрд╛рдВрдирд╛ рдмрд╛рд╣реЗрд░ рдврдХрд▓рдд рд░рд╛рд╣рддрд╛рдд. рдЖрдХрд╛рд░ рд╡рдХреНрд░рд╛рд╡рд░реВрди рдард░рд╡рд╛, "рдкрд░рд╡рдбреЗрд▓ рддрд┐рддрдХрд╛ рдореЛрдард╛" рдпрд╛рд╡рд░реВрди рдирд╛рд╣реА.

тЪая╕П рдиреЗрд╣рдореАрдЪреНрдпрд╛ рдЪреБрдХрд╛

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд

рд╡рд░реНрдЧрд╛рдЪреНрдпрд╛ feed рд╕рд╛рдареА cache-aside, API рдордзреНрдпреЗ (Python рд╕рд╛рд░рдЦрд╛ pseudo-code):

def class_feed(class_id):
    key = f"feed:{class_id}:v1"
    hit = redis.get(key)
    if hit: return json.loads(hit)
    rows = db.query("SELECT тАж FROM notices WHERE class_id = %s ORDER BY created_at DESC LIMIT 20", class_id)
    redis.set(key, json.dumps(rows), ex=60)          # TTL 60 s: the longest staleness we accept
    return rows

def post_notice(class_id, notice):
    db.insert(notice)                                  # the database first тАФ it is the truth
    redis.delete(f"feed:{class_id}:v1")                # then invalidate

On a real account тАФ рдХрд╛рд░реНрдпрд╛рд▓рдпрд╛рд╕рд╛рдареА рдПрдХ рд▓рд╣рд╛рди Redis (ElastiCache Serverless), рдЖрдгрд┐ рд▓рдХреНрд╖ рдареЗрд╡рд╛рдпрдЪрд╛ CloudWatch рдЖрдХрдбрд╛:

aws elasticache create-serverless-cache --serverless-cache-name notice-cache --engine redis
aws cloudwatch get-metric-statistics --namespace AWS/ElastiCache --metric-name CacheHitRate \
    --dimensions Name=CacheClusterId,Value=notice-cache --statistics Average --period 300 \
    --start-time 2026-09-21T00:00:00Z --end-time 2026-09-22T00:00:00Z

рдЕрдзрд┐рдХ рдЦреЛрд▓рд╛рдд: Scaling рд╢рд╛рд│рд╛, рдзрдбрд╛ 09 (Redis рд╕рд╣ cache-aside, TTLs рдЖрдгрд┐ stampede рдерд╛рдВрдмрд╡рдгреЗ), рдЖрдгрд┐ CDN caching рд╕рд╛рдареА рддрд┐рдерд▓реЗ рдзрдбреЗ 03тАУ04.

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд рд╣реЗ рдХрд╛ рдорд╣рддреНрддреНрд╡рд╛рдЪреЗ: рдкреНрд░рддреНрдпреЗрдХ cache рд╕рд╛рдареА TTL рдЖрдгрд┐ invalidation рдЪрд╛ рдирд┐рдпрдо design doc рдордзреНрдпреЗ рд▓рд┐рд╣рд╛. "60 s рдкрд░реНрдпрдВрдд рдЬреБрдиреЗ" рд╣рд╛ product рдирд┐рд░реНрдгрдп рдЖрд╣реЗ тАФ рддрд╛рддрдбреАрдЪреНрдпрд╛ рд╕реВрдЪрдиреЗрд╕рд╛рдареА 60 s рдЪрд╛рд▓реЗрд▓ рдХрд╛ рддреЗ рд╢рд╛рд│реЗрд▓рд╛ рд╡рд┐рдЪрд╛рд░рд╛ (рдХрджрд╛рдЪрд┐рдд рдЪрд╛рд▓рдгрд╛рд░ рдирд╛рд╣реА; рддрд╛рддрдбреАрдЪреНрдпрд╛ рд╕реВрдЪрдирд╛ cache рд╡рдЧрд│реВ рд╢рдХрддрд╛рдд).

тПня╕П рдкреБрдвреЗ

Cache рдЖрдгрд┐ API рдЕрдиреЗрдХ servers рд╡рд░ рдЪрд╛рд▓рддрд╛рдд. рдкреНрд░рддреНрдпреЗрдХ request рдХреЛрдгрддреНрдпрд╛ server рд▓рд╛ рдЬрд╛рддреЛ тАФ рдЖрдгрд┐ рдХреЛрдгрддрд╛ cache server рдХреЛрдгрддрд╛ рд╡рд░реНрдЧ рдареЗрд╡рддреЛ тАФ рд╣реЗ рдХреЛрдг рдард░рд╡рддреЗ?

git checkout lesson-06-load-balancing

ЁЯУМ Lesson 05 тАФ Caching: copies close to the reader

ЁЯУН You are here: Lesson 05 of 18 ┬╖ Previous: lesson-04-data-model ┬╖ Next: lesson-06-load-balancing


ЁЯУж What's in this branch

Lessons 01тАУ04, plus the first building block of Part 2: the cache. Where copies can live (browser, CDN, app, database), how big a cache must be (the hit-ratio curve), and how copies stay fresh enough (cache-aside, write-through, TTL, invalidation). LRUCache and zipf_stream() in design/blocks.py run 20,000 feed reads through caches of different sizes; cache() in design/demo.py prints the curve.

ЁЯзТ Explain like I'm 5

Every morning, hundreds of parents walk to the office and ask: "What are the notices for class 3A?" The clerk walks to the big register at the back, finds the 3A tab, and reads them out. Again. And again. The same answer, hundreds of times.

So Dipika puts a small notice board ЁЯУМ by the office door. The clerk pins a copy of the most-asked classes there. Now most parents read the board and leave. The clerk walks to the register only when a class is not on the board тАФ and then pins that one too.

The board is small. When it is full, the clerk takes down the copy nobody has looked at for the longest time (least recently used) to make room.

Two problems Dipika must plan for:

ЁЯЧ║я╕П Diagram

flowchart LR
    ph["ЁЯУ▒ app<br/>browser / phone cache"] --> cdn["ЁЯМН CDN<br/>edge copies"]
    cdn --> api["ЁЯФМ API servers"]
    api -->|"1 get"| rc[("ЁЯУМ Redis<br/>cache-aside")]
    rc -.->|"hit: answer"| api
    api -->|"2 miss: read"| db[("ЁЯРШ database<br/>+ its buffer cache")]
    api -->|"3 put, TTL 60 s"| rc
    post["тЬНя╕П new notice"] -->|"delete key 'feed:3A'"| rc
    curve["ЁЯУИ hit ratio<br/>1% of classes тЖТ 33.8%<br/>10% тЖТ 64.0% ┬╖ 50% тЖТ 83.1%"]

ЁЯЧ║я╕П Drawn version + a lab: https://school-edh.pages.dev/system-design/lesson-diagrams.html#l05

тЭУ What

ЁЯдФ Why

Because lesson 02 said 100 reads for every write and a peak of 3,507 req/s. Every read the cache answers is a read the database does not do, and it is answered in about a millisecond instead of ten. A cache is the cheapest way to meet "read p99 < 200 ms" for a read-heavy system. But it adds a second copy тАФ and two copies can disagree. So every cache in a design needs two written answers: how stale may it be? and what deletes it?

ЁЯФз How (in this repo)

zipf_stream(n, items) in design/blocks.py makes a fixed list of reads where item i is asked about 1/(i+1) as often as the most popular one тАФ a few classes are very popular, most are quiet. LRUCache(size) counts hits and misses and forgets the least recently used entry when full. cache() in design/demo.py runs 20,000 reads over 5,000 classes through caches holding 1%, 10% and 50% of them. The snippet draws more points of the curve.

ЁЯзк Try it

python3 design/demo.py cache
python3 - <<'EOF'
import sys; sys.path.insert(0, "design"); from blocks import LRUCache, zipf_stream
stream = zipf_stream(20_000, 5_000)
for size in (10, 50, 250, 500, 1_000, 2_500, 5_000):
    c = LRUCache(size)
    for k in stream: c.get(k)
    print(f"cache {size:>5} classes тЖТ hit ratio {c.hit_ratio():6.1%} ┬╖ database reads {c.misses:>6,}")
top = sum(1 for k in stream if k < 50)
print(f"the 50 most popular classes get {top:,} of 20,000 reads ({top / 200:.1f}%)")
EOF

тЬЕ Verify тАФ what you should see

cache prints:

тФАтФА 20,000 feed reads over 5,000 classes (a few classes are very popular) through an LRU cache
   cache holds    50 classes (1% of them) тЖТ hit ratio  33.8%
   cache holds   500 classes (10% of them) тЖТ hit ratio  64.0%
   cache holds  2500 classes (50% of them) тЖТ hit ratio  83.1%
   a cache holding 1% of the classes already catches a third of the reads тАФ sizing is a curve, measure it
   where: browser ┬╖ CDN ┬╖ API (Redis) ┬╖ database buffer ┬╖ invalidate on write or use a short TTL

Your snippet prints:

cache    10 classes тЖТ hit ratio  14.9% ┬╖ database reads 17,018
cache    50 classes тЖТ hit ratio  33.8% ┬╖ database reads 13,240
cache   250 classes тЖТ hit ratio  55.0% ┬╖ database reads  9,008
cache   500 classes тЖТ hit ratio  64.0% ┬╖ database reads  7,203
cache  1000 classes тЖТ hit ratio  73.1% ┬╖ database reads  5,382
cache  2500 classes тЖТ hit ratio  83.1% ┬╖ database reads  3,375
cache  5000 classes тЖТ hit ratio  83.9% ┬╖ database reads  3,212
the 50 most popular classes get 9,822 of 20,000 reads (49.1%)

ЁЯПБ What you just proved

The curve bends: going from 10 to 500 classes cuts database reads from 17,018 to 7,203; going from 2,500 to all 5,000 saves only 163 more. Even a cache that holds everything stops at 83.9% тАФ the first read of each class is always a miss. And the 50 most popular classes get 49.1% of reads, but a 50-slot LRU catches only 33.8%: the quiet classes keep pushing popular ones out. Size by the curve, not by "as big as we can afford".

тЪая╕П Common mistakes

ЁЯПн In production

Cache-aside for the class feed, in the API (Python-like pseudo-code):

def class_feed(class_id):
    key = f"feed:{class_id}:v1"
    hit = redis.get(key)
    if hit: return json.loads(hit)
    rows = db.query("SELECT тАж FROM notices WHERE class_id = %s ORDER BY created_at DESC LIMIT 20", class_id)
    redis.set(key, json.dumps(rows), ex=60)          # TTL 60 s: the longest staleness we accept
    return rows

def post_notice(class_id, notice):
    db.insert(notice)                                  # the database first тАФ it is the truth
    redis.delete(f"feed:{class_id}:v1")                # then invalidate

On a real account тАФ a small Redis for the office (ElastiCache Serverless), and the CloudWatch number to watch:

aws elasticache create-serverless-cache --serverless-cache-name notice-cache --engine redis
aws cloudwatch get-metric-statistics --namespace AWS/ElastiCache --metric-name CacheHitRate \
    --dimensions Name=CacheClusterId,Value=notice-cache --statistics Average --period 300 \
    --start-time 2026-09-21T00:00:00Z --end-time 2026-09-22T00:00:00Z

Go deeper: the Scaling school, lesson 09 (cache-aside with Redis, TTLs and stopping a stampede), and lessons 03тАУ04 there for CDN caching.

ЁЯПн Why this matters in production: write the TTL and the invalidation rule into the design doc for every cache. "Stale for up to 60 s" is a product decision тАФ ask the school if 60 s is acceptable for an urgent notice (it may not be; urgent notices can skip the cache).

тПня╕П Next

The cache and the API run on many servers. Who decides which server gets each request тАФ and which cache server keeps which class?

git checkout lesson-06-load-balancing
тЖР Previousdata modelNext тЖТload balancing

This page is the lesson's README from the lesson-05-caching branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.