🏫 The School›📈 Scaling›⚖️ धडा 05 — Stateless APIs + load balancers: कोणतीही खिडकी कोणत्याही पालकाला सेवा देऊ शकते
🖼️ See the drawing + lab 🏠 Course home 🌿 Branch on GitHub ✏️ View source
🖼️ आकृती आणि labThe drawing + lab पूर्ण पानावर उघडा ↗Open full page ↗

⚖️ धडा 05 — Stateless APIs + load balancers: कोणतीही खिडकी कोणत्याही पालकाला सेवा देऊ शकते

📍 तुम्ही इथे आहात: 13 पैकी धडा 05 · मागे: lesson-04-dynamic-at-the-edge · पुढे: lesson-06-auto-scaling


📦 या ब्रँचमध्ये काय आहे

धडे 01–04, आणि त्याशिवाय scale out शक्य करणारा नियम: stateless API. Sessions, uploads आणि caches server च्या बाहेर राहतात, म्हणून load balancer प्रत्येक request कोणत्याही चालू (healthy) server कडे पाठवू शकतो. scale/demo.py मधले balance() कतरिनाच्या सहा clicks चा तीन servers वरून मागोवा घेते — एकदा तिचे session एका server च्या memory मध्ये असताना, एकदा shared store मध्ये असताना.

🧒 5 वर्षांच्या मुलाला समजावल्यासारखे

कतरिना जत्रेत येते आणि खिडकी 1 वर sign in करते. खिडकी 1 वरची कारकून तिच्या स्वतःच्या वहीत 📓 "कतरिना — signed in" असे लिहिते.

फाटकावरचा मदतनीस प्रत्येक पुढची भेट पुढच्या खिडकीकडे पाठवतो: 1, 2, 3, 1, 2, 3. कतरिनाची दुसरी भेट खिडकी 2 कडे जाते. ती कारकून आपल्या वहीत पाहते: कतरिना नाही. "कृपया पुन्हा sign in करा." कतरिना वैतागते. तिच्या सहा पैकी चार भेटी अशाच संपतात.

दीपिका हे एका सामायिक वहीने दुरुस्त करते, जी प्रत्येक खिडकीला पोहोचता येईल अशा टेबलावर असते (Redis). आता कोणतीही कारकून "कतरिना — signed in" तपासू शकते. प्रत्येक खिडकी प्रत्येक पालकाला सेवा देऊ शकते. आता दीपिका 3 खिडक्या उघडू शकते किंवा 30.

फाटकावरचा मदतनीस दर काही सेकंदांनी प्रत्येक खिडकीजवळून जातो आणि विचारतो "सगळे ठीक आहे ना?". कारकुनाने उत्तर दिले नाही तर मदतनीस तिथे पालक पाठवणे थांबवतो — हा health check.

🗺️ आकृती

flowchart LR
    k["🙋 Katrina<br/>6 clicks"] --> alb["⚖️ load balancer<br/>round robin · health checks"]
    alb --> s1["🖥️ srv-1"]
    alb --> s2["🖥️ srv-2"]
    alb --> s3["🖥️ srv-3"]
    s1 --> r["📓 shared session store<br/>ElastiCache (Redis/Valkey) · DynamoDB"]
    s2 --> r
    s3 --> r
    m["❌ session in srv-1 memory<br/>clicks 2, 3, 5, 6 → login page"]

🗺️ काढलेली आकृती + एक lab: https://school-edh.pages.dev/scaling/lesson-diagrams.html#l05

❓ काय

🤔 का

कारण servers जोडण्याचा पुढचा प्रत्येक मार्ग — Auto Scaling (धडा 06), Kubernetes (धडा 07), Lambda (धडा 08) — प्रती सतत जोडतो आणि काढतो. एखाद्या प्रतीकडे इतर कुठेच नसलेली गोष्ट असेल, तर ती प्रत काढल्यावर ती गोष्ट हरवते: login, upload, अर्धवट बनलेला report. आधी stateless; मग scale out म्हणजे फक्त एक आकडा.

🔧 कसे (या repo मध्ये)

balance() कतरिनाचे clicks round robin ने srv-1, srv-2, srv-3 कडे पाठवते. तिचे session फक्त memory["srv-1"] मध्ये असते, आणि shared dictionary मध्ये (Redis च्या जागी). प्रत्येक click साठी प्रत्येक जागा काय सांगते ते ते print करते. हा section साधा Python आहे — ते दाखवायला sim.py model ची गरज नाही. तुमचा snippet एक failed health check आणि दुसरा user जोडतो.

🧪 करून पाहा

python3 scale/demo.py balance
python3 - <<'EOF'
import sys; sys.path.insert(0, "scale"); from sim import servers_needed
servers = ["srv-1", "srv-2", "srv-3", "srv-4"]
healthy = {"srv-1": True, "srv-2": False, "srv-3": True, "srv-4": True}   # srv-2 failed its health check
sessions = {"katrina": "logged in", "dipika": "logged in"}                 # shared store (Redis / DynamoDB)
live = [s for s in servers if healthy[s]]
for i, user in enumerate(["katrina", "dipika", "katrina", "dipika", "katrina", "dipika"]):
    s = live[i % len(live)]
    print(f"click {i + 1}: {user:<8} → {s}  session: {sessions[user]}")
print("healthy targets:", live)
print("400 req/s needs", servers_needed(400, 150), "servers at 60% · healthy capacity now", len(live) * 150, "req/s")
EOF

✅ तपासा — तुम्हाला काय दिसायला हवे

balance सहा ओळी print करते. Clicks 1 आणि 4 srv-1 वर पडतात आणि logged in म्हणतात; clicks 2, 3, 5 आणि 6 "server memory" column मध्ये NOT FOUND → login page म्हणतात — आणि प्रत्येक ओळ Redis column मध्ये logged in म्हणते:

       1  srv-1   logged in                     logged in
       2  srv-2   NOT FOUND → login page        logged in

तुमचा snippet हे print करतो:

click 1: katrina  → srv-1  session: logged in
click 2: dipika   → srv-3  session: logged in
click 3: katrina  → srv-4  session: logged in
click 4: dipika   → srv-1  session: logged in
click 5: katrina  → srv-3  session: logged in
click 6: dipika   → srv-4  session: logged in
healthy targets: ['srv-1', 'srv-3', 'srv-4']
400 req/s needs 5 servers at 60% · healthy capacity now 450 req/s

🏁 तुम्ही आत्ताच काय सिद्ध केले

Session एका server च्या memory मध्ये असताना, round robin कतरिनाला 6 clicks मध्ये 4 वेळा log out करतो. Shared store असताना प्रत्येक server उत्तर देऊ शकतो, आणि health check मध्ये अपयशी ठरलेल्या server ला फक्त clicks मिळत नाहीत. शेवटची ओळ इशारा आहे: 3 healthy servers 450 req/s पेलू शकतात, पण 400 req/s ला ते 60% च्या खूप वर आहे — आणखी एक अपयश आणि बाकीच्यांवर जास्त भार पडतो. जादा servers ठेवा.

⚠️ नेहमीच्या चुका

🏭 प्रत्यक्ष वापरात

On a real account — उथळ health check, least outstanding requests, आणि stickiness बंद असलेला target group (Terraform):

resource "aws_lb_target_group" "api" {
  name                          = "school-api"
  port                          = 8080
  protocol                      = "HTTP"
  vpc_id                        = aws_vpc.main.id
  load_balancing_algorithm_type = "least_outstanding_requests"
  deregistration_delay          = 30

  health_check {
    path                = "/health"
    interval            = 10
    timeout             = 5
    healthy_threshold   = 2
    unhealthy_threshold = 3
    matcher             = "200"
  }

  stickiness {
    type    = "lb_cookie"
    enabled = false
  }
}

आत्ता कोणते targets healthy आहेत ते पाहा:

aws elbv2 describe-target-health \
    --target-group-arn arn:aws:elasticloadbalancing:ap-south-1:111122223333:targetgroup/school-api/6d0ecf831eec9f09

ElastiCache मधले sessions (Python, redis client) — कोणताही server ते वाचू शकतो:

import json, redis
r = redis.Redis(host="school-sessions.abc123.cache.amazonaws.com", port=6379, ssl=True)
r.set(f"session:{sid}", json.dumps({"user": "katrina"}), ex=3600)   # 1 hour
user = json.loads(r.get(f"session:{sid}") or "null")

🏭 प्रत्यक्ष वापरात हे का महत्त्वाचे: ते test करा. Load balancer मागे दोन प्रती चालवा, log in करा, आणि ज्या प्रतीवर log in केले ती थांबवा. तुम्ही अजूनही logged in असाल, तर API stateless आहे.

⏭️ पुढे

कोणतीही खिडकी कोणत्याही पालकाला सेवा देऊ शकते. आता: किती खिडक्या, आणि रांग वाढली की जास्त खिडक्या कोण उघडतो? Auto Scaling.

git checkout lesson-06-auto-scaling

⚖️ Lesson 05 — Stateless APIs + load balancers: any counter can serve any parent

📍 You are here: Lesson 05 of 13 · Previous: lesson-04-dynamic-at-the-edge · Next: lesson-06-auto-scaling


📦 What's in this branch

Lessons 01–04, plus the rule that makes scaling out possible: a stateless API. Sessions, uploads and caches live outside the server, so a load balancer can send each request to any healthy server. balance() in scale/demo.py follows Katrina's six clicks across three servers — once with her session in one server's memory, once in a shared store.

🧒 Explain like I'm 5

Katrina comes to the fair and signs in at counter 1. The clerk at counter 1 writes "Katrina — signed in" in her own notebook 📓.

The helper at the gate sends each next visit to the next counter: 1, 2, 3, 1, 2, 3. Katrina's second visit goes to counter 2. That clerk looks in her notebook: no Katrina. "Please sign in again." Katrina is annoyed. Four of her six visits end like this.

Dipika fixes it with one shared notebook on a table that every counter can reach (Redis). Now any clerk can check "Katrina — signed in". Every counter can serve every parent. Now Dipika can open 3 counters or 30.

The helper at the gate also walks past every counter every few seconds and asks "are you OK?". If a clerk does not answer, the helper stops sending parents there — that is the health check.

🗺️ Diagram

flowchart LR
    k["🙋 Katrina<br/>6 clicks"] --> alb["⚖️ load balancer<br/>round robin · health checks"]
    alb --> s1["🖥️ srv-1"]
    alb --> s2["🖥️ srv-2"]
    alb --> s3["🖥️ srv-3"]
    s1 --> r["📓 shared session store<br/>ElastiCache (Redis/Valkey) · DynamoDB"]
    s2 --> r
    s3 --> r
    m["❌ session in srv-1 memory<br/>clicks 2, 3, 5, 6 → login page"]

🗺️ Drawn version + a lab: https://school-edh.pages.dev/scaling/lesson-diagrams.html#l05

❓ What

🤔 Why

Because every later way of adding servers — Auto Scaling (lesson 06), Kubernetes (lesson 07), Lambda (lesson 08) — adds and removes copies all the time. If a copy holds something that exists nowhere else, removing it loses that thing: a login, an upload, a half-built report. Stateless first; then scaling out is only a number.

🔧 How (in this repo)

balance() sends Katrina's clicks round robin to srv-1, srv-2, srv-3. Her session exists in memory["srv-1"] only, and in the shared dictionary (standing in for Redis). For each click it prints what each place says. This section is plain Python — no sim.py model is needed to show it. Your snippet adds a failed health check and a second user.

🧪 Try it

python3 scale/demo.py balance
python3 - <<'EOF'
import sys; sys.path.insert(0, "scale"); from sim import servers_needed
servers = ["srv-1", "srv-2", "srv-3", "srv-4"]
healthy = {"srv-1": True, "srv-2": False, "srv-3": True, "srv-4": True}   # srv-2 failed its health check
sessions = {"katrina": "logged in", "dipika": "logged in"}                 # shared store (Redis / DynamoDB)
live = [s for s in servers if healthy[s]]
for i, user in enumerate(["katrina", "dipika", "katrina", "dipika", "katrina", "dipika"]):
    s = live[i % len(live)]
    print(f"click {i + 1}: {user:<8} → {s}  session: {sessions[user]}")
print("healthy targets:", live)
print("400 req/s needs", servers_needed(400, 150), "servers at 60% · healthy capacity now", len(live) * 150, "req/s")
EOF

✅ Verify — what you should see

balance prints six rows. Clicks 1 and 4 land on srv-1 and say logged in; clicks 2, 3, 5 and 6 say NOT FOUND → login page in the "server memory" column — and every row says logged in in the Redis column:

       1  srv-1   logged in                     logged in
       2  srv-2   NOT FOUND → login page        logged in

Your snippet prints:

click 1: katrina  → srv-1  session: logged in
click 2: dipika   → srv-3  session: logged in
click 3: katrina  → srv-4  session: logged in
click 4: dipika   → srv-1  session: logged in
click 5: katrina  → srv-3  session: logged in
click 6: dipika   → srv-4  session: logged in
healthy targets: ['srv-1', 'srv-3', 'srv-4']
400 req/s needs 5 servers at 60% · healthy capacity now 450 req/s

🏁 What you just proved

With the session in one server's memory, round robin logs Katrina out 4 times in 6 clicks. With a shared store, every server can answer, and a server that fails its health check simply gets no clicks. The last line is a warning: 3 healthy servers can carry 450 req/s, but at 400 req/s that is far above 60% — one more failure and the rest are overloaded. Keep spare servers.

⚠️ Common mistakes

🏭 In production

On a real account — a target group with a shallow health check, least outstanding requests, and stickiness off (Terraform):

resource "aws_lb_target_group" "api" {
  name                          = "school-api"
  port                          = 8080
  protocol                      = "HTTP"
  vpc_id                        = aws_vpc.main.id
  load_balancing_algorithm_type = "least_outstanding_requests"
  deregistration_delay          = 30

  health_check {
    path                = "/health"
    interval            = 10
    timeout             = 5
    healthy_threshold   = 2
    unhealthy_threshold = 3
    matcher             = "200"
  }

  stickiness {
    type    = "lb_cookie"
    enabled = false
  }
}

See which targets are healthy right now:

aws elbv2 describe-target-health \
    --target-group-arn arn:aws:elasticloadbalancing:ap-south-1:111122223333:targetgroup/school-api/6d0ecf831eec9f09

Sessions in ElastiCache (Python, redis client) — any server can read them:

import json, redis
r = redis.Redis(host="school-sessions.abc123.cache.amazonaws.com", port=6379, ssl=True)
r.set(f"session:{sid}", json.dumps({"user": "katrina"}), ex=3600)   # 1 hour
user = json.loads(r.get(f"session:{sid}") or "null")

🏭 Why this matters in production: test it. Run two copies behind a load balancer, log in, and stop the copy you logged in on. If you are still logged in, the API is stateless.

⏭️ Next

Any counter can serve any parent. Now: how many counters, and who opens more when the queue grows? Auto Scaling.

git checkout lesson-06-auto-scaling
← Previousdynamic at the edgeNext →auto scaling

This page is the lesson's README from the lesson-05-stateless-load-balancer branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.