🏫 The School›📈 Scaling›⏱️ धडा 02 — Load मोजणे: बांधण्याआधी मोजा
🖼️ See the drawing + lab 🏠 Course home 🌿 Branch on GitHub ✏️ View source
🖼️ आकृती आणि labThe drawing + lab पूर्ण पानावर उघडा ↗Open full page ↗

⏱️ धडा 02 — Load मोजणे: बांधण्याआधी मोजा

📍 तुम्ही इथे आहात: 13 पैकी धडा 02 · मागे: lesson-01-why-scaling · पुढे: lesson-03-cloudfront-s3


📦 या ब्रँचमध्ये काय आहे

धडा 01, आणि त्याशिवाय प्रत्येक scaling निर्णयाची सुरुवात होते ते आकडे: percentiles (p50, p95, p99), सरासरी हळू भेटी का लपवते, Little's law (एकाच वेळी किती requests चालू असतात) आणि "एका server ला सेकंदाला किती requests" हा आकडा देणारा load test. scale/demo.py मधले measure() प्रत्येक गोष्ट दाखवते.

🧒 5 वर्षांच्या मुलाला समजावल्यासारखे

ऐश्वर्या stopwatch ⏱️ घेऊन एका खिडकीवर उभी आहे. ती 20 पालकांची वेळ मोजते.

बहुतेक पालकांना सुमारे 100 milliseconds लागतात. एका पालकाला 1,200 लागतात — तिची पावती हरवली होती आणि कारकुनाला शोधावी लागली.

ऐश्वर्याने फक्त सरासरी लिहिली तर ती "सुमारे 180" लिहील. पण कुणालाच 180 लागले नाहीत! जलद पालकांना कमी लागले, आणि त्या दुर्दैवी पालकाला खूप जास्त. सरासरी तिला लपवते.

म्हणून ऐश्वर्या वेळा सर्वात जलद ते सर्वात हळू अशा रांगेत लावते:

आणि दुसरी युक्ती: दर सेकंदाला 2,000 पालक येत असतील आणि प्रत्येक जण खिडकीवर 0.12 सेकंद थांबत असेल, तर कोणत्याही क्षणी सुमारे 240 जण खिडक्यांवर असतात. जत्रेला एकाच वेळी इतक्या खिडक्या (threads, connections) लागतात.

🗺️ आकृती

flowchart LR
    t["⏱️ 20 timings (ms)<br/>80 … 400 … 1200"] --> s["sort, fastest first"]
    s --> p50["p50 = 100 ms<br/>a normal visit"]
    s --> p95["p95 = 400 ms"]
    s --> p99["p99 = 1,200 ms<br/>the unlucky visit"]
    s --> avg["average = 182.8 ms<br/>hides the slow one"]
    ll["🧮 Little's law<br/>2,000 req/s × 0.120 s"] --> fl["240 in flight<br/>threads · connections · Lambdas"]

🗺️ काढलेली आकृती + एक lab: https://school-edh.pages.dev/scaling/lesson-diagrams.html#l02

❓ काय

🤔 का

कारण पुढच्या प्रत्येक धड्याला एक आकडा हवा: "एक server सेकंदाला किती requests करू शकतो?" (धडे 01, 06), "किती environments?" (धडा 08), "किती connections?" (धडा 10). अंदाजाने चालल्यास एकतर outage होते किंवा मोठे bill येते. आणि सरासरी म्हणून लिहिलेली लक्ष्ये हळू भेटींना लपू देतात. लक्ष्ये p95 किंवा p99 म्हणून लिहा.

🔧 कसे (या repo मध्ये)

percentile(values, p) list क्रमाने लावते आणि rank ceil(p/100 × n) वरची किंमत घेते. littles_law(rps, latency_s) rps × latency_s परत करते. measure() 20 भेटींची वेळ मोजते, p50, p95, p99 आणि सरासरी print करते, मग 2,000 req/s × 0.120 s ला Little's law लावते.

🧪 करून पाहा

python3 scale/demo.py measure
python3 - <<'EOF'
import sys; sys.path.insert(0, "scale"); from sim import percentile, littles_law
lat = [100] * 95 + [900] * 4 + [3000]
for p in (50, 95, 99, 100):
    print(f"p{p:<3} {percentile(lat, p):>5} ms")
print(f"average {sum(lat) / len(lat):.1f} ms")
for s in (0.120, 0.600, 1.200):
    print(f"2,000 req/s × {s:.3f} s → {littles_law(2000, s):.0f} in flight")
EOF

✅ तपासा — तुम्हाला काय दिसायला हवे

measure हे print करते: p50 100 ms, p95 400 ms, p99 1200 ms, मग average 182.8 ms — one 1,200 ms visit hides inside it; percentiles show it आणि ── Little's law: 2,000 req/s × 0.120 s = 240 requests in flight at once.

तुमचा snippet (100 भेटी: 95 जलद, 4 हळू, 1 खूप हळू) हे print करतो:

p50    100 ms
p95    100 ms
p99    900 ms
p100  3000 ms
average 161.0 ms
2,000 req/s × 0.120 s → 240 in flight
2,000 req/s × 0.600 s → 1200 in flight
2,000 req/s × 1.200 s → 2400 in flight

🏁 तुम्ही आत्ताच काय सिद्ध केले

p95 म्हणते "100 ms, सगळे ठीक" — पण p99 ला 900 ms च्या भेटी सापडतात, आणि सरासरी (161 ms) कोणत्याच खऱ्या भेटीशी जुळत नाही. आणि प्रत्येक request ला दहापट वेळ लागला तर दहापट जास्त requests एकाच वेळी आत थांबलेल्या असतात: हळू म्हणजे फक्त हळू नाही, त्याला प्रत्येक गोष्ट जास्त लागते.

⚠️ नेहमीच्या चुका

🏭 प्रत्यक्ष वापरात

On a real account — CloudWatch ला load balancer च्या TargetResponseTime चे p50, p95 आणि p99 मिनिटा-मिनिटाला विचारा:

aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB \
    --metric-name TargetResponseTime \
    --dimensions Name=LoadBalancer,Value=app/school-alb/50dc6c495c0c9188 \
    --extended-statistics p50 p95 p99 --period 60 \
    --start-time 2026-05-20T03:00:00Z --end-time 2026-05-20T05:00:00Z

2,000 req/s पर्यंत वाढणारा आणि p99 500 ms च्या वर गेला तर fail होणारा एक लहान k6 load test:

// results-day.js — run with: k6 run results-day.js
import http from 'k6/http';
export const options = {
  scenarios: { rise: { executor: 'ramping-arrival-rate', startRate: 40, timeUnit: '1s',
    preAllocatedVUs: 300, maxVUs: 3000,
    stages: [{ target: 2000, duration: '5m' }, { target: 2000, duration: '10m' }] } },
  thresholds: { http_req_duration: ['p(99)<500'], http_req_failed: ['rate<0.01'] },
};
export default function () { http.get('https://staging.school.example/results/3A'); }

🏭 प्रत्यक्ष वापरात हे का महत्त्वाचे: मोठ्या दिवसाआधी production च्या प्रतीवर load test चालवा, cloud मधल्या machines वरून, गर्दी खरोखर उघडणार आहे त्या pages वर. तुम्हाला मिळणारा आकडा या कोर्समधल्या प्रत्येक उदाहरणाच्या आकड्याची जागा घेतो.

⏭️ पुढे

आता तुम्ही मोजू शकता. काढायला सर्वात सोपा load म्हणजे तुमच्या servers पर्यंत कधी पोहोचतच नाही तो: प्रत्येक दारावर झेरॉक्स प्रती — CloudFront आणि S3.

git checkout lesson-03-cloudfront-s3

⏱️ Lesson 02 — Measuring load: count before you build

📍 You are here: Lesson 02 of 13 · Previous: lesson-01-why-scaling · Next: lesson-03-cloudfront-s3


📦 What's in this branch

Lesson 01, plus the numbers every scaling decision starts from: percentiles (p50, p95, p99), why the average hides slow visits, Little's law (how many requests are in flight at once) and the load test that gives you "requests per second per server". measure() in scale/demo.py shows each one.

🧒 Explain like I'm 5

Aishwarya stands at a counter with a stopwatch ⏱️. She times 20 parents.

Most parents take about 100 milliseconds. One parent takes 1,200 — she lost her receipt and the clerk had to search.

If Aishwarya only writes the average, she writes "about 180". Nobody took 180! The fast parents took less, and the unlucky parent took far more. The average hides her.

So Aishwarya lines the times up from fastest to slowest:

And a second trick: if 2,000 parents arrive every second and each one stays at a counter for 0.12 seconds, then about 240 are at a counter at any moment. That is how many counters (threads, connections) the fair needs at once.

🗺️ Diagram

flowchart LR
    t["⏱️ 20 timings (ms)<br/>80 … 400 … 1200"] --> s["sort, fastest first"]
    s --> p50["p50 = 100 ms<br/>a normal visit"]
    s --> p95["p95 = 400 ms"]
    s --> p99["p99 = 1,200 ms<br/>the unlucky visit"]
    s --> avg["average = 182.8 ms<br/>hides the slow one"]
    ll["🧮 Little's law<br/>2,000 req/s × 0.120 s"] --> fl["240 in flight<br/>threads · connections · Lambdas"]

🗺️ Drawn version + a lab: https://school-edh.pages.dev/scaling/lesson-diagrams.html#l02

❓ What

🤔 Why

Because every later lesson needs a number: "how many requests per second can one server do?" (lessons 01, 06), "how many environments?" (lesson 08), "how many connections?" (lesson 10). Guessing gives you either an outage or a large bill. And because targets written as averages let slow visits hide. Write targets as p95 or p99.

🔧 How (in this repo)

percentile(values, p) sorts the list and takes the value at rank ceil(p/100 × n). littles_law(rps, latency_s) returns rps × latency_s. measure() times 20 visits, prints p50, p95, p99 and the average, then applies Little's law to 2,000 req/s × 0.120 s.

🧪 Try it

python3 scale/demo.py measure
python3 - <<'EOF'
import sys; sys.path.insert(0, "scale"); from sim import percentile, littles_law
lat = [100] * 95 + [900] * 4 + [3000]
for p in (50, 95, 99, 100):
    print(f"p{p:<3} {percentile(lat, p):>5} ms")
print(f"average {sum(lat) / len(lat):.1f} ms")
for s in (0.120, 0.600, 1.200):
    print(f"2,000 req/s × {s:.3f} s → {littles_law(2000, s):.0f} in flight")
EOF

✅ Verify — what you should see

measure prints p50 100 ms, p95 400 ms, p99 1200 ms, then average 182.8 ms — one 1,200 ms visit hides inside it; percentiles show it and ── Little's law: 2,000 req/s × 0.120 s = 240 requests in flight at once.

Your snippet (100 visits: 95 fast, 4 slow, 1 very slow) prints:

p50    100 ms
p95    100 ms
p99    900 ms
p100  3000 ms
average 161.0 ms
2,000 req/s × 0.120 s → 240 in flight
2,000 req/s × 0.600 s → 1200 in flight
2,000 req/s × 1.200 s → 2400 in flight

🏁 What you just proved

p95 says "100 ms, all fine" — but p99 finds the 900 ms visits, and the average (161 ms) matches no real visit. And when each request takes 10 times longer, 10 times more are waiting inside at once: slow is not only slow, it also needs more of everything.

⚠️ Common mistakes

🏭 In production

On a real account — ask CloudWatch for the load balancer's p50, p95 and p99 of TargetResponseTime, minute by minute:

aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB \
    --metric-name TargetResponseTime \
    --dimensions Name=LoadBalancer,Value=app/school-alb/50dc6c495c0c9188 \
    --extended-statistics p50 p95 p99 --period 60 \
    --start-time 2026-05-20T03:00:00Z --end-time 2026-05-20T05:00:00Z

A small k6 load test that rises to 2,000 req/s and fails if p99 goes above 500 ms:

// results-day.js — run with: k6 run results-day.js
import http from 'k6/http';
export const options = {
  scenarios: { rise: { executor: 'ramping-arrival-rate', startRate: 40, timeUnit: '1s',
    preAllocatedVUs: 300, maxVUs: 3000,
    stages: [{ target: 2000, duration: '5m' }, { target: 2000, duration: '10m' }] } },
  thresholds: { http_req_duration: ['p(99)<500'], http_req_failed: ['rate<0.01'] },
};
export default function () { http.get('https://staging.school.example/results/3A'); }

🏭 Why this matters in production: run the load test on a copy of production before the big day, from machines in the cloud, against the pages the crowd will really open. The number you get replaces every example number in this course.

⏭️ Next

Now you can measure. The easiest load to remove is the load that never reaches your servers: photocopies at every gate — CloudFront and S3.

git checkout lesson-03-cloudfront-s3
← Previouswhy scalingNext →cloudfront s3

This page is the lesson's README from the lesson-02-measuring-load branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.