🏫 The School›🏎️ Performance›⏱️ धडा 01 — Latency आणि percentiles: अंतिम रेषेवरची स्टॉपवॉच
🖼️ See the drawing + lab 🏠 Course home 🌿 Branch on GitHub ✏️ View source
🖼️ आकृती आणि labThe drawing + lab पूर्ण पानावर उघडा ↗Open full page ↗

⏱️ धडा 01 — Latency आणि percentiles: अंतिम रेषेवरची स्टॉपवॉच

📍 तुम्ही इथे आहात: 12 पैकी धडा 01 · पुढे: lesson-02-benchmarking


📦 या ब्रँचमध्ये काय आहे

Performance च्या कामाचा पहिला प्रश्न: हे खरोखर किती वेगवान आहे? आज क्रीडा दिन आहे. प्रत्येक request म्हणजे एक धावपटू, आणि कतरिना हातात स्टॉपवॉच घेऊन अंतिम रेषेवर उभी आहे. एका धावपटूला लागणारा वेळ म्हणजे latency. दर second किती धावपटू पूर्ण करतात ते म्हणजे throughput. आणि बहुतेक लोक जो एक आकडा सांगतात — सरासरी — तो सर्वात महत्त्वाच्या धावपटूंना लपवतो: हळू धावणाऱ्यांना. संपूर्ण कोर्समध्ये तुम्ही वापराल त्या खऱ्या files:

🎒 सुरू करण्याआधी: तुम्हाला फक्त Python 3 लागेल, बाकी काही नाही — cloud account नाही, pip install नाही. sim.py मध्ये शिकवण्यासाठीचे models आहेत. त्यातली मोजमापाची साधने खरी आहेत (cProfile, tracemalloc, SQLite चे EXPLAIN QUERY PLAN), पण खर्च आणि वेळा simulated आणि seeded आहेत, म्हणून प्रत्येक run तेच आकडे छापतो. जेव्हा एखादा धडा खरी घड्याळाची वेळ दाखवतो, तेव्हा तो "तुमचे आकडे वेगळे असतील" असे सांगतो. खऱ्या account वर असे चिन्हांकित commands ना खऱ्या systems लागतात. ही शाळा एक program वेगवान बनवते; machines वाढवणे (CDNs, autoscaling, replicas) हे Scaling school चे काम आहे.

🧒 5 वर्षांच्या मुलाला समजावल्यासारखे

आज क्रीडा दिन आहे. 🏃‍♀️ एक हजार धावपटू प्रत्येकी एक फेरी धावतात. कतरिना अंतिम रेषेवर स्टॉपवॉच ⏱️ घेऊन उभी आहे आणि प्रत्येक वेळ लिहून ठेवते.

बहुतेक धावपटू सुमारे 40 मध्ये पूर्ण करतात (मैदानावर seconds; computer मध्ये milliseconds). काही जण अडखळतात आणि खूप जास्त वेळ घेतात. एक-दोघांचा बूट निसटतो आणि ते खूपच वेळ घेतात.

मुख्याध्यापिका विचारतात: "ते किती वेगाने धावले?" कतरिना सगळ्या वेळा बेरीज करून 1000 ने भागू शकते. ती म्हणजे सरासरी: 51. पण कोणीच 51 मध्ये धावले नाही! बहुतेक सुमारे 40 मध्ये धावले, आणि काहींना 1000 पेक्षा जास्त लागले. सरासरी त्यांना मिसळून असा आकडा बनवते जो कोणाचेच वर्णन करत नाही.

म्हणून कतरिना वेळा सर्वात जलद ते सर्वात हळू अशा रांगेत लावते आणि रांगेतल्या काही ठिकाणी वाचते:

आणि एक वेगळा प्रश्न: दर second किती धावपटू रेषा ओलांडतात? तो म्हणजे throughput. एकाऐवजी चार lanes उघडा, आणि दर second चौपट धावपटू पूर्ण करतात — पण प्रत्येक धावपटूची फेरी जराही जलद होत नाही.

🗺️ आकृती

flowchart LR
    runs["🏃‍♀️ 1000 requests<br/>timed at the finish"] --> sort["📋 sort the times<br/>fastest → slowest"]
    sort --> p50["p50 = 40 ms<br/>the middle runner"]
    sort --> p95["p95 = 59 ms"]
    sort --> p99["p99 = 260 ms<br/>the tail starts"]
    sort --> max["max = 1364 ms"]
    runs --> mean["mean = 51.3 ms<br/>describes nobody"]
    fan["🌐 a page waits for 100 calls"] --> tail["63.4% of page loads<br/>meet a p99-slow call"]

🗺️ काढलेली आकृती + एक lab: https://school-edh.pages.dev/performance/lesson-diagrams.html#l01

❓ काय

हा धडा मुद्दाम लहान आहे. Production मध्ये percentiles कसे गोळा करायचे — histograms, metrics, traces आणि SLOs — ते Observability school मध्ये आहे.

🤔 का

कारण पुढच्या प्रत्येक धड्याला "हे जलद झाले का?" याचे प्रामाणिक उत्तर हवे. तुम्ही mean सांगितलात, तर tail ला मदत करणारा उपाय काहीच नसल्यासारखा दिसू शकतो, आणि शंभरात एका user ला वाईट रीतीने त्रास देणारा बदल विजयासारखा दिसू शकतो. Users ना हळू page लक्षात राहते, आणि busy user अनेक requests करतो, म्हणून जवळजवळ प्रत्येक user कधी ना कधी तुमच्या p99 ला भेटतोच.

🔧 कसे (या repo मध्ये)

perf/sim.py मधले lap_times(n, seed) 1000 seeded latencies बनवते: 95% 20 ते 60 ms दरम्यान, 4% 100 ते 300 ms दरम्यान, आणि सुमारे 1% 800 ते 1500 ms दरम्यान. percentile(xs, p) हा nearest-rank percentile आहे. summary(xs) mean, p50, p95, p99 आणि max देते. fan_out_tail(k) म्हणजे 1 − 0.99ᵏ. perf/demo.py मधले latency() हे सगळे छापते.

🧪 करून पाहा

python3 perf/demo.py latency
python3 - <<'EOF'
import sys; sys.path.insert(0, "perf"); from sim import lap_times, percentile, summary
xs = lap_times()
for p in (50, 90, 95, 99, 99.9):
    print(f"p{p:<4} {percentile(xs, p):>5} ms")
print("5 slowest:", sorted(xs)[-5:])
quick = [x for x in xs if x < 100]
print(f"without the slow {1000 - len(quick)}:", summary(quick))
EOF
python3 perf/test_perf.py

✅ तपासा — तुम्हाला काय दिसायला हवे

latency हे छापते:

── 1000 requests timed by the finish-line stopwatches (ms per request = latency)
   mean 51.3 · p50 40 · p95 59 · p99 260 · max 1364
   37 of 1000 took 100 ms or more · 7 took 800 ms or more — the tail
   the mean (51.3) describes nobody: most finish near 40, a few take far longer
── throughput = finishers per second: one lane at 40 ms a lap → 25/s · four lanes → 100/s · each lap is still 40 ms
   a page that waits for   1 call(s): 1.0% of page loads meet at least one call slower than its p99
   a page that waits for  10 call(s): 9.6% of page loads meet at least one call slower than its p99
   a page that waits for 100 call(s): 63.4% of page loads meet at least one call slower than its p99

तुमचा snippet हे छापतो:

p50      40 ms
p90      57 ms
p95      59 ms
p99     260 ms
p99.9  1364 ms
5 slowest: [991, 1021, 1029, 1177, 1364]
without the slow 37: {'mean': 39.8, 'p50': 39, 'p95': 58, 'p99': 59, 'max': 60}

Tests शेवटी 12/12 passed छापतात.

🏁 तुम्ही आत्ताच काय सिद्ध केले

37 हळू धावपटू काढून टाका आणि mean 51.3 वरून 39.8 वर येतो — त्या 37 जणांनी mean 11 ms पेक्षा जास्त हलवला, तर median जेमतेम हलला (40 → 39). p99 नेमका जिथे हळू धावपटू सुरू होतात तिथे 59 वरून 260 वर उडी मारतो. म्हणजे mean दोन अगदी वेगळ्या गटांना मिसळतो; percentiles दोन्ही दाखवतात. आणि 100 calls च्या fan-out मध्ये, "100 पैकी 1" हळू call 63.4% page loads मध्ये दिसतो.

⚠️ नेहमीच्या चुका

🏭 प्रत्यक्ष वापरात

खऱ्या account वर — curl ने एका request च्या टप्प्यांची वेळ घ्या:

curl -o /dev/null -s -w 'dns %{time_namelookup}s · connect %{time_connect}s · tls %{time_appconnect}s · first byte %{time_starttransfer}s · total %{time_total}s\n' https://results.school.example/

Prometheus histogram सगळ्या instances वरचे percentiles देतो — आधी buckets merge करा, मग quantile काढा:

histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))

Python मध्ये, standard library एका sample चे percentiles मध्ये तुकडे करू शकते (ते interpolate करते, म्हणून लहान samples वर nearest-rank पेक्षा थोडा फरक येतो):

import statistics
cuts = statistics.quantiles(latencies_ms, n=100)   # 99 cut points
p50, p95, p99 = cuts[49], cuts[94], cuts[98]

🏭 Production मध्ये हे का महत्त्वाचे आहे: तुमची उद्दिष्टे (SLOs) percentiles वर ठरवा — उदाहरणार्थ "p99 300 ms च्या आत" — आणि p50, p95 आणि p99 शेजारी शेजारी chart करा. जेव्हा ते एकमेकांपासून दूर जातात, तेव्हा tail मागे शोधण्यासारखे कारण असते.

⏭️ पुढे

आता "जलद" म्हणजे काय हे तुम्ही सांगू शकता. पुढे: एखाद्या बदलाची प्रामाणिकपणे वेळ घेणे — warm-up फेऱ्या, अनेक heats, आणि एका आकड्याऐवजी पसारा (spread).

git checkout lesson-02-benchmarking

⏱️ Lesson 01 — Latency & percentiles: the stopwatches at the finish line

📍 You are here: Lesson 01 of 12 · Next: lesson-02-benchmarking


📦 What's in this branch

The first question of performance work: how fast is it, really? It is sports day. Every request is a runner, and Katrina stands at the finish line with a stopwatch. How long one runner takes is latency. How many runners finish each second is throughput. And the one number most people report — the average — hides the runners who matter most: the slow ones. Real files you will use all the way through:

🎒 Before you start: you need Python 3 and nothing else — no cloud account, no pip install. sim.py holds teaching models. The measuring tools in it are real (cProfile, tracemalloc, SQLite's EXPLAIN QUERY PLAN), but costs and timings are simulated and seeded, so every run prints the same numbers. When a lesson shows a real clock time, it says "your numbers will differ". Commands marked on a real account need real systems. This school makes one program fast; adding machines (CDNs, autoscaling, replicas) is the Scaling school.

🧒 Explain like I'm 5

It is sports day. 🏃‍♀️ A thousand runners run one lap each. Katrina stands at the finish line with a stopwatch ⏱️ and writes down every time.

Most runners finish in about 40 (seconds on the track; milliseconds in the computer). A few trip and take much longer. One or two lose a shoe and take forever.

The head teacher asks: "How fast were they?" Katrina could add all the times and divide by 1000. That is the average: 51. But nobody ran 51! Most ran about 40, and a few ran over 1000. The average mixes them into a number that describes nobody.

So Katrina lines the times up from fastest to slowest and reads them at points along the line:

And a different question: how many runners cross the line each second? That is throughput. Open four lanes instead of one, and four times as many finish each second — but each runner's lap is not any faster.

🗺️ Diagram

flowchart LR
    runs["🏃‍♀️ 1000 requests<br/>timed at the finish"] --> sort["📋 sort the times<br/>fastest → slowest"]
    sort --> p50["p50 = 40 ms<br/>the middle runner"]
    sort --> p95["p95 = 59 ms"]
    sort --> p99["p99 = 260 ms<br/>the tail starts"]
    sort --> max["max = 1364 ms"]
    runs --> mean["mean = 51.3 ms<br/>describes nobody"]
    fan["🌐 a page waits for 100 calls"] --> tail["63.4% of page loads<br/>meet a p99-slow call"]

🗺️ Drawn version + a lab: https://school-edh.pages.dev/performance/lesson-diagrams.html#l01

❓ What

This lesson is short on purpose. How to collect percentiles in production — histograms, metrics, traces and SLOs — is the Observability school.

🤔 Why

Because every later lesson needs an honest answer to "did it get faster?". If you report the mean, a fix that helps the tail can look like nothing, and a change that hurts one user in a hundred badly can look like a win. Users remember the slow page, and a busy user makes many requests, so almost every user meets your p99 sooner or later.

🔧 How (in this repo)

lap_times(n, seed) in perf/sim.py makes 1000 seeded latencies: 95% between 20 and 60 ms, 4% between 100 and 300 ms, and about 1% between 800 and 1500 ms. percentile(xs, p) is the nearest-rank percentile. summary(xs) gives the mean, p50, p95, p99 and max. fan_out_tail(k) is 1 − 0.99ᵏ. latency() in perf/demo.py prints them all.

🧪 Try it

python3 perf/demo.py latency
python3 - <<'EOF'
import sys; sys.path.insert(0, "perf"); from sim import lap_times, percentile, summary
xs = lap_times()
for p in (50, 90, 95, 99, 99.9):
    print(f"p{p:<4} {percentile(xs, p):>5} ms")
print("5 slowest:", sorted(xs)[-5:])
quick = [x for x in xs if x < 100]
print(f"without the slow {1000 - len(quick)}:", summary(quick))
EOF
python3 perf/test_perf.py

✅ Verify — what you should see

latency prints:

── 1000 requests timed by the finish-line stopwatches (ms per request = latency)
   mean 51.3 · p50 40 · p95 59 · p99 260 · max 1364
   37 of 1000 took 100 ms or more · 7 took 800 ms or more — the tail
   the mean (51.3) describes nobody: most finish near 40, a few take far longer
── throughput = finishers per second: one lane at 40 ms a lap → 25/s · four lanes → 100/s · each lap is still 40 ms
   a page that waits for   1 call(s): 1.0% of page loads meet at least one call slower than its p99
   a page that waits for  10 call(s): 9.6% of page loads meet at least one call slower than its p99
   a page that waits for 100 call(s): 63.4% of page loads meet at least one call slower than its p99

Your snippet prints:

p50      40 ms
p90      57 ms
p95      59 ms
p99     260 ms
p99.9  1364 ms
5 slowest: [991, 1021, 1029, 1177, 1364]
without the slow 37: {'mean': 39.8, 'p50': 39, 'p95': 58, 'p99': 59, 'max': 60}

The tests end with 12/12 passed.

🏁 What you just proved

Take away the 37 slow runners and the mean drops from 51.3 to 39.8 — the 37 moved the mean by more than 11 ms, while the median barely moved (40 → 39). The p99 jumped from 59 to 260 exactly where the slow runners begin. So the mean mixes two very different groups; the percentiles show both. And with a fan-out of 100 calls, the "1 in 100" slow call shows up in 63.4% of page loads.

⚠️ Common mistakes

🏭 In production

On a real account — time one request's phases with curl:

curl -o /dev/null -s -w 'dns %{time_namelookup}s · connect %{time_connect}s · tls %{time_appconnect}s · first byte %{time_starttransfer}s · total %{time_total}s\n' https://results.school.example/

A Prometheus histogram gives percentiles over all instances — merge the buckets first, then take the quantile:

histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))

In Python, the standard library can cut a sample into percentiles (it interpolates, so small samples differ a little from nearest-rank):

import statistics
cuts = statistics.quantiles(latencies_ms, n=100)   # 99 cut points
p50, p95, p99 = cuts[49], cuts[94], cuts[98]

🏭 Why this matters in production: set your targets (SLOs) on percentiles — for example "p99 under 300 ms" — and chart p50, p95 and p99 side by side. When they move apart, the tail has a cause worth finding.

⏭️ Next

Now you can say what "fast" means. Next: timing a change honestly — warm-up laps, many heats, and a spread instead of one number.

git checkout lesson-02-benchmarking
← Course homeall lessonsNext →benchmarking

This page is the lesson's README from the lesson-01-latency-percentiles branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.