🏫 The School›🩺 Observability›⏱️ धडा 04 — Percentiles आणि histograms: हळू म्हणजे किती हळू
🖼️ See the drawing + lab 🏠 Course home 🌿 Branch on GitHub ✏️ View source
🖼️ आकृती आणि labThe drawing + lab पूर्ण पानावर उघडा ↗Open full page ↗

⏱️ धडा 04 — Percentiles आणि histograms: हळू म्हणजे किती हळू

📍 तुम्ही इथे आहात: 12 पैकी धडा 04 · मागे: lesson-03-metrics · पुढे: lesson-05-tracing


📦 या ब्रँचमध्ये काय आहे

धडे 01–03, आणि भाग 1 ची शेवटची कल्पना: percentiles. तुम्ही p50, p95 आणि p99 शिकता, सरासरी हळू टोक कसे लपवते, Prometheus histogram buckets मधून percentile चा अंदाज कसा काढतो, आणि percentiles ची सरासरी कधीच का काढता येत नाही. true_percentile() आणि histogram_quantile() obs/signals.py मध्ये आहेत; obs/demo.py मधील percentiles() त्यांची तुलना करते.

🧒 5 वर्षांच्या मुलाला समजावल्यासारखे

आज आरोग्य कक्षाच्या रांगेत वीस विद्यार्थ्यांनी वाट पाहिली. बहुतेकांनी सुमारे 1 किंवा 2 मिनिटे वाट पाहिली. एकाने 20 मिनिटे वाट पाहिली.

दीपिका विचारते: "विद्यार्थी किती वेळ वाट पाहतात?" कतरिनाने "सरासरी सुमारे 3 मिनिटे" असे सांगितले, तर ते ठीक वाटते — आणि 20 मिनिटे वाट पाहणारा विद्यार्थी लपून जातो.

म्हणून कतरिना विद्यार्थ्यांना सर्वात कमी वाट पाहणाऱ्यापासून सर्वात जास्त वाट पाहणाऱ्यापर्यंत रांगेत उभे करते:

आता एक सापळा. लहान वर्गांचा कक्ष सांगतो "आमचा p99 20 मिनिटे आहे". मोठ्या वर्गांचा कक्ष सांगतो "आमचा p99 1 मिनिट आहे". दोघांचा एकत्र p99 10.5 मिनिटे नाही. तो जाणण्यासाठी, दोन्ही कक्षांतले सगळे विद्यार्थी पुन्हा एकाच रांगेत उभे करून मोजावे लागतात.

🗺️ आकृती

flowchart LR
    lat["⏱️ 20 requests (ms)<br/>80 … 400 … 1200"]
    b["🗄️ buckets (le, cumulative)<br/>0.1 s: 10 · 0.25: 18 · 0.5: 19<br/>1.0: 19 · 2.5: 20"]
    q["📐 histogram_quantile(0.99)<br/>target 19.8 of 20 → bucket 1.0–2.5 s<br/>interpolate → 2,200 ms"]
    t["✅ true p99 1,200 ms"]
    lat --> b --> q
    lat --> t
    avg["🚫 average of p99s<br/>(2000 + 100) / 2 = 1050 ✗<br/>merge, then compute → 2000 ✓"]

🗺️ काढलेली आवृत्ती + एक lab: https://school-edh.pages.dev/observability/lesson-diagrams.html#l04

❓ काय

🤔 का

कारण users ना सरासरी जाणवत नाही. निकालासाठी 1.2 सेकंद वाट पाहणाऱ्या पालकाला सरासरी 0.18 s होती याचे काही देणेघेणे नसते. Percentiles सांगतात की वाईट अनुभव किती वाईट आहेत, आणि ते किती आहेत. आणि "सरासरी काढता येत नाही" हा नियम रोज महत्त्वाचा आहे: "pods मधला average p99" दाखवणारा प्रत्येक dashboard काहीच अर्थ नसलेला आकडा दाखवत असतो.

🔧 कसे (या repo मध्ये)

true_percentile(values, p) किमती sort करते आणि rank ceil(p/100 × n) वरची किंमत घेते (nearest-rank पद्धत). histogram_quantile(q, buckets, counts) Prometheus जे करतो तेच करते: bucket शोधते, त्याच्या आत interpolate करते (पहिल्या bucket साठी 0 पासून सुरुवात करून). percentiles() दिवसाच्या 20 नमुना latencies (demo.py मधील LAT) 0.1, 0.25, 0.5, 1.0 आणि 2.5 सेकंदांच्या buckets मध्ये टाकते, मग खऱ्या p50/p95/p99 ची तुलना bucket अंदाजांशी करते. ती दोन servers सुद्धा बनवते (a: 95 जलद + 5 हळू; b: 100 जलद) आणि त्यांच्या p99s च्या सरासरीची तुलना खऱ्या p99 शी करते.

🧪 करून पाहा

Bucket boundaries हलवा, मग तुमचे स्वतःचे "दोन servers" बनवा:

python3 obs/demo.py percentiles
python3 - <<'EOF'
import sys; sys.path.insert(0, "obs"); from demo import LAT; from signals import Histogram, histogram_quantile, true_percentile
for buckets in ([0.1, 0.25, 0.5, 1.0, 2.5], [0.1, 0.25, 0.5, 1.0, 1.25, 2.5], [0.05, 0.1, 0.2, 0.4, 0.8, 1.6]):
    h = Histogram("lat", "", buckets).observe_all([x / 1000 for x in LAT])
    est = [round(histogram_quantile(p / 100, h.buckets, h.counts) * 1000) for p in (50, 95, 99)]
    print(f"buckets {buckets} → p50/p95/p99 ≈ {est} ms (true {[true_percentile(LAT, p) for p in (50, 95, 99)]})")
fast, slow = [100] * 100, [100] * 90 + [3000] * 10
print(f"p99 fast server {true_percentile(fast, 99)} · slow server {true_percentile(slow, 99)} · average {(true_percentile(fast, 99) + true_percentile(slow, 99)) / 2:.0f} · real {true_percentile(fast + slow, 99)} ms")
print(f"p50 of the same two: average {(true_percentile(fast, 50) + true_percentile(slow, 50)) / 2:.0f} · real {true_percentile(fast + slow, 50)} ms · mean {sum(fast + slow) / 200:.0f} ms")
EOF

✅ तपासा — तुम्हाला काय दिसायला हवे

percentiles हे print करते:

── 20 requests (ms): 80 85 90 92 95 98 100 105 110 120 130 150 180 240 400 95 88 102 97 1200
   p50  true   100 ms · from buckets   100.0 ms
   p95  true   400 ms · from buckets   500.0 ms
   p99  true  1200 ms · from buckets  2200.0 ms
── two servers' p99: 2000 and 100 ms · their average 1050 ms · the real p99 of both: 2000 ms
   you cannot average percentiles — merge the histograms and compute again · buckets set the precision

तुमचा snippet हे print करतो:

buckets [0.1, 0.25, 0.5, 1.0, 2.5] → p50/p95/p99 ≈ [100, 500, 2200] ms (true [100, 400, 1200])
buckets [0.1, 0.25, 0.5, 1.0, 1.25, 2.5] → p50/p95/p99 ≈ [100, 500, 1200] ms (true [100, 400, 1200])
buckets [0.05, 0.1, 0.2, 0.4, 0.8, 1.6] → p50/p95/p99 ≈ [100, 400, 1440] ms (true [100, 400, 1200])
p99 fast server 100 · slow server 3000 · average 1550 · real 3000 ms
p50 of the same two: average 100 · real 100 ms · mean 245 ms

🏁 तुम्ही आत्ताच काय सिद्ध केले

Default buckets सह, खऱ्या 1,200 ms साठी p99 चा अंदाज 2,200 ms आला: लक्ष्य rank 19.8 रुंद 1.0–2.5 s bucket मध्ये पडला, आणि interpolation ने तो त्या bucket मध्ये 80% अंतरावर ठेवला (1.0 + 1.5 × 0.8 = 2.2 s). 1.25 s ला एक boundary जोडल्याने bucket अरुंद झाला (1.0 + 0.25 × 0.8 = 1.2 s) — precision buckets ठरवतात. p99s ची सरासरी 1,550 ms आली, असा आकडा जो कोणाचेच वर्णन करत नाही; एकत्र केलेला data 3,000 ms सांगतो. आणि mean (245 ms) median च्या दुपटीपेक्षा जास्त होता — दहा हळू requests नी तो वर ओढला, तर p50 जागचा हलला नाही.

⚠️ नेहमीच्या चुका

🏭 प्रत्यक्ष वापरात

On a real account — 5 मिनिटांतला प्रत्येक route चा p99, सगळ्या pods मधून एकत्र करून (PromQL):

histogram_quantile(0.99,
  sum by (le, route) (rate(http_request_duration_seconds_bucket{job="results-api"}[5m])))

0.5 s पेक्षा जलद requests चा हिस्सा — धडा 09 मध्ये वापरता येणारा SLI (त्यासाठी नेमका 0.5 वर bucket boundary लागतो):

sum(rate(http_request_duration_seconds_bucket{job="results-api", le="0.5"}[5m]))
/
sum(rate(http_request_duration_seconds_count{job="results-api"}[5m]))

Prometheus च्या नव्या आवृत्त्या native histograms सुद्धा देतात (आपोआप निवडलेले exponential buckets, कमी खर्चात खूप बारीक precision). CloudWatch percentiles extended statistics म्हणून मोजते:

aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name TargetResponseTime \
    --dimensions Name=LoadBalancer,Value=app/school-alb/0123456789abcdef \
    --start-time 2026-09-27T11:00:00Z --end-time 2026-09-27T12:00:00Z --period 60 \
    --extended-statistics p50 p99

Datadog मध्ये, distribution metric hosts पलीकडे बरोबर असणारे percentiles ठेवतो: p99:results.request.duration{service:results-api}.

🏭 Production मध्ये हे का महत्त्वाचे: तुमच्या latency SLO मर्यादेवर एक bucket boundary ठेवा, आणि p99 सोबत "मर्यादेपेक्षा जलद असलेला हिस्सा" सुद्धा वाचा. हा हिस्सा अचूक असतो; p99 हा अंदाज असतो.

⏭️ पुढे

भाग 1 पूर्ण झाला: logs, metrics, percentiles. p99 सांगतो requests हळू आहेत. पाच services मध्ये वेळ कुठे गेला? एका विद्यार्थ्याच्या मार्ग कार्डाचा पाठलाग करा.

git checkout lesson-05-tracing

⏱️ Lesson 04 — Percentiles & histograms: how slow is slow

📍 You are here: Lesson 04 of 12 · Previous: lesson-03-metrics · Next: lesson-05-tracing


📦 What's in this branch

Lessons 01–03, plus the last idea of Part 1: percentiles. You learn p50, p95 and p99, why the mean hides the slow tail, how Prometheus estimates a percentile from histogram buckets, and why you can never average percentiles. true_percentile() and histogram_quantile() live in obs/signals.py; percentiles() in obs/demo.py compares them.

🧒 Explain like I'm 5

Twenty pupils waited in the health room queue today. Most waited about 1 or 2 minutes. One waited 20 minutes.

Dipika asks: "How long do pupils wait?" If Katrina says "the average is about 3 minutes", it sounds fine — and it hides the pupil who waited 20.

So Katrina lines the pupils up from shortest wait to longest:

Now a trap. The juniors' room says "our p99 is 20 minutes". The seniors' room says "our p99 is 1 minute". Their p99 together is not 10.5 minutes. To know it, you must put all the pupils from both rooms in one line again and count.

🗺️ Diagram

flowchart LR
    lat["⏱️ 20 requests (ms)<br/>80 … 400 … 1200"]
    b["🗄️ buckets (le, cumulative)<br/>0.1 s: 10 · 0.25: 18 · 0.5: 19<br/>1.0: 19 · 2.5: 20"]
    q["📐 histogram_quantile(0.99)<br/>target 19.8 of 20 → bucket 1.0–2.5 s<br/>interpolate → 2,200 ms"]
    t["✅ true p99 1,200 ms"]
    lat --> b --> q
    lat --> t
    avg["🚫 average of p99s<br/>(2000 + 100) / 2 = 1050 ✗<br/>merge, then compute → 2000 ✓"]

🗺️ Drawn version + a lab: https://school-edh.pages.dev/observability/lesson-diagrams.html#l04

❓ What

🤔 Why

Because users do not feel the average. A parent who waits 1.2 seconds for results does not care that the average was 0.18 s. Percentiles tell you how bad the bad experiences are, and how many of them there are. And the "cannot average" rule matters every day: every dashboard that shows "average p99 across pods" is showing a number that means nothing.

🔧 How (in this repo)

true_percentile(values, p) sorts the values and takes the value at rank ceil(p/100 × n) (the nearest-rank method). histogram_quantile(q, buckets, counts) does what Prometheus does: find the bucket, interpolate inside it (starting from 0 for the first bucket). percentiles() puts the day's 20 sample latencies (LAT in demo.py) into buckets of 0.1, 0.25, 0.5, 1.0 and 2.5 seconds, then compares the true p50/p95/p99 with the bucket estimates. It also builds two servers (a: 95 fast + 5 slow; b: 100 fast) and compares the average of their p99s with the real p99.

🧪 Try it

Move the bucket boundaries, then build your own "two servers":

python3 obs/demo.py percentiles
python3 - <<'EOF'
import sys; sys.path.insert(0, "obs"); from demo import LAT; from signals import Histogram, histogram_quantile, true_percentile
for buckets in ([0.1, 0.25, 0.5, 1.0, 2.5], [0.1, 0.25, 0.5, 1.0, 1.25, 2.5], [0.05, 0.1, 0.2, 0.4, 0.8, 1.6]):
    h = Histogram("lat", "", buckets).observe_all([x / 1000 for x in LAT])
    est = [round(histogram_quantile(p / 100, h.buckets, h.counts) * 1000) for p in (50, 95, 99)]
    print(f"buckets {buckets} → p50/p95/p99 ≈ {est} ms (true {[true_percentile(LAT, p) for p in (50, 95, 99)]})")
fast, slow = [100] * 100, [100] * 90 + [3000] * 10
print(f"p99 fast server {true_percentile(fast, 99)} · slow server {true_percentile(slow, 99)} · average {(true_percentile(fast, 99) + true_percentile(slow, 99)) / 2:.0f} · real {true_percentile(fast + slow, 99)} ms")
print(f"p50 of the same two: average {(true_percentile(fast, 50) + true_percentile(slow, 50)) / 2:.0f} · real {true_percentile(fast + slow, 50)} ms · mean {sum(fast + slow) / 200:.0f} ms")
EOF

✅ Verify — what you should see

percentiles prints:

── 20 requests (ms): 80 85 90 92 95 98 100 105 110 120 130 150 180 240 400 95 88 102 97 1200
   p50  true   100 ms · from buckets   100.0 ms
   p95  true   400 ms · from buckets   500.0 ms
   p99  true  1200 ms · from buckets  2200.0 ms
── two servers' p99: 2000 and 100 ms · their average 1050 ms · the real p99 of both: 2000 ms
   you cannot average percentiles — merge the histograms and compute again · buckets set the precision

Your snippet prints:

buckets [0.1, 0.25, 0.5, 1.0, 2.5] → p50/p95/p99 ≈ [100, 500, 2200] ms (true [100, 400, 1200])
buckets [0.1, 0.25, 0.5, 1.0, 1.25, 2.5] → p50/p95/p99 ≈ [100, 500, 1200] ms (true [100, 400, 1200])
buckets [0.05, 0.1, 0.2, 0.4, 0.8, 1.6] → p50/p95/p99 ≈ [100, 400, 1440] ms (true [100, 400, 1200])
p99 fast server 100 · slow server 3000 · average 1550 · real 3000 ms
p50 of the same two: average 100 · real 100 ms · mean 245 ms

🏁 What you just proved

With the default buckets, the p99 estimate was 2,200 ms for a true 1,200 ms: the target rank 19.8 fell in the wide 1.0–2.5 s bucket, and interpolation put it 80% of the way across (1.0 + 1.5 × 0.8 = 2.2 s). Adding one boundary at 1.25 s made the bucket narrow (1.0 + 0.25 × 0.8 = 1.2 s) — buckets set the precision. Averaging p99s gave 1,550 ms, a number that describes nobody; the merged data says 3,000 ms. And the mean (245 ms) was more than twice the median — ten slow requests pulled it up while the p50 did not move.

⚠️ Common mistakes

🏭 In production

On a real account — p99 per route over 5 minutes, merged across every pod (PromQL):

histogram_quantile(0.99,
  sum by (le, route) (rate(http_request_duration_seconds_bucket{job="results-api"}[5m])))

The share of requests faster than 0.5 s — an SLI you can use in lesson 09 (it needs a bucket boundary at exactly 0.5):

sum(rate(http_request_duration_seconds_bucket{job="results-api", le="0.5"}[5m]))
/
sum(rate(http_request_duration_seconds_count{job="results-api"}[5m]))

Newer Prometheus versions also offer native histograms (exponential buckets chosen automatically, much finer precision at lower cost). CloudWatch computes percentiles as extended statistics:

aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name TargetResponseTime \
    --dimensions Name=LoadBalancer,Value=app/school-alb/0123456789abcdef \
    --start-time 2026-09-27T11:00:00Z --end-time 2026-09-27T12:00:00Z --period 60 \
    --extended-statistics p50 p99

In Datadog, a distribution metric keeps percentiles that are correct across hosts: p99:results.request.duration{service:results-api}.

🏭 Why this matters in production: put a bucket boundary at your latency SLO threshold, and read the "fraction faster than the threshold" as well as the p99. The fraction is exact; the p99 is an estimate.

⏭️ Next

Part 1 is done: logs, metrics, percentiles. The p99 says requests are slow. Where did the time go, across five services? Follow one pupil's route card.

git checkout lesson-05-tracing
← PreviousmetricsNext →tracing

This page is the lesson's README from the lesson-04-percentiles branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.