⏱️ धडा 04 — Percentiles आणि histograms: हळू म्हणजे किती हळू
📍 तुम्ही इथे आहात: 12 पैकी धडा 04 · मागे: lesson-03-metrics · पुढे: lesson-05-tracing
📦 या ब्रँचमध्ये काय आहे
धडे 01–03, आणि भाग 1 ची शेवटची कल्पना: percentiles. तुम्ही p50, p95 आणि p99 शिकता,
सरासरी हळू टोक कसे लपवते, Prometheus histogram buckets मधून percentile चा अंदाज कसा
काढतो, आणि percentiles ची सरासरी कधीच का काढता येत नाही. true_percentile() आणि
histogram_quantile() obs/signals.py मध्ये आहेत;
obs/demo.py मधील percentiles() त्यांची तुलना करते.
🧒 5 वर्षांच्या मुलाला समजावल्यासारखे
आज आरोग्य कक्षाच्या रांगेत वीस विद्यार्थ्यांनी वाट पाहिली. बहुतेकांनी सुमारे 1 किंवा 2 मिनिटे वाट पाहिली. एकाने 20 मिनिटे वाट पाहिली.
दीपिका विचारते: "विद्यार्थी किती वेळ वाट पाहतात?" कतरिनाने "सरासरी सुमारे 3 मिनिटे" असे सांगितले, तर ते ठीक वाटते — आणि 20 मिनिटे वाट पाहणारा विद्यार्थी लपून जातो.
म्हणून कतरिना विद्यार्थ्यांना सर्वात कमी वाट पाहणाऱ्यापासून सर्वात जास्त वाट पाहणाऱ्यापर्यंत रांगेत उभे करते:
- मधला विद्यार्थी (20 पैकी 10 वा) — तो p50: अर्ध्यांनी कमी वाट पाहिली;
- 20 पैकी 19 वा विद्यार्थी — तो p95: 100 पैकी 95 जणांनी कमी वाट पाहिली;
- 100 पैकी सर्वात हळू काही — ते p99.
आता एक सापळा. लहान वर्गांचा कक्ष सांगतो "आमचा p99 20 मिनिटे आहे". मोठ्या वर्गांचा कक्ष सांगतो "आमचा p99 1 मिनिट आहे". दोघांचा एकत्र p99 10.5 मिनिटे नाही. तो जाणण्यासाठी, दोन्ही कक्षांतले सगळे विद्यार्थी पुन्हा एकाच रांगेत उभे करून मोजावे लागतात.
🗺️ आकृती
flowchart LR
lat["⏱️ 20 requests (ms)<br/>80 … 400 … 1200"]
b["🗄️ buckets (le, cumulative)<br/>0.1 s: 10 · 0.25: 18 · 0.5: 19<br/>1.0: 19 · 2.5: 20"]
q["📐 histogram_quantile(0.99)<br/>target 19.8 of 20 → bucket 1.0–2.5 s<br/>interpolate → 2,200 ms"]
t["✅ true p99 1,200 ms"]
lat --> b --> q
lat --> t
avg["🚫 average of p99s<br/>(2000 + 100) / 2 = 1050 ✗<br/>merge, then compute → 2000 ✓"]
🗺️ काढलेली आवृत्ती + एक lab: https://school-edh.pages.dev/observability/lesson-diagrams.html#l04
❓ काय
- Latency — एका request ला किती वेळ लागतो, शक्य असेल तर user च्या बाजूने मोजलेला.
- Mean (सरासरी) — बेरीज भागिले संख्या. एक खूप हळू request तिला थोडी हलवते; दहा खूप हळू requests तिच्यात पूर्णपणे लपून जातात.
- Percentile (quantile) — pN म्हणजे अशी किंमत की N% निरीक्षणे तिच्या बरोबरीची किंवा तिच्यापेक्षा कमी असतात. p50 म्हणजे median (मध्यक); p95 आणि p99 tail चे वर्णन करतात — हळू requests. "p99 = 1.2 s" म्हणजे 100 पैकी 1 request ला 1.2 s पेक्षा जास्त वेळ लागला. 30 calls करणाऱ्या page वर, सुमारे 4 पैकी 1 load (1 − 0.99³⁰ ≈ 26%) किमान एका p99-हळू call ची वाट पाहतो.
- Quantile विरुद्ध percentile — तीच कल्पना वेगळ्या मापात: quantile 0.99 = 99 वा percentile.
- Percentiles ची सरासरी काढता येत नाही — दोन servers चा p99 म्हणजे त्यांच्या p99s ची सरासरी नाही. तुम्हाला कच्चा data (किंवा histogram buckets) एकत्र करून पुन्हा मोजावे लागते. वेळेनुसार p99 ची सरासरी काढण्यालाही हेच लागू होते (एका तासाचा p99 म्हणजे साठ एक-मिनिटांच्या p99s ची सरासरी नाही).
- Histograms एकत्र करता येतात — bucket counts हे फक्त counters आहेत, म्हणून सगळ्या pods चे buckets बेरीज करून मग एक percentile मोजता येतो. म्हणूनच अनेक copies असलेल्या services साठी histograms summaries पेक्षा सरस ठरतात (धडा 03).
histogram_quantile(q, buckets)— Prometheus buckets मधून percentile चा अंदाज असा काढतो:- लक्ष्य rank म्हणजे
q × total count(20 च्या p99 साठी: 19.8); - ज्याचा cumulative count लक्ष्यापर्यंत पोहोचतो असा पहिला bucket शोधा;
- त्या bucket च्या आत रेषीय interpolation करा, म्हणजे किमती त्याच्या खालच्या आणि वरच्या
मर्यादेमध्ये समान पसरलेल्या आहेत असे मानून.
म्हणून उत्तर buckets इतकेच अचूक असते. लक्ष्य
+Infbucket मध्ये पडले, तर Prometheus सर्वात वरच्या finite bucket ची वरची मर्यादा परत करतो.
- लक्ष्य rank म्हणजे
- Buckets निवडणे — boundaries तुम्हाला महत्त्वाच्या आकड्यांजवळ ठेवा: तुमची SLO मर्यादा (उदाहरणार्थ 0.5 s) आणि तुम्हाला अपेक्षित p99 च्या आसपासच्या किमती.
🤔 का
कारण users ना सरासरी जाणवत नाही. निकालासाठी 1.2 सेकंद वाट पाहणाऱ्या पालकाला सरासरी 0.18 s होती याचे काही देणेघेणे नसते. Percentiles सांगतात की वाईट अनुभव किती वाईट आहेत, आणि ते किती आहेत. आणि "सरासरी काढता येत नाही" हा नियम रोज महत्त्वाचा आहे: "pods मधला average p99" दाखवणारा प्रत्येक dashboard काहीच अर्थ नसलेला आकडा दाखवत असतो.
🔧 कसे (या repo मध्ये)
true_percentile(values, p) किमती sort करते आणि rank ceil(p/100 × n) वरची किंमत घेते
(nearest-rank पद्धत). histogram_quantile(q, buckets, counts) Prometheus जे करतो तेच करते:
bucket शोधते, त्याच्या आत interpolate करते (पहिल्या bucket साठी 0 पासून सुरुवात करून).
percentiles() दिवसाच्या 20 नमुना latencies (demo.py मधील LAT) 0.1, 0.25, 0.5, 1.0 आणि
2.5 सेकंदांच्या buckets मध्ये टाकते, मग खऱ्या p50/p95/p99 ची तुलना bucket अंदाजांशी करते.
ती दोन servers सुद्धा बनवते (a: 95 जलद + 5 हळू; b: 100 जलद) आणि त्यांच्या p99s च्या
सरासरीची तुलना खऱ्या p99 शी करते.
🧪 करून पाहा
Bucket boundaries हलवा, मग तुमचे स्वतःचे "दोन servers" बनवा:
python3 obs/demo.py percentiles
python3 - <<'EOF'
import sys; sys.path.insert(0, "obs"); from demo import LAT; from signals import Histogram, histogram_quantile, true_percentile
for buckets in ([0.1, 0.25, 0.5, 1.0, 2.5], [0.1, 0.25, 0.5, 1.0, 1.25, 2.5], [0.05, 0.1, 0.2, 0.4, 0.8, 1.6]):
h = Histogram("lat", "", buckets).observe_all([x / 1000 for x in LAT])
est = [round(histogram_quantile(p / 100, h.buckets, h.counts) * 1000) for p in (50, 95, 99)]
print(f"buckets {buckets} → p50/p95/p99 ≈ {est} ms (true {[true_percentile(LAT, p) for p in (50, 95, 99)]})")
fast, slow = [100] * 100, [100] * 90 + [3000] * 10
print(f"p99 fast server {true_percentile(fast, 99)} · slow server {true_percentile(slow, 99)} · average {(true_percentile(fast, 99) + true_percentile(slow, 99)) / 2:.0f} · real {true_percentile(fast + slow, 99)} ms")
print(f"p50 of the same two: average {(true_percentile(fast, 50) + true_percentile(slow, 50)) / 2:.0f} · real {true_percentile(fast + slow, 50)} ms · mean {sum(fast + slow) / 200:.0f} ms")
EOF
✅ तपासा — तुम्हाला काय दिसायला हवे
percentiles हे print करते:
── 20 requests (ms): 80 85 90 92 95 98 100 105 110 120 130 150 180 240 400 95 88 102 97 1200
p50 true 100 ms · from buckets 100.0 ms
p95 true 400 ms · from buckets 500.0 ms
p99 true 1200 ms · from buckets 2200.0 ms
── two servers' p99: 2000 and 100 ms · their average 1050 ms · the real p99 of both: 2000 ms
you cannot average percentiles — merge the histograms and compute again · buckets set the precision
तुमचा snippet हे print करतो:
buckets [0.1, 0.25, 0.5, 1.0, 2.5] → p50/p95/p99 ≈ [100, 500, 2200] ms (true [100, 400, 1200])
buckets [0.1, 0.25, 0.5, 1.0, 1.25, 2.5] → p50/p95/p99 ≈ [100, 500, 1200] ms (true [100, 400, 1200])
buckets [0.05, 0.1, 0.2, 0.4, 0.8, 1.6] → p50/p95/p99 ≈ [100, 400, 1440] ms (true [100, 400, 1200])
p99 fast server 100 · slow server 3000 · average 1550 · real 3000 ms
p50 of the same two: average 100 · real 100 ms · mean 245 ms
🏁 तुम्ही आत्ताच काय सिद्ध केले
Default buckets सह, खऱ्या 1,200 ms साठी p99 चा अंदाज 2,200 ms आला: लक्ष्य rank 19.8 रुंद 1.0–2.5 s bucket मध्ये पडला, आणि interpolation ने तो त्या bucket मध्ये 80% अंतरावर ठेवला (1.0 + 1.5 × 0.8 = 2.2 s). 1.25 s ला एक boundary जोडल्याने bucket अरुंद झाला (1.0 + 0.25 × 0.8 = 1.2 s) — precision buckets ठरवतात. p99s ची सरासरी 1,550 ms आली, असा आकडा जो कोणाचेच वर्णन करत नाही; एकत्र केलेला data 3,000 ms सांगतो. आणि mean (245 ms) median च्या दुपटीपेक्षा जास्त होता — दहा हळू requests नी तो वर ओढला, तर p50 जागचा हलला नाही.
⚠️ नेहमीच्या चुका
- "pods मधला average p99" असा dashboard panel — आधी
sum by (le)ने buckets एकत्र करा (धडा 07) - mean latency वर alert करणे — mean जवळजवळ न हलता हळू tail दुप्पट होऊ शकते
- तुमच्या SLO मर्यादेजवळ एकही boundary नसलेले default buckets
- अनेक labels असलेल्या metric वर खूप जास्त buckets — प्रत्येक bucket ही एक series असते (धडा 03)
- फक्त server वर मोजणे: पालकाच्या वाट पाहण्यात network आणि browser चा वेळही असतो
- छोट्या नमुन्यावरचा (20 requests) p99 स्थिर असल्यासारखा वाचणे
🏭 प्रत्यक्ष वापरात
On a real account — 5 मिनिटांतला प्रत्येक route चा p99, सगळ्या pods मधून एकत्र करून (PromQL):
histogram_quantile(0.99,
sum by (le, route) (rate(http_request_duration_seconds_bucket{job="results-api"}[5m])))
0.5 s पेक्षा जलद requests चा हिस्सा — धडा 09 मध्ये वापरता येणारा SLI (त्यासाठी नेमका 0.5 वर bucket boundary लागतो):
sum(rate(http_request_duration_seconds_bucket{job="results-api", le="0.5"}[5m]))
/
sum(rate(http_request_duration_seconds_count{job="results-api"}[5m]))
Prometheus च्या नव्या आवृत्त्या native histograms सुद्धा देतात (आपोआप निवडलेले exponential buckets, कमी खर्चात खूप बारीक precision). CloudWatch percentiles extended statistics म्हणून मोजते:
aws cloudwatch get-metric-statistics --namespace AWS/ApplicationELB --metric-name TargetResponseTime \
--dimensions Name=LoadBalancer,Value=app/school-alb/0123456789abcdef \
--start-time 2026-09-27T11:00:00Z --end-time 2026-09-27T12:00:00Z --period 60 \
--extended-statistics p50 p99
Datadog मध्ये, distribution metric hosts पलीकडे बरोबर असणारे percentiles ठेवतो:
p99:results.request.duration{service:results-api}.
🏭 Production मध्ये हे का महत्त्वाचे: तुमच्या latency SLO मर्यादेवर एक bucket boundary ठेवा, आणि p99 सोबत "मर्यादेपेक्षा जलद असलेला हिस्सा" सुद्धा वाचा. हा हिस्सा अचूक असतो; p99 हा अंदाज असतो.
⏭️ पुढे
भाग 1 पूर्ण झाला: logs, metrics, percentiles. p99 सांगतो requests हळू आहेत. पाच services मध्ये वेळ कुठे गेला? एका विद्यार्थ्याच्या मार्ग कार्डाचा पाठलाग करा.
git checkout lesson-05-tracing