🏫 The School›🏢 APIs›17 · 📈 Performance आणि monitoring — काउंटरवरचे स्टॉपवॉच
🖼️ See the drawing + lab 🏠 Course home 🌿 Branch on GitHub ✏️ View source
🖼️ आकृती आणि labThe drawing + lab पूर्ण पानावर उघडा ↗Open full page ↗

17 · 📈 Performance आणि monitoring — काउंटरवरचे स्टॉपवॉच

भाग 4 — production APIs. धडे 01–16 ने काउंटर बांधला आणि त्याभोवतीचे सगळे नकाशात मांडले. भाग 4 तो खऱ्या traffic खाली चालवतो: production जसे मोजते तसे मोजा, आणि टीम्स खरोखर वापरतात त्या साधनांनी त्यावर लक्ष ठेवा.

📦 या ब्रँचमध्ये काय आहे

धडे 01–17. काउंटर आता स्वतःलाच मोजतो:

🧒 5 वर्षांच्या मुलाला समजावल्यासारखे

मुख्याध्यापिका काउंटरवर एक स्टॉपवॉच ठेवतात आणि प्रत्येक पाहुण्याने किती वेळ वाट पाहिली ते लिहून ठेवतात. दिवसाच्या शेवटी त्या वेळा सर्वात कमी ते सर्वात जास्त अशा रांगेत लावतात.

सरासरी सगळे बेरीज करून भागाकार करते — आणि ती दुर्दैवी पाहुण्यांना लपवते. आमच्या run मध्ये report route ची सरासरी 18.9 ms होती, तर p95 होता 135.8 ms: बहुतेक पाहुण्यांनी ~10 ms वाट पाहिली, पण सुमारे 15 पैकी 1 जणाने ~150 ms. सरासरी म्हणाली "ठीक आहे". पाहुणे तसे म्हणाले नाहीत.

🗺️ आकृती

flowchart LR
    c["🧑‍💻 callers"] --> api["🏢 counter<br/>times every request<br/>per route"]
    api -->|"GET /metrics (pull)"| prom["📊 Prometheus → Grafana"]
    api -->|"DogStatsD over UDP (push)"| dd["🐶 Datadog agent → Datadog"]
    api -->|"EMF log line (push)"| cw["☁️ CloudWatch Logs → metric"]
    prom --> al["🔔 one alarm: p95 &gt; 100 ms for 15 min"]
    dd --> al
    cw --> al

❓ काय

🤔 का

कारण लोकांना tail जाणवते, सरासरी नव्हे. 93 callers साठी जलद आणि 7 साठी हळू असलेला route त्या 7 जणांना बिघडलेला वाटतो — आणि तेच सोडून जातात. प्रत्येक route साठी percentiles मोजणे हाच त्यांनी तक्रार लिहिण्याआधी त्यांना शोधण्याचा मार्ग आहे.

🔧 कसे (या repo मध्ये) — आकडे प्रवास करण्याचे तीन मार्ग

साधन आकडा कसा मिळवते perf.py काय दाखवतो
📊 Prometheus + Grafana pull: दर ~15 s ने GET /metrics scrape करते school_api_request_duration_seconds_bucket{route="/v1/reports/grades",le="0.25"} 104
🐶 Datadog push: app UDP वरून local agent ला DogStatsD packets पाठवते school_api.request.duration:10.16|d|#method:GET,route:/v1/reports/grades,status:200
☁️ CloudWatch push: प्रत्येक request साठी एक Embedded Metric Format JSON log ओळ; CloudWatch तिचे metric बनवते {"_aws": {… "Metrics": [{"Name": "Latency", "Unit": "Milliseconds"}]}, "Route": "GET /v1/reports/grades", "Latency": 10.4}

तोच alarm — report route चा p95 15 मिनिटे 100 ms च्या वर — प्रत्येक साधनात:

CloudWatch  aws cloudwatch put-metric-alarm --alarm-name school-api-report-p95 --namespace SchoolAPI \
              --metric-name Latency --dimensions Name=Route,Value='GET /v1/reports/grades' --extended-statistic p95 \
              --period 300 --evaluation-periods 3 --threshold 100 --comparison-operator GreaterThanThreshold
Datadog     percentile(last_15m):p95:school_api.request.duration{route:/v1/reports/grades} > 100
Prometheus  histogram_quantile(0.95, sum by (le) (rate(school_api_request_duration_seconds_bucket{route="/v1/reports/grades"}[5m]))) > 0.1

🧪 करून पाहा

python3 api/perf.py                       # 400 requests, 8 callers
python3 api/perf.py 2000 32               # more traffic, more callers — watch p99 move
python3 api/school_api.py &               # then, in another terminal:
curl -s localhost:8080/v1/reports/grades; curl -s localhost:8080/metrics | grep reports
METRICS_EMF=1 python3 api/school_api.py 8081   # every request prints a CloudWatch EMF line on stdout

✅ तपासा — तुम्हाला काय दिसायला हवे

प्रत्येक route चा तक्ता GET /v1/reports/grades ला सर्वात वर ठेवतो, p90 सुमारे 11 ms आणि p95/p99 सुमारे 136–156 ms सह, तर student routes सुमारे 1 ms जवळ राहतात. सगळ्या routes चा p95 ~11 ms असतो आणि फक्त सगळ्या routes चा p99 ~150 ms पर्यंत उडी मारतो. histogram_quantile caller च्या p95 जवळ येतो पण नेमका त्यावर नाही — buckets ढोबळ असतात. सुमारे 400 DogStatsD packets पकडले जातात, प्रत्येक request साठी एक.

🏁 तुम्ही आत्ताच काय सिद्ध केले

तुम्ही production जसे मोजते तसे API मोजू शकता — प्रत्येक route साठी percentiles — सरासरीने लपवलेला हळू route शोधू शकता, आणि CloudWatch, Datadog आणि Prometheus मध्ये तोच SLO alarm लिहू शकता.

⚠️ नेहमीच्या चुका

🏭 प्रत्यक्ष वापरात हे का महत्त्वाचे: प्रत्येक on-call टीमकडे प्रत्येक route च्या RED metrics चा dashboard असतो आणि percentiles मध्ये लिहिलेल्या SLOs वर alerts असतात — Prometheus/Grafana, Datadog किंवा CloudWatch मध्ये (AWS शाळेचा धडा 18 आणि Kubernetes शाळेचा धडा 26 पाहा).

⏭️ पुढे

काउंटर बांधला, नकाशात मांडला आणि मोजला. अभ्यास आराखड्यात या धड्यासाठी आठवडा 6 आहे; quiz मध्ये भाग 4 आहे.

17 · 📈 Performance & monitoring — the stopwatch at the counter

Part 4 — production APIs. Lessons 01–16 built the counter and mapped everything around it. Part 4 runs it under real traffic: measure it the way production does, and watch it with the tools teams actually use.

📦 What's in this branch

Lessons 01–17. The counter now measures itself:

🧒 Explain like I'm 5

The head teacher puts a stopwatch at the counter and writes down how long every visitor waited. At the end of the day she lines the times up, shortest to longest.

The average adds everything up and divides — and it hides the unlucky visitors. In our run the report route's average was 18.9 ms while p95 was 135.8 ms: most visitors waited ~10 ms, but about 1 in 15 waited ~150 ms. The average said "fine". The visitors did not.

🗺️ Diagram

flowchart LR
    c["🧑‍💻 callers"] --> api["🏢 counter<br/>times every request<br/>per route"]
    api -->|"GET /metrics (pull)"| prom["📊 Prometheus → Grafana"]
    api -->|"DogStatsD over UDP (push)"| dd["🐶 Datadog agent → Datadog"]
    api -->|"EMF log line (push)"| cw["☁️ CloudWatch Logs → metric"]
    prom --> al["🔔 one alarm: p95 &gt; 100 ms for 15 min"]
    dd --> al
    cw --> al

❓ What

🤔 Why

Because people feel the tail, not the average. A route that is fast for 93 callers and slow for 7 feels broken to those 7 — and they are the ones who leave. Measuring per route with percentiles is how you find them before they write in.

🔧 How (in this repo) — three ways the numbers travel

Tool How it gets the number What perf.py shows
📊 Prometheus + Grafana pulls: scrapes GET /metrics every ~15 s school_api_request_duration_seconds_bucket{route="/v1/reports/grades",le="0.25"} 104
🐶 Datadog push: the app sends DogStatsD packets over UDP to the local agent school_api.request.duration:10.16|d|#method:GET,route:/v1/reports/grades,status:200
☁️ CloudWatch push: one Embedded Metric Format JSON log line per request; CloudWatch turns it into a metric {"_aws": {… "Metrics": [{"Name": "Latency", "Unit": "Milliseconds"}]}, "Route": "GET /v1/reports/grades", "Latency": 10.4}

The same alarm — p95 of the report route above 100 ms for 15 minutes — in each tool:

CloudWatch  aws cloudwatch put-metric-alarm --alarm-name school-api-report-p95 --namespace SchoolAPI \
              --metric-name Latency --dimensions Name=Route,Value='GET /v1/reports/grades' --extended-statistic p95 \
              --period 300 --evaluation-periods 3 --threshold 100 --comparison-operator GreaterThanThreshold
Datadog     percentile(last_15m):p95:school_api.request.duration{route:/v1/reports/grades} > 100
Prometheus  histogram_quantile(0.95, sum by (le) (rate(school_api_request_duration_seconds_bucket{route="/v1/reports/grades"}[5m]))) > 0.1

🧪 Try it

python3 api/perf.py                       # 400 requests, 8 callers
python3 api/perf.py 2000 32               # more traffic, more callers — watch p99 move
python3 api/school_api.py &               # then, in another terminal:
curl -s localhost:8080/v1/reports/grades; curl -s localhost:8080/metrics | grep reports
METRICS_EMF=1 python3 api/school_api.py 8081   # every request prints a CloudWatch EMF line on stdout

✅ Verify — what you should see

The per-route table puts GET /v1/reports/grades on top with p90 around 11 ms and p95/p99 around 136–156 ms, while the student routes stay near 1 ms. The all-routes p95 is ~11 ms and only the all-routes p99 jumps to ~150 ms. histogram_quantile lands near the caller's p95 but not exactly on it — buckets are coarse. About 400 DogStatsD packets are caught, one per request.

🏁 What you just proved

You can measure an API the way production does — percentiles per route — find the slow route the average hid, and write the same SLO alarm in CloudWatch, Datadog and Prometheus.

⚠️ Common mistakes

🏭 Why this matters in production: every on-call team has a dashboard of RED metrics per route and alerts on SLOs written in percentiles — in Prometheus/Grafana, Datadog or CloudWatch (see the AWS school's lesson 18 and the Kubernetes school's lesson 26).

⏭️ Next

The counter is built, mapped and measured. The study plan has week 6 for this lesson; the quiz has a Part 4.

← Previoustesting typesFinished! Take the quiz →check what stuck

This page is the lesson's README from the lesson-17-performance-monitoring branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.