ЁЯПл The SchoolтА║ЁЯй║ ObservabilityтА║ЁЯУИ рдзрдбрд╛ 07 тАФ Prometheus рдЖрдгрд┐ Grafana: scrape рдХрд░рд╛, рд╕рд╛рдард╡рд╛, рд╡рд┐рдЪрд╛рд░рд╛, рдХрд╛рдврд╛
ЁЯЦ╝я╕П See the drawing + lab ЁЯПа Course home ЁЯМ┐ Branch on GitHub тЬПя╕П View source
ЁЯЦ╝я╕П рдЖрдХреГрддреА рдЖрдгрд┐ labThe drawing + lab рдкреВрд░реНрдг рдкрд╛рдирд╛рд╡рд░ рдЙрдШрдбрд╛ тЖЧOpen full page тЖЧ

ЁЯУИ рдзрдбрд╛ 07 тАФ Prometheus рдЖрдгрд┐ Grafana: scrape рдХрд░рд╛, рд╕рд╛рдард╡рд╛, рд╡рд┐рдЪрд╛рд░рд╛, рдХрд╛рдврд╛

ЁЯУН рддреБрдореНрд╣реА рдЗрдереЗ рдЖрд╣рд╛рдд: 12 рдкреИрдХреА рдзрдбрд╛ 07 ┬╖ рдорд╛рдЧреЗ: lesson-06-opentelemetry ┬╖ рдкреБрдвреЗ: lesson-08-alerting


ЁЯУж рдпрд╛ рдмреНрд░рдБрдЪрдордзреНрдпреЗ рдХрд╛рдп рдЖрд╣реЗ

рдзрдбреЗ 01тАУ06, рдЖрдгрд┐ рд╕рд░реНрд╡рд╛рдд рдкреНрд░рдЪрд▓рд┐рдд open-source metrics stack: Prometheus (scrape рдХрд░рдгреЗ рдЖрдгрд┐ рд╕рд╛рдард╡рдгреЗ, PromQL рдиреЗ рд╡рд┐рдЪрд╛рд░рдгреЗ) рдЖрдгрд┐ Grafana (рдЪрд┐рддреНрд░ рдХрд╛рдврдгреЗ). рддреБрдореНрд╣реА pull model, scrape interval, rate() рд╣реЗ рддреБрдореНрд╣реА type рдХрд░рддрд╛ рддреЗ рдкрд╣рд┐рд▓реЗ function рдХрд╛ рдЕрд╕рддреЗ, recording rules, рдЖрдгрд┐ Alertmanager рдЪреА рдкрд╣рд┐рд▓реА рдУрд│рдЦ рд╢рд┐рдХрддрд╛. obs/signals.py рдордзреАрд▓ rate() counter reset Prometheus рдкреНрд░рдорд╛рдгреЗрдЪ рд╣рд╛рддрд╛рд│рддреЗ; obs/demo.py рдордзреАрд▓ prometheus() рддреЗ рдЖрдгрд┐ рд░реЛрдЬрдЪреНрдпрд╛ рддреАрди queries рджрд╛рдЦрд╡рддреЗ.

ЁЯзТ 5 рд╡рд░реНрд╖рд╛рдВрдЪреНрдпрд╛ рдореБрд▓рд╛рд▓рд╛ рд╕рдордЬрд╛рд╡рд▓реНрдпрд╛рд╕рд╛рд░рдЦреЗ

рдХрддрд░рд┐рдирд╛ рдкреНрд░рддреНрдпреЗрдХ рд╡рд░реНрдЧрд╛рдд рдмрд╕реВ рд╢рдХрдд рдирд╛рд╣реА. рдореНрд╣рдгреВрди рджрд░ 15 рд╕реЗрдХрдВрджрд╛рдВрдиреА рддрд┐рдЪреА рдорджрддрдиреАрд╕ clipboard ЁЯУЛ рдШреЗрдКрди рд╡реНрд╣рд░рд╛рдВрдбреНрдпрд╛рддреВрди рдлрд┐рд░рддреЗ рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХ рд╡рд░реНрдЧрд╛рдЪреНрдпрд╛ рджрд╛рд░рд╛рд╡рд░рдЪрд╛ рдореЛрдЬрдгреА рдХрд╛рдЧрдж рд╡рд╛рдЪрддреЗ. рддреА рдЖрдХрдбрд╛ рдЖрдгрд┐ рд╡реЗрд│ рд▓рд┐рд╣реВрди рдШреЗрддреЗ. рд╣реЗ рдореНрд╣рдгрдЬреЗ pulling: рдорджрддрдиреАрд╕ рдЦреЛрд▓реНрдпрд╛рдВрдХрдбреЗ рдЬрд╛рддреЗ; рдЦреЛрд▓реНрдпрд╛рдВрдирд╛ рдХрд╛рд╣реАрд╣реА рдкрд╛рдард╡рд╛рд╡реЗ рд▓рд╛рдЧрдд рдирд╛рд╣реА. рдПрдЦрд╛рджреНрдпрд╛ рджрд╛рд░рд╛рд╡рд░ рдХрд╛рдЧрджрдЪ рдирд╕реЗрд▓, рдХрд┐рдВрд╡рд╛ рдЦреЛрд▓реАрд▓рд╛ рдХреБрд▓реВрдк рдЕрд╕реЗрд▓, рддрд░ рдорджрддрдиреАрд╕ рддреЗрд╣реА рдиреЛрдВрджрд╡рддреЗ тАФ рд╣рд░рд╡рд▓реЗрд▓реА рдЦреЛрд▓реА рд╣реА рд╕реБрджреНрдзрд╛ рдмрд╛рддрдореА рдЕрд╕рддреЗ.

рдПрдХрд╛ рд╕рдХрд╛рд│реА 3A рдЪрд╛ рдореЛрдЬрдгреА рдХрд╛рдЧрдж 2,200 рд╡рд░реВрди 150 рд╡рд░ рдЬрд╛рддреЛ. 2,050 рд╡рд┐рджреНрдпрд╛рд░реНрдереА рдирд┐рдШреВрди рдЧреЗрд▓реЗ рдХрд╛? рдирд╛рд╣реА тАФ рдХрд╛рдЧрдж рдЦрд╛рд▓реА рдкрдбрд▓рд╛ рдЖрдгрд┐ рдирд╡рд╛ рдХрд╛рдЧрдж рд╢реВрдиреНрдпрд╛рдкрд╛рд╕реВрди рд╕реБрд░реВ рдЭрд╛рд▓рд╛. рд╣реБрд╢рд╛рд░ рдорджрддрдиреАрд╕рд▓рд╛ рдорд╛рд╣реАрдд рдЕрд╕рддреЗ рдХреА рдореЛрдЬрдгреА рдХрдзреАрдЪ рдХрдореА рд╣реЛрдд рдирд╛рд╣реА, рдореНрд╣рдгреВрди рддреА рддреНрдпрд╛ 150 рдЦреБрдгрд╛ рдирд╡реНрдпрд╛ рдореНрд╣рдгреВрди рдореЛрдЬрддреЗ.

рдордЧ рджреАрдкрд┐рдХрд╛ рдПрдХрд╛ рдЫреЛрдЯреНрдпрд╛, рдиреЗрдордХреНрдпрд╛ рднрд╛рд╖реЗрдд рдкреНрд░рд╢реНрди рд╡рд┐рдЪрд╛рд░рддреЗ: "рдЧреЗрд▓реНрдпрд╛ 5 рдорд┐рдирд┐рдЯрд╛рдВрдд, рдкреНрд░рддреНрдпреЗрдХ рд╡рд░реНрдЧрд╛рд╕рд╛рдареА, рджрд░ рд╕реЗрдХрдВрджрд╛рд▓рд╛ рдХрд┐рддреА рдЦреБрдгрд╛?" рдЖрдгрд┐ рдЙрддреНрддрд░реЗ рд╢рд┐рдХреНрд╖рдХ рдХрдХреНрд╖рд╛рддрд▓реНрдпрд╛ рдПрдХрд╛ рдореЛрдареНрдпрд╛ рдлрд▓рдХрд╛рд╡рд░ рд░реЗрд╖рд╛ рдореНрд╣рдгреВрди рдХрд╛рдврд▓реА рдЬрд╛рддрд╛рдд тАФ рддреЛ рдлрд▓рдХ рдореНрд╣рдгрдЬреЗ Grafana.

ЁЯЧ║я╕П рдЖрдХреГрддреА

flowchart LR
    subgraph targets["ЁЯОп targets (each serves /metrics)"]
      a1["results-api pod 1"]
      a2["results-api pod 2"]
    end
    p["ЁЯФе Prometheus<br/>scrape every 15 s (pull)<br/>TSDB ┬╖ rules"]
    p -->|"GET /metrics"| a1
    p -->|"GET /metrics"| a2
    g["ЁЯУК Grafana<br/>dashboards"] -->|"PromQL"| p
    p -->|"firing alerts"| am["ЁЯЪи Alertmanager<br/>group ┬╖ route ┬╖ silence"]
    r["counter: 1000, 1600, 2200, 150 (restart), 750<br/>rate() = 8.1/s тАФ not negative"]

ЁЯЧ║я╕П рдХрд╛рдврд▓реЗрд▓реА рдЖрд╡реГрддреНрддреА + рдПрдХ lab: https://school-edh.pages.dev/observability/lesson-diagrams.html#l07

тЭУ рдХрд╛рдп

ЁЯдФ рдХрд╛

рдХрд╛рд░рдг counter рдЪреА рдХрдЪреНрдЪреА рдХрд┐рдВрдордд graph рд╡рд░ рдирд┐рд░реБрдкрдпреЛрдЧреА рдЕрд╕рддреЗ тАФ рддреА рдлрдХреНрдд рдЪрдврдд рдЬрд╛рддреЗ тАФ рдЖрдгрд┐ рд╕рд╛рдзреЗ "рд╢реЗрд╡рдЯрдЪреЗ рд╡рдЬрд╛ рдкрд╣рд┐рд▓реЗ" рдкреНрд░рддреНрдпреЗрдХ restart рд▓рд╛ рдореЛрдбрддреЗ (deploy рдкреНрд░рддреНрдпреЗрдХ pod restart рдХрд░рддреЛ). rate() counter рдЪреЗ рддреБрдореНрд╣рд╛рд▓рд╛ рд╣рд╡реНрдпрд╛ рдЕрд╕рд▓реЗрд▓реНрдпрд╛ рдЖрдХрдбреНрдпрд╛рдд рд░реВрдкрд╛рдВрддрд░ рдХрд░рддреЗ, "рджрд░ рд╕реЗрдХрдВрджрд╛рд▓рд╛, рдЖрддреНрддрд╛", рдЖрдгрд┐ restarts рдордзреВрди рдЯрд┐рдХрддреЗ. Pull model рдореБрд│реЗ рдХреЛрдгрддреЗ targets рдЕрд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗрдд рд╣реЗ monitoring system рд▓рд╛ рдорд╛рд╣реАрдд рдЕрд╕рддреЗ, рдореНрд╣рдгреВрди рд╢рд╛рдВрддрддреЗрдорд╛рдЧреЗ рдмрдВрдж рдкрдбрд▓реЗрд▓реА service рд▓рдкреВ рд╢рдХрдд рдирд╛рд╣реА.

ЁЯФз рдХрд╕реЗ (рдпрд╛ repo рдордзреНрдпреЗ)

obs/signals.py рдордзреАрд▓ rate(samples) [(time, value), тАж] рдШреЗрддреЗ, рд╢реЗрдЬрд╛рд░рдЪреНрдпрд╛ samples рдордзрд▓реА рд╡рд╛рдв рдмреЗрд░реАрдЬ рдХрд░рддреЗ тАФ рдЖрдгрд┐ рдПрдЦрд╛рджреА рдХрд┐рдВрдордд рдХрдореА рдЭрд╛рд▓реА, рддрд░ рдЛрдг рдЖрдХрдбреНрдпрд╛рдРрд╡рдЬреА рдирд╡реА рдХрд┐рдВрдордд (restart рдирдВрддрд░рдЪреА рдореЛрдЬрдгреА) рдЬреЛрдбрддреЗ тАФ рдордЧ рддреА рдЭрд╛рдХрд▓реЗрд▓реНрдпрд╛ рд╡реЗрд│реЗрдиреЗ рднрд╛рдЧрддреЗ. (рдЦрд░рд╛ Prometheus range рдЪреНрдпрд╛ рдХрдбрд╛рдВрдкрд░реНрдпрдВрдд extrapolate рд╕реБрджреНрдзрд╛ рдХрд░рддреЛ, рдореНрд╣рдгреВрди рддреНрдпрд╛рдЪреЗ рдЖрдХрдбреЗ рдереЛрдбреЗ рдЬрд╛рд╕реНрдд рдЕрд╕реВ рд╢рдХрддрд╛рдд.) prometheus() рддреНрдпрд╛рд▓рд╛ рджрд░ 60 s рд▓рд╛ scrape рдХреЗрд▓реЗрд▓рд╛ рдЖрдгрд┐ 180 s рд▓рд╛ restart рдЭрд╛рд▓реЗрд▓рд╛ counter рджреЗрддреЗ.

ЁЯзк рдХрд░реВрди рдкрд╛рд╣рд╛

рд╢реВрдиреНрдп, рдПрдХ рдЖрдгрд┐ рджреЛрди restarts рдордзреВрди rate() рдЪреА рд╕рд╛рдзреНрдпрд╛ "(last тИТ first) / time" рд╢реА рддреБрд▓рдирд╛ рдХрд░рд╛:

python3 obs/demo.py prometheus
python3 - <<'EOF'
import sys; sys.path.insert(0, "obs"); from signals import rate
calm = [(0, 1000), (60, 1600), (120, 2200), (180, 2800), (240, 3400)]
restart = [(0, 1000), (60, 1600), (120, 2200), (180, 150), (240, 750)]
two = [(0, 1000), (60, 1600), (120, 100), (180, 700), (240, 50)]
for name, s in (("no restart", calm), ("one restart", restart), ("two restarts", two)):
    naive = (s[-1][1] - s[0][1]) / (s[-1][0] - s[0][0])
    print(f"{name:<13} rate() {rate(s):5.1f}/s ┬╖ naive (last - first) / time {naive:6.1f}/s")
EOF
python3 obs/test_obs.py

тЬЕ рддрдкрд╛рд╕рд╛ тАФ рддреБрдореНрд╣рд╛рд▓рд╛ рдХрд╛рдп рджрд┐рд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ

prometheus рд╣реЗ print рдХрд░рддреЗ:

тФАтФА a counter scraped every 60 s: [(0, 1000), (60, 1600), (120, 2200), (180, 150), (240, 750)] тАФ the app restarted at 180 s
   rate over 4 min = 8.1 requests/s (the reset is handled, not a negative spike)
тФАтФА PromQL you will type every week:
   sum by (status) (rate(http_requests_total[5m]))
   sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m]))
   histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))
   Prometheus PULLS /metrics every 15тАУ60 s and stores series ┬╖ Grafana draws them ┬╖ recording rules pre-compute

рддреБрдордЪрд╛ snippet рд╣реЗ print рдХрд░рддреЛ:

no restart    rate()  10.0/s ┬╖ naive (last - first) / time   10.0/s
one restart   rate()   8.1/s ┬╖ naive (last - first) / time   -1.0/s
two restarts  rate()   5.6/s ┬╖ naive (last - first) / time   -4.0/s

Tests рдордзреНрдпреЗ тЬЕ L07 rate() survives a counter reset рдЕрд╕рддреЗ.

ЁЯПБ рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдХрд╛рдп рд╕рд┐рджреНрдз рдХреЗрд▓реЗ

Restart рдирд╕рддрд╛рдирд╛ рджреЛрдиреНрд╣реА рдкрджреНрдзрддреА рдЬреБрд│рддрд╛рдд (рджрд░ рд╕реЗрдХрдВрджрд╛рд▓рд╛ 10 requests). рдПрдХрд╛ restart рд╕рд╣ рд╕рд╛рдзреА рдкрджреНрдзрдд рджрд░ рд╕реЗрдХрдВрджрд╛рд▓рд╛ тИТ1 рд╕рд╛рдВрдЧрддреЗ тАФ counter рд╕рд╛рдареА рдЕрд╢рдХреНрдп тАФ рддрд░ rate() 8.1 рд╕рд╛рдВрдЧрддреЗ: 240 s рдордзреНрдпреЗ 600 + 600 + 150 + 600 = 1,950 requests. Reset рдирдВрддрд░рдЪреЗ 150 рдирд╡реНрдпрд╛ requests рдореНрд╣рдгреВрди рдореЛрдЬрд▓реЗ рдЧреЗрд▓реЗ. рджреЛрди restarts: рд╕рд╛рдзреА рдкрджреНрдзрдд тИТ4.0, rate() 5.6. рдкреНрд░рддреНрдпреЗрдХ deploy рдореНрд╣рдгрдЬреЗ restart, рдореНрд╣рдгреВрди рд╣реЗ рджрд░ рдЖрдард╡рдбреНрдпрд╛рд▓рд╛ рдШрдбрддреЗ.

тЪая╕П рдиреЗрд╣рдореАрдЪреНрдпрд╛ рдЪреБрдХрд╛

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд

On a real account тАФ static target, Kubernetes pod discovery, rule files рдЖрдгрд┐ Alertmanager рд╕рд╣ prometheus.yml:

global:
  scrape_interval: 15s
  evaluation_interval: 15s

rule_files: ["rules/*.yml"]

alerting:
  alertmanagers:
    - static_configs:
        - targets: ["alertmanager:9093"]

scrape_configs:
  - job_name: results-api
    static_configs:
      - targets: ["results-api:8000"]
  - job_name: kubernetes-pods
    kubernetes_sd_configs:
      - role: pod
    relabel_configs:
      - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
        action: keep
        regex: "true"

Recording rules (rules/results-api.yml) тАФ error ratio рдЖрдгрд┐ p99 рдПрдХрджрд╛рдЪ рдЖрдзреА рдореЛрдЬреВрди рдареЗрд╡рд╛:

groups:
  - name: results-api-recording
    rules:
      - record: job:http_requests:rate5m
        expr: sum by (job) (rate(http_requests_total[5m]))
      - record: job:http_errors:ratio_rate5m
        expr: |
          sum by (job) (rate(http_requests_total{status=~"5.."}[5m]))
          /
          sum by (job) (rate(http_requests_total[5m]))
      - record: job:http_request_duration_seconds:p99_5m
        expr: histogram_quantile(0.99, sum by (job, le) (rate(http_request_duration_seconds_bucket[5m])))

Files load рдХрд░рдгреНрдпрд╛рдкреВрд░реНрд╡реА рддрдкрд╛рд╕рд╛, рдЖрдгрд┐ рдереЗрдЯ HTTP API рд▓рд╛ рд╡рд┐рдЪрд╛рд░рд╛:

promtool check config prometheus.yml
promtool check rules rules/results-api.yml
curl -s 'http://prometheus:9090/api/v1/query' --data-urlencode 'query=job:http_errors:ratio_rate5m'

Grafana тАФ provisioning file рдиреЗ Prometheus data source рдореНрд╣рдгреВрди рдЬреЛрдбрд╛ (provisioning/datasources/prometheus.yml), рдордЧ $job variable (label_values(up, job)) рдЖрдгрд┐ рддреАрди recorded series рд╡рд░рдЪреНрдпрд╛ panels рд╕рд╣ рдПрдХ RED dashboard рдмрдирд╡рд╛:

apiVersion: 1
datasources:
  - name: Prometheus
    type: prometheus
    access: proxy
    url: http://prometheus:9090
    isDefault: true

ЁЯПн Production рдордзреНрдпреЗ рд╣реЗ рдХрд╛ рдорд╣рддреНрддреНрд╡рд╛рдЪреЗ: рдХреЛрдгреАрд╣реА custom dashboard design рдХрд░рдгреНрдпрд╛рдкреВрд░реНрд╡реА, рдкреНрд░рддреНрдпреЗрдХ service рд▓рд╛ рддреЛрдЪ рдкрд╣рд┐рд▓рд╛ dashboard рджреНрдпрд╛ тАФ рджрд░ рд╕реЗрдХрдВрджрд╛рдЪреНрдпрд╛ requests, error ratio, p99, recording rules рдордзреВрди. Incident рджрд░рдореНрдпрд╛рди, рдУрд│рдЦреАрдЪрд╛ dashboard рд╣рд╛рдЪ рдЬрд▓рдж рдЕрд╕рддреЛ.

тПня╕П рдкреБрдвреЗ

Graphs рдХрд╛рдврд▓реЗ рдЖрд╣реЗрдд. рдкрд╣рд╛рдЯреЗ 3 рд╡рд╛рдЬрддрд╛ рдХреЛрдгреАрд╣реА graphs рдкрд╛рд╣рдд рдирд╛рд╣реА. Alarm рдиреЗ рдХрдзреА рдХреЛрдгрд╛рд▓рд╛ рдЙрдард╡рд╛рд╡реЗ тАФ рдЖрдгрд┐ рдХрдзреА рд╢рд╛рдВрдд рд░рд╛рд╣рд┐рд▓реЗрдЪ рдкрд╛рд╣рд┐рдЬреЗ?

git checkout lesson-08-alerting

ЁЯУИ Lesson 07 тАФ Prometheus & Grafana: scrape, store, ask, draw

ЁЯУН You are here: Lesson 07 of 12 ┬╖ Previous: lesson-06-opentelemetry ┬╖ Next: lesson-08-alerting


ЁЯУж What's in this branch

Lessons 01тАУ06, plus the most common open-source metrics stack: Prometheus (scrape and store, ask with PromQL) and Grafana (draw). You learn the pull model, the scrape interval, why rate() is the first function you type, recording rules, and a first look at Alertmanager. rate() in obs/signals.py handles a counter reset the way Prometheus does; prometheus() in obs/demo.py shows it and three everyday queries.

ЁЯзТ Explain like I'm 5

Katrina cannot sit in every classroom. So every 15 seconds her helper walks the corridor with a clipboard ЁЯУЛ and reads the tally sheet on each classroom door. She writes down the number and the time. That is pulling: the helper goes to the rooms; the rooms do not have to post anything. If a door has no sheet, or the room is locked, the helper notes that too тАФ a missing room is news.

One morning the 3A tally sheet goes from 2,200 to 150. Did 2,050 pupils leave? No тАФ the sheet fell down and a new one was started from zero. A smart helper knows a tally never goes down, so she counts the 150 as new marks.

Then Dipika asks questions in a short, exact language: "marks per second, per class, over the last 5 minutes?" And the answers are drawn as lines on a big board in the staff room тАФ that board is Grafana.

ЁЯЧ║я╕П Diagram

flowchart LR
    subgraph targets["ЁЯОп targets (each serves /metrics)"]
      a1["results-api pod 1"]
      a2["results-api pod 2"]
    end
    p["ЁЯФе Prometheus<br/>scrape every 15 s (pull)<br/>TSDB ┬╖ rules"]
    p -->|"GET /metrics"| a1
    p -->|"GET /metrics"| a2
    g["ЁЯУК Grafana<br/>dashboards"] -->|"PromQL"| p
    p -->|"firing alerts"| am["ЁЯЪи Alertmanager<br/>group ┬╖ route ┬╖ silence"]
    r["counter: 1000, 1600, 2200, 150 (restart), 750<br/>rate() = 8.1/s тАФ not negative"]

ЁЯЧ║я╕П Drawn version + a lab: https://school-edh.pages.dev/observability/lesson-diagrams.html#l07

тЭУ What

ЁЯдФ Why

Because the counter's raw value is useless on a graph тАФ it only climbs тАФ and a naive "last minus first" breaks at every restart (a deploy restarts every pod). rate() turns a counter into the number you want, "per second, now", and survives restarts. The pull model makes the monitoring system the one that knows which targets should exist, so silence cannot hide a dead service.

ЁЯФз How (in this repo)

rate(samples) in obs/signals.py takes [(time, value), тАж], adds up the increase between neighbouring samples тАФ and when a value drops, adds the new value (the count since the restart) instead of a negative number тАФ then divides by the time covered. (Real Prometheus also extrapolates to the edges of the range, so its numbers can be slightly higher.) prometheus() feeds it a counter scraped every 60 s that restarted at 180 s.

ЁЯзк Try it

Compare rate() with a naive "(last тИТ first) / time" through zero, one and two restarts:

python3 obs/demo.py prometheus
python3 - <<'EOF'
import sys; sys.path.insert(0, "obs"); from signals import rate
calm = [(0, 1000), (60, 1600), (120, 2200), (180, 2800), (240, 3400)]
restart = [(0, 1000), (60, 1600), (120, 2200), (180, 150), (240, 750)]
two = [(0, 1000), (60, 1600), (120, 100), (180, 700), (240, 50)]
for name, s in (("no restart", calm), ("one restart", restart), ("two restarts", two)):
    naive = (s[-1][1] - s[0][1]) / (s[-1][0] - s[0][0])
    print(f"{name:<13} rate() {rate(s):5.1f}/s ┬╖ naive (last - first) / time {naive:6.1f}/s")
EOF
python3 obs/test_obs.py

тЬЕ Verify тАФ what you should see

prometheus prints:

тФАтФА a counter scraped every 60 s: [(0, 1000), (60, 1600), (120, 2200), (180, 150), (240, 750)] тАФ the app restarted at 180 s
   rate over 4 min = 8.1 requests/s (the reset is handled, not a negative spike)
тФАтФА PromQL you will type every week:
   sum by (status) (rate(http_requests_total[5m]))
   sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m]))
   histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))
   Prometheus PULLS /metrics every 15тАУ60 s and stores series ┬╖ Grafana draws them ┬╖ recording rules pre-compute

Your snippet prints:

no restart    rate()  10.0/s ┬╖ naive (last - first) / time   10.0/s
one restart   rate()   8.1/s ┬╖ naive (last - first) / time   -1.0/s
two restarts  rate()   5.6/s ┬╖ naive (last - first) / time   -4.0/s

The tests include тЬЕ L07 rate() survives a counter reset.

ЁЯПБ What you just proved

With no restart both methods agree (10 requests a second). With one restart the naive method says тИТ1 per second тАФ impossible for a counter тАФ while rate() says 8.1: 600 + 600 + 150 + 600 = 1,950 requests in 240 s. The 150 after the reset was counted as new requests. Two restarts: naive тИТ4.0, rate() 5.6. Every deploy is a restart, so this happens every week.

тЪая╕П Common mistakes

ЁЯПн In production

On a real account тАФ prometheus.yml with a static target, Kubernetes pod discovery, rule files and Alertmanager:

global:
  scrape_interval: 15s
  evaluation_interval: 15s

rule_files: ["rules/*.yml"]

alerting:
  alertmanagers:
    - static_configs:
        - targets: ["alertmanager:9093"]

scrape_configs:
  - job_name: results-api
    static_configs:
      - targets: ["results-api:8000"]
  - job_name: kubernetes-pods
    kubernetes_sd_configs:
      - role: pod
    relabel_configs:
      - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
        action: keep
        regex: "true"

Recording rules (rules/results-api.yml) тАФ pre-compute the error ratio and p99 once:

groups:
  - name: results-api-recording
    rules:
      - record: job:http_requests:rate5m
        expr: sum by (job) (rate(http_requests_total[5m]))
      - record: job:http_errors:ratio_rate5m
        expr: |
          sum by (job) (rate(http_requests_total{status=~"5.."}[5m]))
          /
          sum by (job) (rate(http_requests_total[5m]))
      - record: job:http_request_duration_seconds:p99_5m
        expr: histogram_quantile(0.99, sum by (job, le) (rate(http_request_duration_seconds_bucket[5m])))

Check the files before you load them, and ask the HTTP API directly:

promtool check config prometheus.yml
promtool check rules rules/results-api.yml
curl -s 'http://prometheus:9090/api/v1/query' --data-urlencode 'query=job:http_errors:ratio_rate5m'

Grafana тАФ add Prometheus as a data source by provisioning file (provisioning/datasources/prometheus.yml), then build a RED dashboard with a $job variable (label_values(up, job)) and panels on the three recorded series:

apiVersion: 1
datasources:
  - name: Prometheus
    type: prometheus
    access: proxy
    url: http://prometheus:9090
    isDefault: true

ЁЯПн Why this matters in production: give every service the same first dashboard тАФ requests per second, error ratio, p99, from recording rules тАФ before anyone designs a custom one. During an incident, the familiar dashboard is the fast one.

тПня╕П Next

The graphs are drawn. Nobody watches graphs at 3 a.m. When should the alarm wake someone тАФ and when must it stay quiet?

git checkout lesson-08-alerting
тЖР PreviousopentelemetryNext тЖТalerting

This page is the lesson's README from the lesson-07-prometheus-grafana branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.