ЁЯПл The SchoolтА║ЁЯЫОя╕П API GatewayтА║ЁЯУИ рдзрдбрд╛ 11 тАФ Monitoring: dashboard рдЖрдгрд┐ stopwatch
ЁЯЦ╝я╕П See the drawing + lab ЁЯПа Course home ЁЯМ┐ Branch on GitHub тЬПя╕П View source
ЁЯЦ╝я╕П рдЖрдХреГрддреА рдЖрдгрд┐ labThe drawing + lab рдкреВрд░реНрдг рдкрд╛рдирд╛рд╡рд░ рдЙрдШрдбрд╛ тЖЧOpen full page тЖЧ

ЁЯУИ рдзрдбрд╛ 11 тАФ Monitoring: dashboard рдЖрдгрд┐ stopwatch

ЁЯУН рддреБрдореНрд╣реА рдЗрдереЗ рдЖрд╣рд╛рдд: 12 рдкреИрдХреА рдзрдбрд╛ 11 ┬╖ рдорд╛рдЧреЗ: lesson-10-custom-domains ┬╖ рдкреБрдвреЗ: lesson-12-private-waf-limits


ЁЯУж рдпрд╛ рдмреНрд░рдБрдЪрдордзреНрдпреЗ рдХрд╛рдп рдЖрд╣реЗ

рдзрдбреЗ 01тАУ10, рдЖрдгрд┐ office рдЪреНрдпрд╛ рдЖрдд рдХрд╕реЗ рдкрд╛рд╣рд╛рдпрдЪреЗ: CloudWatch metrics (Count, 4XXError, 5XXError, Latency, IntegrationLatency, CacheHitCount, CacheMissCount), $context variables рдЕрд╕рд▓реЗрд▓реЗ access logs, рдЖрдгрд┐ X-Ray traces. рдореБрдЦреНрдп рдХреМрд╢рд▓реНрдп: рд╡реЗрд│ рдХреБрдареЗ рдЬрд╛рддреЛ рддреЗ рд╢реЛрдзрдгреНрдпрд╛рд╕рд╛рдареА Latency рд╡рд┐рд░реБрджреНрдз IntegrationLatency рд╡рд╛рдЪрдгреЗ. apigw/demo.py рдордзрд▓реЗ monitor() рд╕рд╣рд╛ requests рдкрд╛рдард╡рддреЗ рдЖрдгрд┐ dashboard рд╡ visitor book рдЫрд╛рдкрддреЗ.

ЁЯзТ 5 рд╡рд░реНрд╖рд╛рдВрдЪреНрдпрд╛ рдореБрд▓рд╛рд▓рд╛ рд╕рдордЬрд╛рд╡рд▓реНрдпрд╛рд╕рд╛рд░рдЦреЗ

Front office рдЪреНрдпрд╛ рднрд┐рдВрддреАрд╡рд░ рдПрдХ dashboard ЁЯУК рдЖрд╣реЗ: рдЖрдЬ рдХрд┐рддреА visitors рдЖрд▓реЗ, рдХрд┐рддреА рдЬрдгрд╛рдВрдирд╛ рддреНрдпрд╛рдВрдЪреНрдпрд╛рдЪ рдЪреБрдХреАрдореБрд│реЗ рдкрд░рдд рдкрд╛рдард╡рд▓реЗ (4XX), рдХрд┐рддреА рдЬрдгрд╛рдВрдЪреА kitchen рдореБрд│реЗ рдирд┐рд░рд╛рд╢рд╛ рдЭрд╛рд▓реА (5XX), рдХрд┐рддреА рдЙрддреНрддрд░реЗ рд╕реВрдЪрдирд╛ рдлрд▓рдХрд╛рд╡рд░реВрди рдЖрд▓реА.

Clerk рдХрдбреЗ рджреЛрди stopwatches тП▒я╕П рд╕реБрджреНрдзрд╛ рдЖрд╣реЗрдд. рдкрд╣рд┐рд▓реЗ visitor рдЖрдд рдЖрд▓реНрдпрд╛рд╡рд░ рд╕реБрд░реВ рд╣реЛрддреЗ рдЖрдгрд┐ рддреЛ рдмрд╛рд╣реЗрд░ рдкрдбрд▓реНрдпрд╛рд╡рд░ рдерд╛рдВрдмрддреЗ тАФ рд╕рдВрдкреВрд░реНрдг рднреЗрдЯ (Latency). рджреБрд╕рд░реЗ рдлрдХреНрдд visitor kitchen рдЪреА рд╡рд╛рдЯ рдкрд╛рд╣рдд рдЕрд╕рддрд╛рдирд╛рдЪ рдЪрд╛рд▓рддреЗ (IntegrationLatency).

рд╕рдВрдкреВрд░реНрдг рднреЗрдЯреАрд▓рд╛ 42 рд╕реЗрдХрдВрдж рд▓рд╛рдЧрд▓реЗ рдЖрдгрд┐ kitchen рд▓рд╛ 38, рддрд░ kitchen рд╕рд╛рд╡рдХрд╛рд╢ рдЖрд╣реЗ. рд╕рдВрдкреВрд░реНрдг рднреЗрдЯреАрд▓рд╛ 42 рд▓рд╛рдЧрд▓реЗ рдЖрдгрд┐ kitchen рд▓рд╛ 2, рддрд░ рд╕рдорд╕реНрдпрд╛ desk рд╡рд░ рдЖрд╣реЗ тАФ рдХрджрд╛рдЪрд┐рдд badge рддрдкрд╛рд╕рдгреА рд╕рд╛рд╡рдХрд╛рд╢ рдЖрд╣реЗ.

рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХ рднреЗрдЯ рдореНрд╣рдгрдЬреЗ visitor book ЁЯУТ рдордзрд▓реА рдПрдХ рдУрд│ тАФ рдХреЛрдг, рдХреЗрд╡реНрд╣рд╛, рдХреЛрдгрддреЗ рджрд╛рд░, рдХрд╛рдп рдЙрддреНрддрд░, рдХрд┐рддреА рд╡реЗрд│.

ЁЯЧ║я╕П рдЖрдХреГрддреА

flowchart LR
    in["ЁЯЩЛ request in"] --> office["ЁЯЫОя╕П office work<br/>auth ┬╖ throttle ┬╖ mapping<br/>тЙИ 4 ms in the model"]
    office --> kitchen["ЁЯН│ kitchen<br/>IntegrationLatency 38 ms"]
    kitchen --> out["ЁЯУд answer out"]
    lat["тП▒я╕П Latency = 42 ms<br/>whole trip"] -.-> in
    lat -.-> out
    out --> cw["ЁЯУК CloudWatch<br/>Count ┬╖ 4XXError ┬╖ 5XXError ┬╖ Cache*"]
    out --> log["ЁЯУТ access log<br/>$context.requestId ┬╖ status ┬╖ latencies"]

ЁЯЧ║я╕П рдХрд╛рдврд▓реЗрд▓реА рдЖрдХреГрддреА + рдПрдХ lab: https://school-edh.pages.dev/apigateway/lesson-diagrams.html#l11

тЭУ рдХрд╛рдп

ЁЯдФ рдХрд╛

рдХрд╛рд░рдг "API рд╕рд╛рд╡рдХрд╛рд╢ рдЖрд╣реЗ" рдХрд┐рдВрд╡рд╛ "users рдирд╛ errors рдпреЗрддрд╛рдд" рд╣реЗ рддреБрдореНрд╣реА рджреБрд░реБрд╕реНрдд рдХрд░реВ рд╢рдХрдд рдирд╛рд╣реА. рддреБрдореНрд╣рд╛рд▓рд╛ рдХреЛрдгрддрд╛ route, рдХреЛрдгрддрд╛ status рдЖрдгрд┐ рдкреНрд░рд╡рд╛рд╕рд╛рдЪрд╛ рдХреЛрдгрддрд╛ рднрд╛рдЧ рд╣реЗ рдорд╛рд╣реАрдд рдЕрд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ. рджреЛрди latency metrics рдХрд╛рд╣реА рд╕реЗрдХрдВрджрд╛рдВрдд рд╕рд╛рдВрдЧрддрд╛рдд рдХреА back-end team рд▓рд╛ рдмреЛрд▓рд╡рд╛рдпрдЪреЗ рдХреА authorizer рдкрд╛рд╣рд╛рдпрдЪрд╛. requestId рдЕрд╕рд▓реЗрд▓реЗ access logs user рдЪреА рддрдХреНрд░рд╛рд░ рдПрдХрд╛ рдиреЗрдордХреНрдпрд╛ request рд╢реА рдЬреЛрдбрддрд╛рдд.

ЁЯФз рдХрд╕реЗ (рдпрд╛ repo рдордзреНрдпреЗ)

apigw/gateway.py рдордзрд▓рд╛ Gateway.handle() gw.metrics рдордзреНрдпреЗ Count, 4XXError, 5XXError, CacheHitCount рдЖрдгрд┐ CacheMissCount рдореЛрдЬрддреЛ, рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХ request рд╕рд╛рдареА requestId, stage, httpMethod, path, status, principal, latencyMs рдЖрдгрд┐ integrationLatencyMs рдЕрд╕рд▓реЗрд▓реА рдПрдХ access-log dict gw.logs рдордзреНрдпреЗ рдЬреЛрдбрддреЛ. Kitchen рдЪрд╛ рд╡реЗрд│ рдореНрд╣рдгрдЬреЗ route рдЪреА latency (/students/{id} рд╕рд╛рдареА 38 ms; /report рд╕рд╛рдареА 31 s, 29 s timeout рд▓рд╛ рдХрд╛рдкрд▓рд╛ рдЬрд╛рддреЛ), рдЖрдгрд┐ model рд╕реНрд╡рддрдГрдЪреНрдпрд╛ рдХрд╛рдорд╛рд╕рд╛рдареА OFFICE_MS = 4 рдЬреЛрдбрддреЗ тАФ рдореНрд╣рдгреВрди рдЗрдереЗ latencyMs - integrationLatencyMs рдиреЗрд╣рдореА 4 рдЕрд╕рддреЛ.

ЁЯзк рдХрд░реВрди рдкрд╛рд╣рд╛

python3 apigw/demo.py monitor
python3 - <<'EOF'
import sys; sys.path.insert(0, "apigw"); from demo import office, Clock
c = Clock(); gw, _ = office(c)
for p in ("/students/7", "/students/7", "/students/9", "/students/404", "/broken", "/report", "/teachers", "/files/a.pdf"):
    gw.handle("GET", p)
m = gw.metrics
print(m)
print(f"4XX rate {m['4XXError'] / m['Count']:.0%} ┬╖ 5XX rate {m['5XXError'] / m['Count']:.0%}")
for l in gw.logs:
    print(f"{l['requestId']} {l['path']:<14} {l['status']}  Latency {l['latencyMs']:>5} ms  Integration {l['integrationLatencyMs']:>5} ms  office {l['latencyMs'] - l['integrationLatencyMs']} ms")
EOF

тЬЕ рддрдкрд╛рд╕рд╛ тАФ рддреБрдореНрд╣рд╛рд▓рд╛ рдХрд╛рдп рджрд┐рд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ

monitor рдЫрд╛рдкрддреЗ Count=6 ┬╖ 4XXError=2 ┬╖ 5XXError=2 ┬╖ CacheHitCount=1 ┬╖ CacheMissCount=2 рдЖрдгрд┐ рд╕рд╣рд╛ log рдУрд│реА, рдЙрджрд╛. {'requestId': 'req-0001', 'stage': 'prod', 'httpMethod': 'GET', 'path': '/students/7', 'status': 200, 'principal': None, 'latencyMs': 42, 'integrationLatencyMs': 38} рдЖрдгрд┐, рд╕рд╛рд╡рдХрд╛рд╢ report рд╕рд╛рдареА, {'requestId': 'req-0005', тАж, 'path': '/report', 'status': 504, тАж, 'latencyMs': 29004, 'integrationLatencyMs': 29000}.

рддреБрдордЪрд╛ snippet рдЫрд╛рдкрддреЛ {'Count': 8, '4XXError': 2, '5XXError': 2, 'CacheHitCount': 1, 'CacheMissCount': 3}, 4XX rate 25% ┬╖ 5XX rate 25%, рдЖрдгрд┐ рдПрдХ table:

req-0001 /students/7    200  Latency    42 ms  Integration    38 ms  office 4 ms
req-0002 /students/7    200  Latency     4 ms  Integration     0 ms  office 4 ms
req-0003 /students/9    200  Latency    42 ms  Integration    38 ms  office 4 ms
req-0004 /students/404  404  Latency    42 ms  Integration    38 ms  office 4 ms
req-0005 /broken        502  Latency     4 ms  Integration     0 ms  office 4 ms
req-0006 /report        504  Latency 29004 ms  Integration 29000 ms  office 4 ms
req-0007 /teachers      404  Latency     4 ms  Integration     0 ms  office 4 ms
req-0008 /files/a.pdf   200  Latency     4 ms  Integration     0 ms  office 4 ms

Cache hit (req-0002) рд▓рд╛ kitchen рдЪрд╛ рд╡реЗрд│ рдЕрдЬрд┐рдмрд╛рдд рд▓рд╛рдЧрд▓рд╛ рдирд╛рд╣реА; 504 рдиреЗ 29 s kitchen рдордзреНрдпреЗ рдШрд╛рд▓рд╡рд▓реЗ.

ЁЯПБ рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдХрд╛рдп рд╕рд┐рджреНрдз рдХреЗрд▓реЗ

рджреЛрди рдЖрдХрдбреНрдпрд╛рдВрд╡рд░реВрди рддреБрдореНрд╣реА рд╡реЗрд│ рдХреБрдареЗ рдЧреЗрд▓рд╛ рддреЗ рд╕рд╛рдВрдЧреВ рд╢рдХрддрд╛, рдЖрдгрд┐ log рдУрд│реАрд╡рд░реВрди user рдЬреНрдпрд╛ рдПрдХрд╛ request рдмрджреНрджрд▓ рддрдХреНрд░рд╛рд░ рдХрд░рддреЛ рддреА рд╢реЛрдзреВ рд╢рдХрддрд╛ тАФ рддрд┐рдЪреНрдпрд╛ requestId рдиреЗ.

тЪая╕П рдиреЗрд╣рдореАрдЪреНрдпрд╛ рдЪреБрдХрд╛

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд

рдЦрд▒реНрдпрд╛ account рд╡рд░ тАФ JSON рдордзреНрдпреЗ access logs рдЖрдгрд┐ X-Ray рдЪрд╛рд▓реВ (Terraform), рдЖрдгрд┐ 5XXError рд╡рд░ alarm (CLI):

resource "aws_api_gateway_stage" "prod" {
  rest_api_id          = aws_api_gateway_rest_api.school.id
  deployment_id        = aws_api_gateway_deployment.school.id
  stage_name           = "prod"
  xray_tracing_enabled = true
  access_log_settings {
    destination_arn = aws_cloudwatch_log_group.school_api_access.arn
    format = jsonencode({
      requestId          = "$context.requestId"
      ip                 = "$context.identity.sourceIp"
      method             = "$context.httpMethod"
      path               = "$context.resourcePath"
      status             = "$context.status"
      latency            = "$context.responseLatency"
      integrationLatency = "$context.integrationLatency"
      principal          = "$context.authorizer.principalId"
    })
  }
}
aws cloudwatch put-metric-alarm --alarm-name school-api-prod-5xx \
    --namespace AWS/ApiGateway --metric-name 5XXError \
    --dimensions Name=ApiName,Value=school-api Name=Stage,Value=prod \
    --statistic Sum --period 60 --evaluation-periods 3 --threshold 5 \
    --comparison-operator GreaterThanOrEqualToThreshold \
    --alarm-actions arn:aws:sns:ap-south-1:111122223333:on-call

рдЧреЗрд▓реНрдпрд╛ рддрд╛рд╕рд╛рд╕рд╛рдареА рджреЛрдиреНрд╣реА stopwatches рдЪреА рддреБрд▓рдирд╛ рдХрд░рд╛:

for m in Latency IntegrationLatency; do
  aws cloudwatch get-metric-statistics --namespace AWS/ApiGateway --metric-name $m \
      --dimensions Name=ApiName,Value=school-api Name=Stage,Value=prod \
      --start-time $(date -u -v-1H +%FT%TZ) --end-time $(date -u +%FT%TZ) \
      --period 3600 --extended-statistics p99
done

(date -v-1H рд╣реЗ macOS рдЪреЗ рдЖрд╣реЗ; Linux рд╡рд░ date -u -d '-1 hour' +%FT%TZ рд╡рд╛рдкрд░рд╛.)

ЁЯПн Production рдордзреНрдпреЗ рд╣реЗ рдХрд╛ рдорд╣рддреНрддреНрд╡рд╛рдЪреЗ рдЖрд╣реЗ: рд╕рд╛рдзреНрдпрд╛ API dashboard рдордзреНрдпреЗ рдЪрд╛рд░ panels рдЕрд╕рддрд╛рдд тАФ requests, 4XX, 5XX, рдЖрдгрд┐ p99 IntegrationLatency рдЪреНрдпрд╛ рд╢реЗрдЬрд╛рд░реА p99 Latency тАФ рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХ stage рд╕рд╛рдареА 5XX рд╡ latency рд╡рд░ alarms. рддреБрдореНрд╣реА рдкрд░рдд рджреЗрдд рдЕрд╕рд▓реЗрд▓реНрдпрд╛ рдкреНрд░рддреНрдпреЗрдХ error message рдордзреНрдпреЗ requestId рдШрд╛рд▓рд╛.

тПня╕П рдкреБрдвреЗ

рддреБрдореНрд╣рд╛рд▓рд╛ рдЖрддрд╛ рд╕рдЧрд│реЗ рджрд┐рд╕рддреЗ. рд╢реЗрд╡рдЯрдЪрд╛ рдзрдбрд╛: рдХрдбрдХ рдорд░реНрдпрд╛рджрд╛ тАФ 29-second timeout, 10 MB limit, private APIs, WAF рдЖрдгрд┐ bill.

git checkout lesson-12-private-waf-limits

ЁЯУИ Lesson 11 тАФ Monitoring: the dashboard and the stopwatch

ЁЯУН You are here: Lesson 11 of 12 ┬╖ Previous: lesson-10-custom-domains ┬╖ Next: lesson-12-private-waf-limits


ЁЯУж What's in this branch

Lessons 01тАУ10, plus how to see inside the office: the CloudWatch metrics (Count, 4XXError, 5XXError, Latency, IntegrationLatency, CacheHitCount, CacheMissCount), access logs with $context variables, and X-Ray traces. The key skill: read Latency vs IntegrationLatency to find where the time goes. monitor() in apigw/demo.py sends six requests and prints the dashboard and the visitor book.

ЁЯзТ Explain like I'm 5

The front office has a dashboard ЁЯУК on the wall: how many visitors today, how many were sent away for their own mistakes (4XX), how many were let down by a kitchen (5XX), how many answers came from the notice board.

The clerk also has two stopwatches тП▒я╕П. The first starts when the visitor walks in and stops when they leave тАФ the whole visit (Latency). The second runs only while the visitor waits for the kitchen (IntegrationLatency).

If the whole visit took 42 seconds and the kitchen took 38, the kitchen is slow. If the whole visit took 42 and the kitchen took 2, the problem is at the desk тАФ maybe the badge check is slow.

And every visit is one line in the visitor book ЁЯУТ тАФ who, when, which door, what answer, how long.

ЁЯЧ║я╕П Diagram

flowchart LR
    in["ЁЯЩЛ request in"] --> office["ЁЯЫОя╕П office work<br/>auth ┬╖ throttle ┬╖ mapping<br/>тЙИ 4 ms in the model"]
    office --> kitchen["ЁЯН│ kitchen<br/>IntegrationLatency 38 ms"]
    kitchen --> out["ЁЯУд answer out"]
    lat["тП▒я╕П Latency = 42 ms<br/>whole trip"] -.-> in
    lat -.-> out
    out --> cw["ЁЯУК CloudWatch<br/>Count ┬╖ 4XXError ┬╖ 5XXError ┬╖ Cache*"]
    out --> log["ЁЯУТ access log<br/>$context.requestId ┬╖ status ┬╖ latencies"]

ЁЯЧ║я╕П Drawn version + a lab: https://school-edh.pages.dev/apigateway/lesson-diagrams.html#l11

тЭУ What

ЁЯдФ Why

Because "the API is slow" or "users get errors" is not something you can fix. You need to know which route, which status and which part of the trip. The two latency metrics tell you in seconds whether to call the back-end team or look at the authorizer. Access logs with the requestId connect a user's complaint to one exact request.

ЁЯФз How (in this repo)

Gateway.handle() in apigw/gateway.py counts Count, 4XXError, 5XXError, CacheHitCount and CacheMissCount in gw.metrics, and appends one access-log dict per request to gw.logs with requestId, stage, httpMethod, path, status, principal, latencyMs and integrationLatencyMs. The kitchen's time is the route's latency (38 ms for /students/{id}; 31 s for /report, cut at the 29 s timeout), and the model adds OFFICE_MS = 4 for its own work тАФ so latencyMs - integrationLatencyMs is always 4 here.

ЁЯзк Try it

python3 apigw/demo.py monitor
python3 - <<'EOF'
import sys; sys.path.insert(0, "apigw"); from demo import office, Clock
c = Clock(); gw, _ = office(c)
for p in ("/students/7", "/students/7", "/students/9", "/students/404", "/broken", "/report", "/teachers", "/files/a.pdf"):
    gw.handle("GET", p)
m = gw.metrics
print(m)
print(f"4XX rate {m['4XXError'] / m['Count']:.0%} ┬╖ 5XX rate {m['5XXError'] / m['Count']:.0%}")
for l in gw.logs:
    print(f"{l['requestId']} {l['path']:<14} {l['status']}  Latency {l['latencyMs']:>5} ms  Integration {l['integrationLatencyMs']:>5} ms  office {l['latencyMs'] - l['integrationLatencyMs']} ms")
EOF

тЬЕ Verify тАФ what you should see

monitor prints Count=6 ┬╖ 4XXError=2 ┬╖ 5XXError=2 ┬╖ CacheHitCount=1 ┬╖ CacheMissCount=2 and six log lines, e.g. {'requestId': 'req-0001', 'stage': 'prod', 'httpMethod': 'GET', 'path': '/students/7', 'status': 200, 'principal': None, 'latencyMs': 42, 'integrationLatencyMs': 38} and, for the slow report, {'requestId': 'req-0005', тАж, 'path': '/report', 'status': 504, тАж, 'latencyMs': 29004, 'integrationLatencyMs': 29000}.

Your snippet prints {'Count': 8, '4XXError': 2, '5XXError': 2, 'CacheHitCount': 1, 'CacheMissCount': 3}, 4XX rate 25% ┬╖ 5XX rate 25%, and a table:

req-0001 /students/7    200  Latency    42 ms  Integration    38 ms  office 4 ms
req-0002 /students/7    200  Latency     4 ms  Integration     0 ms  office 4 ms
req-0003 /students/9    200  Latency    42 ms  Integration    38 ms  office 4 ms
req-0004 /students/404  404  Latency    42 ms  Integration    38 ms  office 4 ms
req-0005 /broken        502  Latency     4 ms  Integration     0 ms  office 4 ms
req-0006 /report        504  Latency 29004 ms  Integration 29000 ms  office 4 ms
req-0007 /teachers      404  Latency     4 ms  Integration     0 ms  office 4 ms
req-0008 /files/a.pdf   200  Latency     4 ms  Integration     0 ms  office 4 ms

The cache hit (req-0002) has no kitchen time at all; the 504 spent 29 s in the kitchen.

ЁЯПБ What you just proved

From two numbers you can say where the time went, and from the log line you can find the one request a user complains about тАФ by its requestId.

тЪая╕П Common mistakes

ЁЯПн In production

On a real account тАФ access logs in JSON and X-Ray on (Terraform), and an alarm on 5XXError (CLI):

resource "aws_api_gateway_stage" "prod" {
  rest_api_id          = aws_api_gateway_rest_api.school.id
  deployment_id        = aws_api_gateway_deployment.school.id
  stage_name           = "prod"
  xray_tracing_enabled = true
  access_log_settings {
    destination_arn = aws_cloudwatch_log_group.school_api_access.arn
    format = jsonencode({
      requestId          = "$context.requestId"
      ip                 = "$context.identity.sourceIp"
      method             = "$context.httpMethod"
      path               = "$context.resourcePath"
      status             = "$context.status"
      latency            = "$context.responseLatency"
      integrationLatency = "$context.integrationLatency"
      principal          = "$context.authorizer.principalId"
    })
  }
}
aws cloudwatch put-metric-alarm --alarm-name school-api-prod-5xx \
    --namespace AWS/ApiGateway --metric-name 5XXError \
    --dimensions Name=ApiName,Value=school-api Name=Stage,Value=prod \
    --statistic Sum --period 60 --evaluation-periods 3 --threshold 5 \
    --comparison-operator GreaterThanOrEqualToThreshold \
    --alarm-actions arn:aws:sns:ap-south-1:111122223333:on-call

Compare the two stopwatches for the last hour:

for m in Latency IntegrationLatency; do
  aws cloudwatch get-metric-statistics --namespace AWS/ApiGateway --metric-name $m \
      --dimensions Name=ApiName,Value=school-api Name=Stage,Value=prod \
      --start-time $(date -u -v-1H +%FT%TZ) --end-time $(date -u +%FT%TZ) \
      --period 3600 --extended-statistics p99
done

(date -v-1H is macOS; on Linux use date -u -d '-1 hour' +%FT%TZ.)

ЁЯПн Why this matters in production: a basic API dashboard has four panels тАФ requests, 4XX, 5XX, p99 Latency next to p99 IntegrationLatency тАФ and alarms on 5XX and latency per stage. Put the requestId in every error message you return.

тПня╕П Next

You can see everything. Last lesson: the hard edges тАФ the 29-second timeout, the 10 MB limit, private APIs, WAF and the bill.

git checkout lesson-12-private-waf-limits
тЖР Previouscustom domainsNext тЖТprivate waf limits

This page is the lesson's README from the lesson-11-monitoring branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.