ЁЯПл The SchoolтА║ЁЯЦея╕П EC2тА║ЁЯУИ рдзрдбрд╛ 11 тАФ Monitoring рдЖрдгрд┐ troubleshooting: рджреЛрди status checks, рджреЛрди рдЙрдкрд╛рдп
ЁЯЦ╝я╕П See the drawing + lab ЁЯПа Course home ЁЯМ┐ Branch on GitHub тЬПя╕П View source
ЁЯЦ╝я╕П рдЖрдХреГрддреА рдЖрдгрд┐ labThe drawing + lab рдкреВрд░реНрдг рдкрд╛рдирд╛рд╡рд░ рдЙрдШрдбрд╛ тЖЧOpen full page тЖЧ

ЁЯУИ рдзрдбрд╛ 11 тАФ Monitoring рдЖрдгрд┐ troubleshooting: рджреЛрди status checks, рджреЛрди рдЙрдкрд╛рдп

ЁЯУН рддреБрдореНрд╣реА рдЗрдереЗ рдЖрд╣рд╛рдд: 12 рдкреИрдХреА рдзрдбрд╛ 11 ┬╖ рдорд╛рдЧреАрд▓: lesson-10-load-balancing ┬╖ рдкреБрдвреАрд▓: lesson-12-lifecycle


ЁЯУж рдпрд╛ рдмреНрд░рдБрдЪрдордзреНрдпреЗ рдХрд╛рдп рдЖрд╣реЗ

рдзрдбреЗ 01тАУ10, рдЖрдгрд┐ рд╕рдорд╕реНрдпрд╛ рдХреЛрдгрд╛рдЪреА рдЖрд╣реЗ рд╣реЗ рдХрд╕реЗ рдУрд│рдЦрд╛рдпрдЪреЗ: system status check (AWS рдЪреА рдмрд╛рдЬреВ) рд╡рд┐рд░реБрджреНрдз instance status check (рддреБрдордЪреА рдмрд╛рдЬреВ), рдлреБрдХрдЯ рдорд┐рд│рдгрд╛рд░реЗ metrics рд╡рд┐рд░реБрджреНрдз рдлрдХреНрдд CloudWatch agent рд▓рд╛рдЪ рджрд┐рд╕рдгрд╛рд░реЗ metrics, ec2/fleet.py рдордзреАрд▓ fleet.credits() рд╡рд░реВрди рдмрдирд╡рд▓реЗрд▓рд╛ CPUCreditBalance рд╡рд░рдЪрд╛ рдПрдХ alarm, рдЖрдгрд┐ "рдорд╛рдЭреНрдпрд╛ рдбреЗрд╕реНрдХрдкрд░реНрдпрдВрдд рдкреЛрд╣реЛрдЪрддрд╛ рдпреЗрдд рдирд╛рд╣реА" рд╕рд╛рдареА troubleshooting рдЪрд╛ рдЫреЛрдЯрд╛ рдирдХрд╛рд╢рд╛.

ЁЯзТ 5 рд╡рд░реНрд╖рд╛рдВрдЪреНрдпрд╛ рдореБрд▓рд╛рд▓рд╛ рд╕рдордЬрд╛рд╡рд▓реНрдпрд╛рд╕рд╛рд░рдЦреЗ

рдкреНрд░рддреНрдпреЗрдХ рдбреЗрд╕реНрдХ рджрд░ рдорд┐рдирд┐рдЯрд╛рд▓рд╛ рджреЛрди рд╡реЗрдЧрд╡реЗрдЧрд│реЗ рдирд┐рд░реАрдХреНрд╖рдХ рддрдкрд╛рд╕рддрд╛рдд.

рдЗрдорд╛рд░рдд рдирд┐рд░реАрдХреНрд╖рдХ ЁЯПЧя╕П (system status check) рдбреЗрд╕реНрдХрдЪреНрдпрд╛ рдЦрд╛рд▓реА рдХрд╛рдп рдЖрд╣реЗ рддреЗ рдкрд╛рд╣рддреЛ: рдлрд░рд╢реА, рд╡реАрдЬреЗрдЪрд╛ socket, network cable тАФ рдХрдВрдкрдиреАрдЪреНрдпрд╛ рдорд╛рд▓рдХреАрдЪреНрдпрд╛ рдЧреЛрд╖реНрдЯреА. рддреЗ рдирд╛рдкрд╛рд╕ рдЭрд╛рд▓реЗ рддрд░ рдбреЗрд╕реНрдХрд╡рд░ рдмрд╕реВрди рддреБрдореНрд╣реА рдХрд╛рд╣реАрдЪ рджреБрд░реБрд╕реНрдд рдХрд░реВ рд╢рдХрдд рдирд╛рд╣реА. рддреБрдореНрд╣реА рдбреЗрд╕реНрдХ рджреБрд╕рд▒реНрдпрд╛ рдЬрд╛рдЧреА рд╣рд▓рд╡рддрд╛ (stop, рдордЧ start тАФ рдбреЗрд╕реНрдХ рдирд┐рд░реЛрдЧреА hardware рд╡рд░ рдкрд░рдд рдпреЗрддреЛ).

рдбреЗрд╕реНрдХ рдирд┐рд░реАрдХреНрд╖рдХ ЁЯФН (instance status check) рддреБрдордЪрд╛ рдбреЗрд╕реНрдХ рдкрд╛рд╣рддреЛ: рдХреЛрдгреА рдЙрддреНрддрд░ рджреЗрддреЗрдп рдХрд╛, рдбреНрд░реЙрд╡рд░ рдЗрддрдХрд╛ рднрд░рд▓рд╛рдп рдХрд╛ рдХреА рдХрд╛рд╣реАрдЪ рдЙрдШрдбрдд рдирд╛рд╣реА, рдХреЛрдгреА settings рдмрд┐рдШрдбрд╡рд▓реНрдпрд╛ рдХрд╛? рд╣реЗ рджреБрд░реБрд╕реНрдд рдХрд░рдгреЗ рддреБрдордЪреЗ рдХрд╛рдо: restart рдХрд░рд╛, рд╕реБрд░реВ рд╣реЛрддрд╛рдирд╛ рдбреЗрд╕реНрдХрдиреЗ рдХрд╛рдп рдЫрд╛рдкрд▓реЗ рддреЗ рд╡рд╛рдЪрд╛ (console log), рджреБрд░реБрд╕реНрдд рдХрд░рд╛.

рдЖрдгрд┐ рдХрд╛рд╣реА рдЧреЛрд╖реНрдЯреА рдХреЛрдгрддреНрдпрд╛рдЪ рдирд┐рд░реАрдХреНрд╖рдХрд╛рд▓рд╛ рдмрд╛рд╣реЗрд░реВрди рджрд┐рд╕рдд рдирд╛рд╣реАрдд тАФ рддреБрдордЪрд╛ рдбреНрд░реЙрд╡рд░ рдХрд┐рддреА рднрд░рд▓рд╛рдп, рддреБрдордЪреЗ programs рдХрд┐рддреА memory рд╡рд╛рдкрд░рддрд╛рдд. рддреНрдпрд╛рдВрдЪреНрдпрд╛рд╕рд╛рдареА рддреБрдореНрд╣реА рдбреЗрд╕реНрдХрд╡рд░ рдПрдХ рдЫреЛрдЯрд╛ reporter рдареЗрд╡рддрд╛ (CloudWatch agent) рдЬреЛ рддреЗ рд▓рд┐рд╣реВрди рдкрд╛рдард╡рддреЛ.

ЁЯЧ║я╕П рдЖрдХреГрддреА

flowchart LR
    sys["тЭМ system status check<br/>AWS hardware, power, network under the desk"]
    inst["тЭМ instance status check<br/>your OS: full disk, bad config, kernel, memory"]
    fix1["тП╣я╕П stop + start тЖТ new host<br/>or automatic recovery"]
    fix2["ЁЯФБ reboot ┬╖ read the console log ┬╖ fix the OS"]
    sys -->|"1 AWS's side"| fix1
    inst -->|"2 your side"| fix2
    free["ЁЯУК free every 5 min: CPU, network, disk I/O, CPUCreditBalance"]
    agent["ЁЯз╛ CloudWatch agent: memory, disk space used, logs"]
    alarm["ЁЯФФ alarm: CPUCreditBalance below 50<br/>fires an hour before the throttle"]
    free -->|"3"| alarm
    agent -.->|"4 what the hypervisor cannot see"| alarm

ЁЯЧ║я╕П рдХрд╛рдврд▓реЗрд▓реА рдЖрдХреГрддреА + рдПрдХ lab: https://school-edh.pages.dev/ec2/lesson-diagrams.html#l11

тЭУ рдХрд╛рдп

ЁЯдФ рдХрд╛

рдХрд╛рд░рдг incident рдЪреЗ рдкрд╣рд┐рд▓реЗ рдорд┐рдирд┐рдЯ рд╕рдорд╕реНрдпрд╛ рдХреЛрдгрд╛рдЪреА рдЖрд╣реЗ рд╣реЗ рдард░рд╡рдгреНрдпрд╛рдд рдЬрд╛рддреЗ тАФ рдЖрдгрд┐ рддреБрдореНрд╣реА terminal рдЙрдШрдбрдгреНрдпрд╛рдЖрдзреАрдЪ рджреЛрди status checks рддреНрдпрд╛рдЪреЗ рдЙрддреНрддрд░ рджреЗрддрд╛рдд. рдЖрдгрд┐ рдХрд╛рд░рдг рдЬреЗ graphs рддреБрдореНрд╣реА рдЧреЛрд│рд╛ рдХреЗрд▓реЗ рдирд╛рд╣реАрдд рддреЗрдЪ рддреБрдореНрд╣рд╛рд▓рд╛ рд╣рд╡реЗ рдЕрд╕рддрд╛рдд: memory рдЖрдгрд┐ disk space рд╣реА рдЖрдЬрд╛рд░реА рдбреЗрд╕реНрдХрдЪреА рджреЛрди рд╕рд░реНрд╡рд╛рдд рд╕рд╛рдорд╛рдиреНрдп рдХрд╛рд░рдгреЗ рдЖрд╣реЗрдд, рдЖрдгрд┐ default рдиреЗ рдпрд╛рддрд▓реЗ рдХрд╛рд╣реАрдЪ рдЧреЛрд│рд╛ рд╣реЛрдд рдирд╛рд╣реА.

ЁЯФз рдХрд╕реЗ (рдпрд╛ repo рдордзреНрдпреЗ)

ec2/demo.py рдордзреАрд▓ monitoring() рджреЛрди checks рддреНрдпрд╛рдВрдЪреНрдпрд╛ рдЙрдкрд╛рдпрд╛рдВрд╕рд╣ рдЫрд╛рдкрддреЗ, рдХрд╛рдп рдлреБрдХрдЯ рдЧреЛрд│рд╛ рд╣реЛрддреЗ рдЖрдгрд┐ рдХрд╛рдп рдирд╛рд╣реА рддреЗ рд╕рд╛рдВрдЧрддреЗ, рдЖрдгрд┐ рдордЧ рдзрдбрд╛ 02 рдордзрд▓рд╛ busy t3.micro run (fleet.credits(3)) рдкреБрдиреНрд╣рд╛ рд╡рд╛рдкрд░реВрди CPUCreditBalance < 20 рд╡рд░рдЪрд╛ alarm рдХреЛрдгрддреНрдпрд╛ рддрд╛рд╕рд╛рд▓рд╛ рд╡рд╛рдЬреЗрд▓ рддреЗ рд╢реЛрдзрддреЗ. рдЦрд╛рд▓рдЪрд╛ snippet threshold 50 рд╡рд░ рдиреЗрддреЛ рдЖрдгрд┐ рдкреВрд░реНрдг рднрд░рд▓реЗрд▓реНрдпрд╛ рдЯрд╛рдХреАрдкрд╛рд╕реВрди рд╕реБрд░реВ рдХрд░рддреЛ, рдореНрд╣рдгрдЬреЗ throttle рдЖрдзреАрдЪ рдЗрд╢рд╛рд░рд╛ рджреЗрдгрд╛рд░рд╛ alarm рджрд┐рд╕рддреЛ.

ЁЯзк рдХрд░реВрди рдкрд╛рд╣рд╛

python3 ec2/demo.py monitoring
python3 - <<'EOF'
import sys; sys.path.insert(0, "ec2"); import fleet
# a t3.micro at 40% on both vCPUs, starting with a full tank (288). When does CPUCreditBalance < 50 fire?
for h, bal, got in fleet.credits(9, busy_pct=40, start_balance=288):
    print(f"hour {h}: {bal:>5} credits ┬╖ CPU {got}%{'  ЁЯФФ ALARM: CPUCreditBalance < 50' if bal < 50 else ''}")
EOF

рдЦрд▒реНрдпрд╛ account рд╡рд░ (AWS credentials рд▓рд╛рдЧрддрд╛рдд; alarms рд╕рд╛рдареА рджрд░ рдорд╣рд┐рдиреНрдпрд╛рд▓рд╛ рдереЛрдбрд╛ рдЦрд░реНрдЪ рдпреЗрддреЛ тАФ рддреБрдордЪрд╛ рд╕реНрд╡рддрдГрдЪрд╛ instance ID рдЖрдгрд┐ SNS topic рд╡рд╛рдкрд░рд╛):

aws ec2 describe-instance-status --instance-ids i-0123456789abcdef0 --include-all-instances
aws ec2 get-console-output --instance-id i-0123456789abcdef0 --latest --output text
aws cloudwatch put-metric-alarm --alarm-name web-a-credits-low \
    --namespace AWS/EC2 --metric-name CPUCreditBalance \
    --dimensions Name=InstanceId,Value=i-0123456789abcdef0 \
    --statistic Average --period 300 --evaluation-periods 2 \
    --threshold 50 --comparison-operator LessThanThreshold \
    --alarm-actions arn:aws:sns:us-east-1:123456789012:school-oncall
aws cloudwatch put-metric-alarm --alarm-name web-a-system-check \
    --namespace AWS/EC2 --metric-name StatusCheckFailed_System \
    --dimensions Name=InstanceId,Value=i-0123456789abcdef0 \
    --statistic Maximum --period 60 --evaluation-periods 2 \
    --threshold 1 --comparison-operator GreaterThanOrEqualToThreshold \
    --alarm-actions arn:aws:automate:us-east-1:ec2:recover

тЬЕ рддрдкрд╛рд╕рд╛ тАФ рддреБрдореНрд╣рд╛рд▓рд╛ рдХрд╛рдп рджрд┐рд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ

monitoring рдЫрд╛рдкрддреЗ тЭМ system status check тЖТ AWS's hardware or network under your desk тЖТ stop + start (moves to a new host), тЭМ instance status check тЖТ YOUR operating system (full disk, bad config, kernel) тЖТ reboot, read the console log, fix, memory рдЖрдгрд┐ disk-space-used рд╕рд╛рдареА CloudWatch agent рд▓рд╛рдЧрддреЛ рд╣реА рд╕реВрдЪрдирд╛, рдЖрдгрд┐ ЁЯФФ alarm idea: CPUCreditBalance < 20 тЖТ in the busy run above it fires from hour 2. рддреБрдордЪрд╛ snippet рджрд░ рддрд╛рд╕рд╛рд▓рд╛ 36 credits рд╕рдВрдкрд╡рддреЛ тАФ 252.0, 216.0, тАж тАФ рддрд╛рд╕ 7 рд▓рд╛ ЁЯФФ ALARM рд╡рд╛рдЬрддреЛ (36.0 credits ┬╖ CPU 40%), рдЖрдгрд┐ рдлрдХреНрдд рддрд╛рд╕ 8 рд▓рд╛ CPU 10% рдкрд░реНрдпрдВрдд рдШрд╕рд░рддреЛ (0.6 credits). Alarm рдиреЗ рддреБрдореНрд╣рд╛рд▓рд╛ рдПрдХ рддрд╛рд╕ рдорд┐рд│рд╡реВрди рджрд┐рд▓рд╛.

ЁЯПБ рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдХрд╛рдп рд╕рд┐рджреНрдз рдХреЗрд▓реЗ

AWS рдЪреА рд╕рдорд╕реНрдпрд╛ рдЖрдгрд┐ рддреБрдордЪреА рд╕рдорд╕реНрдпрд╛ рддреБрдореНрд╣реА рдПрдХрд╛ рдирдЬрд░реЗрдд рдУрд│рдЦреВ рд╢рдХрддрд╛, рдХреЛрдгрддреЗ рдЖрдХрдбреЗ рдлреБрдХрдЯ рдорд┐рд│рддрд╛рдд рдЖрдгрд┐ рдХреЛрдгрддреНрдпрд╛рдВрд╕рд╛рдареА agent рд▓рд╛рдЧрддреЛ рд╣реЗ рддреБрдореНрд╣рд╛рд▓рд╛ рдорд╛рд╣реАрдд рдЖрд╣реЗ, рдЖрдгрд┐ рдХреГрддреА рдХрд░рд╛рдпрд▓рд╛ рдЕрдЬреВрди рд╡реЗрд│ рдЕрд╕рддрд╛рдирд╛рдЪ рд╡рд╛рдЬрдгрд╛рд░рд╛ alarm рддреБрдореНрд╣реА рд▓рд╛рд╡рд▓рд╛.

тЪая╕П рдиреЗрд╣рдореАрдЪреНрдпрд╛ рдЪреБрдХрд╛

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд рд╣реЗ рдХрд╛ рдорд╣рддреНрддреНрд╡рд╛рдЪреЗ: рдПрдХрд╛ рд╕рд╛рдорд╛рдиреНрдп EC2 dashboard рд╡рд░ status checks, CPU, burstable types рд╕рд╛рдареА credit balance, agent рдХрдбреВрди memory рдЖрдгрд┐ disk, рдЖрдгрд┐ load balancer рдЪрд╛ healthy-host count рд╡ 5xx rate рдЕрд╕рддрд╛рдд тАФ рдЬреНрдпрд╛рд▓рд╛ рдорд╛рдгреВрд╕ рд▓рд╛рдЧрддреЛ рддреНрдпрд╛рд╡рд░ page рдХрд░рдгрд╛рд░реЗ рдЖрдгрд┐ рдЬреНрдпрд╛рд▓рд╛ рд▓рд╛рдЧрдд рдирд╛рд╣реА рддреЗ рдЖрдкреЛрдЖрдк recover рдХрд░рдгрд╛рд░реЗ alarms рд╕рд╣.

тПня╕П рдкреБрдвреЗ

рд╢реЗрд╡рдЯрдЪрд╛ рдзрдбрд╛: рд░рд╛рддреНрд░реАрдЪреА photocopy, рд╕реЛрдиреЗрд░реА рд╕рд╛рдЪрд╛, рдЖрдгрд┐ stop, hibernate рдЖрдгрд┐ terminate рдЦрд░реЛрдЦрд░ рдХрд╛рдп рдареЗрд╡рддрд╛рдд. Backup, patching рдЖрдгрд┐ lifecycle.

git checkout lesson-12-lifecycle

ЁЯУИ Lesson 11 тАФ Monitoring & troubleshooting: two status checks, two fixes

ЁЯУН You are here: Lesson 11 of 12 ┬╖ Previous: lesson-10-load-balancing ┬╖ Next: lesson-12-lifecycle


ЁЯУж What's in this branch

Lessons 01тАУ10, plus how to tell whose problem it is: the system status check (AWS's side) vs the instance status check (your side), the free metrics vs the ones only the CloudWatch agent can see, an alarm on CPUCreditBalance built from fleet.credits() in ec2/fleet.py, and a short troubleshooting map for "I cannot reach my desk".

ЁЯзТ Explain like I'm 5

Every desk is checked every minute by two different inspectors.

The building inspector ЁЯПЧя╕П (the system status check) looks at what is under the desk: the floor, the power socket, the network cable тАФ things the company owns. If that fails, there is nothing you can fix on the desk. You move to a different desk spot (stop, then start тАФ the desk comes back on healthy hardware).

The desk inspector ЁЯФН (the instance status check) looks at your desk: is anyone answering, is the drawer so full nothing opens, did someone break the settings? That one is yours to fix: restart, read what the desk printed while starting up (the console log), repair.

And some things neither inspector can see from outside тАФ how full your drawer is, how much memory your programs use. For those, you put a small reporter on the desk (the CloudWatch agent) that writes them down.

ЁЯЧ║я╕П Diagram

flowchart LR
    sys["тЭМ system status check<br/>AWS hardware, power, network under the desk"]
    inst["тЭМ instance status check<br/>your OS: full disk, bad config, kernel, memory"]
    fix1["тП╣я╕П stop + start тЖТ new host<br/>or automatic recovery"]
    fix2["ЁЯФБ reboot ┬╖ read the console log ┬╖ fix the OS"]
    sys -->|"1 AWS's side"| fix1
    inst -->|"2 your side"| fix2
    free["ЁЯУК free every 5 min: CPU, network, disk I/O, CPUCreditBalance"]
    agent["ЁЯз╛ CloudWatch agent: memory, disk space used, logs"]
    alarm["ЁЯФФ alarm: CPUCreditBalance below 50<br/>fires an hour before the throttle"]
    free -->|"3"| alarm
    agent -.->|"4 what the hypervisor cannot see"| alarm

ЁЯЧ║я╕П Drawn version + a lab: https://school-edh.pages.dev/ec2/lesson-diagrams.html#l11

тЭУ What

ЁЯдФ Why

Because the first minute of an incident is spent deciding whose problem it is тАФ and the two status checks answer that before you open a terminal. And because the graphs you did not collect are the ones you need: memory and disk space are the two most common causes of a sick desk, and neither is collected by default.

ЁЯФз How (in this repo)

monitoring() in ec2/demo.py prints the two checks with their fixes, what is and is not collected for free, and then reuses the busy t3.micro run from lesson 02 (fleet.credits(3)) to find the hour an alarm on CPUCreditBalance < 20 would fire. The snippet below moves the threshold to 50 and starts from a full tank, to show an alarm that gives you warning before the throttle.

ЁЯзк Try it

python3 ec2/demo.py monitoring
python3 - <<'EOF'
import sys; sys.path.insert(0, "ec2"); import fleet
# a t3.micro at 40% on both vCPUs, starting with a full tank (288). When does CPUCreditBalance < 50 fire?
for h, bal, got in fleet.credits(9, busy_pct=40, start_balance=288):
    print(f"hour {h}: {bal:>5} credits ┬╖ CPU {got}%{'  ЁЯФФ ALARM: CPUCreditBalance < 50' if bal < 50 else ''}")
EOF

On a real account (needs AWS credentials; alarms cost a little each month тАФ use your own instance ID and SNS topic):

aws ec2 describe-instance-status --instance-ids i-0123456789abcdef0 --include-all-instances
aws ec2 get-console-output --instance-id i-0123456789abcdef0 --latest --output text
aws cloudwatch put-metric-alarm --alarm-name web-a-credits-low \
    --namespace AWS/EC2 --metric-name CPUCreditBalance \
    --dimensions Name=InstanceId,Value=i-0123456789abcdef0 \
    --statistic Average --period 300 --evaluation-periods 2 \
    --threshold 50 --comparison-operator LessThanThreshold \
    --alarm-actions arn:aws:sns:us-east-1:123456789012:school-oncall
aws cloudwatch put-metric-alarm --alarm-name web-a-system-check \
    --namespace AWS/EC2 --metric-name StatusCheckFailed_System \
    --dimensions Name=InstanceId,Value=i-0123456789abcdef0 \
    --statistic Maximum --period 60 --evaluation-periods 2 \
    --threshold 1 --comparison-operator GreaterThanOrEqualToThreshold \
    --alarm-actions arn:aws:automate:us-east-1:ec2:recover

тЬЕ Verify тАФ what you should see

monitoring prints тЭМ system status check тЖТ AWS's hardware or network under your desk тЖТ stop + start (moves to a new host), тЭМ instance status check тЖТ YOUR operating system (full disk, bad config, kernel) тЖТ reboot, read the console log, fix, the note that memory and disk-space-used need the CloudWatch agent, and ЁЯФФ alarm idea: CPUCreditBalance < 20 тЖТ in the busy run above it fires from hour 2. Your snippet drains 36 credits an hour тАФ 252.0, 216.0, тАж тАФ fires ЁЯФФ ALARM at hour 7 (36.0 credits ┬╖ CPU 40%), and only at hour 8 does the CPU fall to 10% (0.6 credits). The alarm bought you an hour.

ЁЯПБ What you just proved

You can tell AWS's problem from yours in one glance, you know which numbers you get for free and which need an agent, and you set an alarm that fires while there is still time to act.

тЪая╕П Common mistakes

ЁЯПн Why this matters in production: a standard EC2 dashboard has status checks, CPU, credit balance for burstable types, memory and disk from the agent, and the load balancer's healthy-host count and 5xx rate тАФ with alarms that page on what needs a human and auto-recover what does not.

тПня╕П Next

Last lesson: the nightly photocopy, the golden template, and what stop, hibernate and terminate really keep. Backup, patching and lifecycle.

git checkout lesson-12-lifecycle
тЖР Previousload balancingNext тЖТlifecycle

This page is the lesson's README from the lesson-11-monitoring branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.