ЁЯПл The SchoolтА║ЁЯЪк NATтА║ЁЯФО рдзрдбрд╛ 08 тАФ рдлрд╛рдЯрдХрд╛рдЪреЗ troubleshooting
ЁЯЦ╝я╕П See the drawing + lab ЁЯПа Course home ЁЯМ┐ Branch on GitHub тЬПя╕П View source
ЁЯЦ╝я╕П рдЖрдХреГрддреА рдЖрдгрд┐ labThe drawing + lab рдкреВрд░реНрдг рдкрд╛рдирд╛рд╡рд░ рдЙрдШрдбрд╛ тЖЧOpen full page тЖЧ

ЁЯФО рдзрдбрд╛ 08 тАФ рдлрд╛рдЯрдХрд╛рдЪреЗ troubleshooting

ЁЯУН рддреБрдореНрд╣реА рдЗрдереЗ рдЖрд╣рд╛рдд: 8 рдкреИрдХреА рдзрдбрд╛ 08 ┬╖ рдорд╛рдЧреЗ: lesson-07-alternatives


ЁЯУж рдпрд╛ рдмреНрд░рдБрдЪрдордзреНрдпреЗ рдХрд╛рдп рдЖрд╣реЗ

рдзрдбреЗ 01тАУ07, рдЖрдгрд┐ рддреБрдореНрд╣рд╛рд▓рд╛ рдкреНрд░рддреНрдпрдХреНрд╖рд╛рдд рджрд┐рд╕рдгрд╛рд░реЗ рджреЛрди рдмрд┐рдШрд╛рдб: port exhaustion (рдПрдХрд╛ destination рд╕рд╛рдареА рдиреЛрдВрджрд╡рд╣реА рднрд░рд▓реА тЖТ ErrorPortAllocation) рдЖрдгрд┐ 350-second idle timeout (рд╢рд╛рдВрдд рдУрд│ рд╡рд┐рд╕рд░рд▓реА рдЬрд╛рддреЗ тЖТ connection reset рд╣реЛрддреЗ). рдЖрдгрд┐ рджреЛрдиреНрд╣реА рджрд╛рдЦрд╡рдгрд╛рд░реЗ CloudWatch metrics. nat/demo.py рдордзреАрд▓ troubleshoot() рдореБрджреНрджрд╛рдо рдлрд╛рдЯрдХ рдмрд┐рдШрдбрд╡рддреЗ.

ЁЯзТ 5 рд╡рд░реНрд╖рд╛рдВрдЪреНрдпрд╛ рдореБрд▓рд╛рд▓рд╛ рд╕рдордЬрд╛рд╡рд▓реНрдпрд╛рд╕рд╛рд░рдЦреЗ

рдлрд╛рдЯрдХрд╛рд╡рд░ рджреЛрди рдЧреЛрд╖реНрдЯреА рдЪреБрдХреВ рд╢рдХрддрд╛рдд:

ЁЯЧ║я╕П рдЖрдХреГрддреА

flowchart TB
    subgraph full["ЁЯУТ port exhaustion"]
      c["connections 1тАУ5 тЖТ pypi"] --> ok["тЬЕ ports 1024тАУ1028"]
      c6["connection 6 тЖТ pypi"] --> err["тЭМ ErrorPortAllocation"]
      gh["connection тЖТ github"] --> ok2["тЬЕ different destination"]
    end
    subgraph idle["ЁЯХ░я╕П idle timeout"]
      t0["t=0 row written"] --> t400["t=400 s quiet > 350 s тЖТ row wiped"]
      t400 --> late["late answer тЖТ dropped, app sees a reset"]
    end

ЁЯЧ║я╕П рдХрд╛рдврд▓реЗрд▓реА рдЖрдХреГрддреА + рдПрдХ lab: https://school-edh.pages.dev/nat/lesson-diagrams.html#l08

тЭУ рдХрд╛рдп

ЁЯдФ рдХрд╛

рдХрд╛рд░рдг рджреЛрдиреНрд╣реА рдмрд┐рдШрд╛рдб рдЕрдзреВрдирдордзреВрди рдпреЗрдгрд╛рд▒реНрдпрд╛ application errors рд╕рд╛рд░рдЦреЗ рджрд┐рд╕рддрд╛рдд: "load рдЦрд╛рд▓реА timeouts" рдЖрдгрд┐ "рд╢рд╛рдВрдд рдХрд╛рд▓рд╛рд╡рдзреАрдирдВрддрд░ connection reset". рдиреЛрдВрджрд╡рд╣реА рдорд╛рд╣реАрдд рдирд╕реЗрд▓ рддрд░ teams servers restart рдХрд░рддрд╛рдд рдХрд┐рдВрд╡рд╛ рдореЛрдареА instances рдЬреЛрдбрддрд╛рдд, рдЖрдгрд┐ рдХрд╛рд╣реАрдЪ рдмрджрд▓рдд рдирд╛рд╣реА. Metrics рдЕрд╕рддреАрд▓ рддрд░ рдХрд╛рд░рдг рдХрд╛рд╣реА рдорд┐рдирд┐рдЯрд╛рдВрдд рджрд┐рд╕рддреЗ.

ЁЯФз рдХрд╕реЗ (рдпрд╛ repo рдордзреНрдпреЗ)

nat/gateway.py рдордзреНрдпреЗ, рдкреНрд░рддреНрдпреЗрдХ IP рдиреЗ рддреНрдпрд╛ destination рдХрдбреЗ ports_per_dest ports рд╡рд╛рдкрд░рд▓реЗ рдХреА send() рд╣реЗ self.errors += 1 рдореЛрдЬрддреЗ (model рдЪреЗ ErrorPortAllocation). tick(now) рдЬреНрдпрд╛ рдУрд│реАрдВрдЪреА last рд╡реЗрд│ idle_timeout seconds рдкреЗрдХреНрд╖рд╛ рдЬрд╛рд╕реНрдд рдЬреБрдиреА рдЖрд╣реЗ рддреНрдпрд╛ рдкреБрд╕рддреЗ, рдЖрдгрд┐ рдУрд│реАрд╡рд░рдЪреЗ рдкреНрд░рддреНрдпреЗрдХ send() рдХрд┐рдВрд╡рд╛ reply() рд╣реЗ last рддрд╛рдЬреЗ рдХрд░рддреЗ тАФ keepalive рдиреЗрдордХреЗ рд╣реЗрдЪ рдХрд░рддреЛ. nat/demo.py рдордзреАрд▓ troubleshoot() рдиреЛрдВрджрд╡рд╣реА 5 ports рдкрд░реНрдпрдВрдд рд▓рд╣рд╛рди рдХрд░рддреЗ, рдореНрд╣рдгрдЬреЗ рддреА screen рд╡рд░рдЪ рднрд░рддреЗ. (Model рдордзреНрдпреЗ рдЦрд░реЗ 55,000 рднрд░рдгреНрдпрд╛рдЪрд╛ рдкреНрд░рдпрддреНрди рдХрд░реВ рдирдХрд╛ тАФ рддреЗ рдкреНрд░рддреНрдпреЗрдХ send рд▓рд╛ рд╕рдВрдкреВрд░реНрдг table scan рдХрд░рддреЗ, рддреНрдпрд╛рдореБрд│реЗ рддреНрдпрд╛рд▓рд╛ рдЦреВрдк рд╡реЗрд│ рд▓рд╛рдЧрддреЛ. рддреНрдпрд╛рдРрд╡рдЬреА ports_per_dest рд▓рд╣рд╛рди рдХрд░рд╛.)

ЁЯзк рдХрд░реВрди рдкрд╛рд╣рд╛

python3 nat/demo.py troubleshoot
python3 - <<'EOF'
import sys; sys.path.insert(0, "nat"); from gateway import NatGateway
PYPI = ("151.101.0.223", 443, "tcp")
for now in (350, 351):
    g = NatGateway("nat-a", ["13.232.10.20"]); ip, port = g.send(("10.20.48.25", 50001), PYPI, t=0)
    print(f"quiet until t={now}: wiped {g.tick(now)} ┬╖ reply тЖТ {g.reply(ip, port, PYPI, now)}")
g = NatGateway("nat-a", ["13.232.10.20"]); ip, port = g.send(("10.20.48.25", 50001), PYPI, t=0)
for t in (300, 600, 900):                       # a keepalive every 300 s
    g.send(("10.20.48.25", 50001), PYPI, t=t); g.tick(t)
print("with keepalive every 300 s, at t=1000:", g.tick(1000), "wiped ┬╖ reply тЖТ", g.reply(ip, port, PYPI, 1000))
EOF
python3 nat/test_nat.py

тЬЕ рддрдкрд╛рд╕рд╛ тАФ рддреБрдореНрд╣рд╛рд▓рд╛ рдХрд╛рдп рджрд┐рд╕рд╛рдпрд▓рд╛ рд╣рд╡реЗ

troubleshoot рд╣реЗ connection 1 тЖТ 13.232.10.20:1024 рдкрд╛рд╕реВрди connection 5 тЖТ 13.232.10.20:1028 рдкрд░реНрдпрдВрдд рдЫрд╛рдкрддреЗ, рдордЧ connection 6 тЖТ тЭМ ErrorPortAllocation рдЖрдгрд┐ connection 7 тЖТ тЭМ ErrorPortAllocation, рдордЧ a DIFFERENT destination still works тЖТ rows: 6 ┬╖ ErrorPortAllocation = 2, рдЖрдгрд┐ t=400 s the server answers late тЖТ wiped 1 row тЖТ reply dropped (the app sees a reset). рддреБрдордЪрд╛ snippet рд╣реЗ рдЫрд╛рдкрддреЛ:

quiet until t=350: wiped 0 ┬╖ reply тЖТ ('10.20.48.25', 50001)
quiet until t=351: wiped 1 ┬╖ reply тЖТ None
with keepalive every 300 s, at t=1000: 0 wiped ┬╖ reply тЖТ ('10.20.48.25', 50001)

рдмрд░реЛрдмрд░ 350 seconds рд▓рд╛ рдУрд│ рдЯрд┐рдХрддреЗ; рдПрдХ second рдЬрд╛рд╕реНрдд рдЭрд╛рд▓рд╛ рдХреА рддреА рдЬрд╛рддреЗ. рджрд░ 300 seconds рд▓рд╛ keepalive рдЕрд╕реЗрд▓ рддрд░ 1,000 seconds рдирдВрддрд░рд╣реА рдУрд│ рддрд┐рдереЗрдЪ рдЕрд╕рддреЗ. Tests 8/8 passed рдиреЗ рд╕рдВрдкрддрд╛рдд.

ЁЯПБ рддреБрдореНрд╣реА рдЖрддреНрддрд╛рдЪ рдХрд╛рдп рд╕рд┐рджреНрдз рдХреЗрд▓реЗ

"Load рдЦрд╛рд▓реА connect рд╣реЛрдд рдирд╛рд╣реА" рдореНрд╣рдгрдЬреЗ рдПрдХрд╛ destination рд╕рд╛рдареА рднрд░рд▓реЗрд▓реА рдиреЛрдВрджрд╡рд╣реА рдЕрд╕реВ рд╢рдХрддреЗ, рдЖрдгрд┐ "рд╢рд╛рдВрдд рдХрд╛рд▓рд╛рд╡рдзреАрдирдВрддрд░ reset" рдореНрд╣рдгрдЬреЗ рд╡рд┐рд╕рд░рд▓реЗрд▓реА рдУрд│ рдЕрд╕реВ рд╢рдХрддреЗ тАФ рдЖрдгрд┐ рдкреНрд░рддреНрдпреЗрдХрд╛рдЪрд╛ metric рдЖрдгрд┐ рдЙрдкрд╛рдп рддреБрдореНрд╣рд╛рд▓рд╛ рдорд╛рд╣реАрдд рдЖрд╣реЗ.

тЪая╕П рдиреЗрд╣рдореАрдЪреНрдпрд╛ рдЪреБрдХрд╛

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд

рдЦрд▒реНрдпрд╛ account рд╡рд░ тАФ port allocation errors рд╡рд░ alarm рд▓рд╛рд╡рд╛, idle-timeout count рд╡рд╛рдЪрд╛, рдЖрдгрд┐ Linux instance рд╡рд░ 350 seconds рдкреЗрдХреНрд╖рд╛ рдХрдореА keepalive рдареЗрд╡рд╛:

aws cloudwatch put-metric-alarm --alarm-name nat-a-port-allocation --namespace AWS/NATGateway --metric-name ErrorPortAllocation --dimensions Name=NatGatewayId,Value=nat-0abc --statistic Sum --period 300 --evaluation-periods 1 --threshold 0 --comparison-operator GreaterThanThreshold --alarm-actions arn:aws:sns:us-east-1:111122223333:ops
aws cloudwatch get-metric-statistics --namespace AWS/NATGateway --metric-name IdleTimeoutCount --dimensions Name=NatGatewayId,Value=nat-0abc --start-time 2026-09-26T00:00:00Z --end-time 2026-09-27T00:00:00Z --period 3600 --statistics Sum
aws cloudwatch get-metric-statistics --namespace AWS/NATGateway --metric-name PacketsDropCount --dimensions Name=NatGatewayId,Value=nat-0abc --start-time 2026-09-26T00:00:00Z --end-time 2026-09-27T00:00:00Z --period 3600 --statistics Sum
sudo sysctl -w net.ipv4.tcp_keepalive_time=300    # on the instance; the app must enable SO_KEEPALIVE
resource "aws_cloudwatch_metric_alarm" "nat_ports" {
  for_each            = aws_nat_gateway.this
  alarm_name          = "nat-${each.key}-port-allocation"
  namespace           = "AWS/NATGateway"
  metric_name         = "ErrorPortAllocation"
  dimensions          = { NatGatewayId = each.value.id }
  statistic           = "Sum"
  period              = 300
  evaluation_periods  = 1
  threshold           = 0
  comparison_operator = "GreaterThanThreshold"
}

ЁЯПн рдкреНрд░рддреНрдпрдХреНрд╖ рд╡рд╛рдкрд░рд╛рдд рд╣реЗ рдХрд╛ рдорд╣рддреНрддреНрд╡рд╛рдЪреЗ: рдкреНрд░рддреНрдпреЗрдХ NAT gateway рд╡рд░ ErrorPortAllocation рдЪрд╛ alarm рдЕрд╕рд╛рдпрд▓рд╛ рд╣рд╡рд╛ рдЖрдгрд┐ ActiveConnectionCount, IdleTimeoutCount, PacketsDropCount рд╡ BytesOutToDestination рдЕрд╕рд▓реЗрд▓рд╛ dashboard рдЕрд╕рд╛рдпрд▓рд╛ рд╣рд╡рд╛. рд╣реЗ рдЪрд╛рд░ graphs рдХреЛрдгреА ticket рдЙрдШрдбрдгреНрдпрд╛рдЖрдзреАрдЪ рдмрд╣реБрддреЗрдХ egress рдкреНрд░рд╢реНрдирд╛рдВрдЪреА рдЙрддреНрддрд░реЗ рджреЗрддрд╛рдд.

ЁЯОУ рдлрд╛рдЯрдХ рдЖрддрд╛ рддреБрдордЪреЗ рдЖрд╣реЗ

рд╢рд╛рд│реЗрдЪреЗ рдлрд╛рдЯрдХ тЖТ рдиреЛрдВрджрд╡рд╣реА тЖТ managed рдлрд╛рдЯрдХ рдХреА рд╕реНрд╡рддрдГрдЪреЗ тЖТ рдкреНрд░рддреНрдпреЗрдХ рдЗрдорд╛рд░рддреАрд▓рд╛ рдПрдХ рдлрд╛рдЯрдХ тЖТ рд╕реНрдерд┐рд░ рдкрд░рддреАрдЪрд╛ рдкрддреНрддрд╛ тЖТ рджреЛрди рдореАрдЯрд░ тЖТ рдлрд╛рдЯрдХ рдЯрд╛рд│рдгрд╛рд▒реНрдпрд╛ рдлреЗрд▒реНрдпрд╛ тЖТ рднрд░рд▓реЗрд▓реНрдпрд╛ рдиреЛрдВрджрд╡рд╣реНрдпрд╛ рдЖрдгрд┐ рд╡рд┐рд╕рд░рд▓реЗрд▓реНрдпрд╛ рдУрд│реА. рддреБрдореНрд╣реА рдлрдХреНрдд NAT рд╢рд┐рдХрд▓рд╛ рдирд╛рд╣реАрдд тАФ рддреБрдореНрд╣реА рд╢рд╛рд│реЗрдЪреЗ рдлрд╛рдЯрдХ рдЪрд╛рд▓рд╡реВ рд╢рдХрддрд╛. ЁЯЪкЁЯОУ

рд╢рд╛рд│реЗрддрд▓реА рдкреБрдврдЪреА рджрд╛рд░реЗ: VPC рд╢рд╛рд│рд╛ рд╣реЗ рдлрд╛рдЯрдХ рдЬреНрдпрд╛ рд╕рдВрдкреВрд░реНрдг campus рдЪрд╛ рднрд╛рдЧ рдЖрд╣реЗ рддреЛ рдмрд╛рдВрдзрддреЗ; Networking рд╢рд╛рд│рд╛ IP, ports рдЖрдгрд┐ TCP рдордзреНрдпреЗ рдЕрдзрд┐рдХ рдЦреЛрд▓рд╡рд░ рдЬрд╛рддреЗ; School portal рдордзреНрдпреЗ рдмрд╛рдХреА рд╕рдЧрд│реЗ рдЖрд╣реЗ.

git checkout main
python3 nat/demo.py     # one last run, for fun

ЁЯФО Lesson 08 тАФ Troubleshooting the gate

ЁЯУН You are here: Lesson 08 of 8 ┬╖ Previous: lesson-07-alternatives


ЁЯУж What's in this branch

Lessons 01тАУ07, plus the two failures you will actually see: port exhaustion (the register is full for one destination тЖТ ErrorPortAllocation) and the 350-second idle timeout (a quiet row is forgotten тЖТ the connection is reset). And the CloudWatch metrics that show both. troubleshoot() in nat/demo.py breaks the gate on purpose.

ЁЯзТ Explain like I'm 5

Two things can go wrong at the gate:

ЁЯЧ║я╕П Diagram

flowchart TB
    subgraph full["ЁЯУТ port exhaustion"]
      c["connections 1тАУ5 тЖТ pypi"] --> ok["тЬЕ ports 1024тАУ1028"]
      c6["connection 6 тЖТ pypi"] --> err["тЭМ ErrorPortAllocation"]
      gh["connection тЖТ github"] --> ok2["тЬЕ different destination"]
    end
    subgraph idle["ЁЯХ░я╕П idle timeout"]
      t0["t=0 row written"] --> t400["t=400 s quiet > 350 s тЖТ row wiped"]
      t400 --> late["late answer тЖТ dropped, app sees a reset"]
    end

ЁЯЧ║я╕П Drawn version + a lab: https://school-edh.pages.dev/nat/lesson-diagrams.html#l08

тЭУ What

ЁЯдФ Why

Because both failures look like random application errors: "timeouts under load" and "connection reset after a quiet period". Without knowing the register, teams restart servers or add bigger instances, and nothing changes. With the metrics, the cause is visible in minutes.

ЁЯФз How (in this repo)

In nat/gateway.py, send() counts self.errors += 1 (the model's ErrorPortAllocation) when every IP has used ports_per_dest ports to that destination. tick(now) wipes rows whose last time is more than idle_timeout seconds old, and every send() or reply() on a row refreshes last тАФ which is exactly what a keepalive does. troubleshoot() in nat/demo.py shrinks the register to 5 ports so it fills on screen. (Do not try to fill the real 55,000 in the model тАФ it scans the whole table on every send, so that takes a very long time. Shrink ports_per_dest instead.)

ЁЯзк Try it

python3 nat/demo.py troubleshoot
python3 - <<'EOF'
import sys; sys.path.insert(0, "nat"); from gateway import NatGateway
PYPI = ("151.101.0.223", 443, "tcp")
for now in (350, 351):
    g = NatGateway("nat-a", ["13.232.10.20"]); ip, port = g.send(("10.20.48.25", 50001), PYPI, t=0)
    print(f"quiet until t={now}: wiped {g.tick(now)} ┬╖ reply тЖТ {g.reply(ip, port, PYPI, now)}")
g = NatGateway("nat-a", ["13.232.10.20"]); ip, port = g.send(("10.20.48.25", 50001), PYPI, t=0)
for t in (300, 600, 900):                       # a keepalive every 300 s
    g.send(("10.20.48.25", 50001), PYPI, t=t); g.tick(t)
print("with keepalive every 300 s, at t=1000:", g.tick(1000), "wiped ┬╖ reply тЖТ", g.reply(ip, port, PYPI, 1000))
EOF
python3 nat/test_nat.py

тЬЕ Verify тАФ what you should see

troubleshoot prints connection 1 тЖТ 13.232.10.20:1024 up to connection 5 тЖТ 13.232.10.20:1028, then connection 6 тЖТ тЭМ ErrorPortAllocation and connection 7 тЖТ тЭМ ErrorPortAllocation, then a DIFFERENT destination still works тЖТ rows: 6 ┬╖ ErrorPortAllocation = 2, and t=400 s the server answers late тЖТ wiped 1 row тЖТ reply dropped (the app sees a reset). Your snippet prints:

quiet until t=350: wiped 0 ┬╖ reply тЖТ ('10.20.48.25', 50001)
quiet until t=351: wiped 1 ┬╖ reply тЖТ None
with keepalive every 300 s, at t=1000: 0 wiped ┬╖ reply тЖТ ('10.20.48.25', 50001)

At exactly 350 seconds the row survives; one second more and it is gone. With a keepalive every 300 seconds, the row is still there after 1,000 seconds. The tests end with 8/8 passed.

ЁЯПБ What you just proved

"Cannot connect under load" can be a full register for one destination, and "reset after a quiet period" can be a forgotten row тАФ and you know the metric and the fix for each.

тЪая╕П Common mistakes

ЁЯПн In production

On a real account тАФ alarm on port allocation errors, read the idle-timeout count, and set a keepalive under 350 seconds on a Linux instance:

aws cloudwatch put-metric-alarm --alarm-name nat-a-port-allocation --namespace AWS/NATGateway --metric-name ErrorPortAllocation --dimensions Name=NatGatewayId,Value=nat-0abc --statistic Sum --period 300 --evaluation-periods 1 --threshold 0 --comparison-operator GreaterThanThreshold --alarm-actions arn:aws:sns:us-east-1:111122223333:ops
aws cloudwatch get-metric-statistics --namespace AWS/NATGateway --metric-name IdleTimeoutCount --dimensions Name=NatGatewayId,Value=nat-0abc --start-time 2026-09-26T00:00:00Z --end-time 2026-09-27T00:00:00Z --period 3600 --statistics Sum
aws cloudwatch get-metric-statistics --namespace AWS/NATGateway --metric-name PacketsDropCount --dimensions Name=NatGatewayId,Value=nat-0abc --start-time 2026-09-26T00:00:00Z --end-time 2026-09-27T00:00:00Z --period 3600 --statistics Sum
sudo sysctl -w net.ipv4.tcp_keepalive_time=300    # on the instance; the app must enable SO_KEEPALIVE
resource "aws_cloudwatch_metric_alarm" "nat_ports" {
  for_each            = aws_nat_gateway.this
  alarm_name          = "nat-${each.key}-port-allocation"
  namespace           = "AWS/NATGateway"
  metric_name         = "ErrorPortAllocation"
  dimensions          = { NatGatewayId = each.value.id }
  statistic           = "Sum"
  period              = 300
  evaluation_periods  = 1
  threshold           = 0
  comparison_operator = "GreaterThanThreshold"
}

ЁЯПн Why this matters in production: every NAT gateway should have an alarm on ErrorPortAllocation and a dashboard with ActiveConnectionCount, IdleTimeoutCount, PacketsDropCount and BytesOutToDestination. Those four graphs answer most egress questions before anyone opens a ticket.

ЁЯОУ The gate is yours

The school gate тЖТ the register тЖТ a managed gate or your own тЖТ one gate per building тЖТ a fixed return address тЖТ two meters тЖТ trips that skip the gate тЖТ full registers and forgotten rows. You didn't just learn NAT тАФ you can run the school gate. ЁЯЪкЁЯОУ

Next doors in the school: the VPC school builds the whole campus this gate belongs to; the Networking school goes deeper into IP, ports and TCP; the School portal has the rest.

git checkout main
python3 nat/demo.py     # one last run, for fun
тЖР PreviousalternativesFinished! Take the quiz тЖТcheck what stuck

This page is the lesson's README from the lesson-08-troubleshooting branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.