🏫 The School›📈 Scaling›☸️ धडा 07 — Kubernetes वर scaling: pods आणि ते ज्या जमिनीवर उभे असतात ती
🖼️ See the drawing + lab 🏠 Course home 🌿 Branch on GitHub ✏️ View source
🖼️ आकृती आणि labThe drawing + lab पूर्ण पानावर उघडा ↗Open full page ↗

☸️ धडा 07 — Kubernetes वर scaling: pods आणि ते ज्या जमिनीवर उभे असतात ती

📍 तुम्ही इथे आहात: 13 पैकी धडा 07 · मागे: lesson-06-auto-scaling · पुढे: lesson-08-serverless-scaling


📦 या ब्रँचमध्ये काय आहे

धडे 01–06, आणि त्याशिवाय Kubernetes वरचे scaling चे दोन थर: HorizontalPodAutoscaler (HPA) pods वाढवतो, आणि Cluster Autoscaler किंवा Karpenter ते चालवण्यासाठी nodes वाढवतो. दोन्ही प्रामाणिक resource requests वर अवलंबून असतात. scale/demo.py मधले kubernetes() CPU target साठी HPA सूत्र चालवते आणि pods nodes वर बसवते.

🧒 5 वर्षांच्या मुलाला समजावल्यासारखे

जत्रेत प्रत्येक कारकून म्हणजे एक pod, आणि प्रत्येक टेबल म्हणजे एक node. एका टेबलावर ठरावीक कामापुरतीच जागा असते.

ऐश्वर्या कारकुनांवर लक्ष ठेवते. "चार कारकून, प्रत्येकी 90% व्यस्त, आणि मला त्या 60% व्यस्त हव्यात. म्हणून मला 4 × 90 ÷ 60 = 6 कारकून लागतील." ती आणखी दोघींना बोलावते. हा HPA.

पण नव्या कारकुनांना टेबलावर जागा हवी. प्रत्येक कारकून तिला किती जागा हवी ते सांगते — तिची request. एका टेबलाला 4 खुर्च्या आहेत. एक खुर्ची मागणारी कारकून इतर तिघींसोबत बसते. दोन खुर्च्यांपेक्षा थोडी जास्त मागणाऱ्या कारकुनाला टेबल एकटीलाच मिळते — उरलेल्या जागेत दुसरे कुणीच बसत नाही, आणि त्या खुर्च्या रिकाम्याच राहतात.

एखाद्या कारकुनाला बसायला टेबलच नसेल, तर ती बाजूला थांबते (Pending). मग दुसरा मदतनीस — Karpenter किंवा Cluster Autoscaler — आणखी एक टेबल आणतो.

🗺️ आकृती

flowchart LR
    m["📊 metrics-server<br/>CPU vs request"] --> hpa["☸️ HPA<br/>ceil(4 × 90 / 60) = 6"]
    hpa -->|"replicas 4 → 6"| dep["📦 Deployment"]
    dep --> p["🟡 new pods Pending<br/>no node has room"]
    p --> k["🚜 Karpenter / Cluster Autoscaler<br/>adds a node"]
    k --> n["🖥️ node 2,000m<br/>4 pods × 500m"]

🗺️ काढलेली आकृती + एक lab: https://school-edh.pages.dev/scaling/lesson-diagrams.html#l07

❓ काय

🤔 का

कारण Kubernetes वर गर्दीसाठी योग्य क्रमाने दोन गोष्टी लागतात: जास्त pods, आणि त्यांच्यासाठी जागा. HPA सेकंदांत प्रतिक्रिया देतो; नव्या node ला सुमारे एक मिनिट किंवा जास्त लागतो. आणि दोन्ही autoscalers requests वर विश्वास ठेवतात: खूप जास्त requests nodes वाया घालवतात; खूप कमी requests एका node वर खूप pods टाकतात, आणि गर्दी आली की ते CPU साठी भांडतात.

🔧 कसे (या repo मध्ये)

scale/sim.py मधले hpa_desired(replicas, current, target, tolerance=0.1, lo=1, hi=100) हे CPU-utilization target चे सूत्र आहे, त्याच्या tolerance आणि min / max सह (current आणि target हे request चे percentages आहेत). त्यात stabilization window नाही — खरा HPA 6 pods वरून 3 वर जाण्याआधी 5 मिनिटांपर्यंत थांबेल. pack(pod_requests_m, node_cpu_m) pods first-fit decreasing ने ठेवते (आधी सर्वात मोठा, जागा असलेल्या पहिल्या node वर) आणि nodes ची संख्या परत करते — scheduler आणि node autoscaler चे एक साधे model.

🧪 करून पाहा

python3 scale/demo.py kubernetes
python3 - <<'EOF'
import sys; sys.path.insert(0, "scale"); from sim import hpa_desired, pack
for reps, cur in ((4, 65), (4, 67), (4, 54), (10, 61), (10, 120), (2, 600)):
    print(f"{reps:>2} pods at {cur:>3}% (target 60%) → {hpa_desired(reps, cur, 60)} pods")
print("50 pods at 300% (max 100) →", hpa_desired(50, 300, 60))
for req in (250, 500, 700, 1000, 1100):
    print(f"16 pods × {req:>4}m on 2,000m nodes → {pack([req] * 16, 2000):>2} nodes · on 4,000m nodes → {pack([req] * 16, 4000):>2}")
EOF

✅ तपासा — तुम्हाला काय दिसायला हवे

kubernetes हे print करते:

── HorizontalPodAutoscaler on CPU utilisation: desired = ceil(replicas × current % / target %), 10% tolerance
   4 pods at 90% CPU, target 60% → 6 pods
   6 pods at 63% CPU, target 60% → 6 pods
   6 pods at 30% CPU, target 60% → 3 pods
   3 pods at 240% CPU, target 60% → 12 pods
── 16 pods requesting 500m CPU on 2,000m nodes → 4 nodes (4 pods fit per node)
   the same pods requesting 1,100m → 16 nodes (1 fits, 900m wasted per node)

तुमचा snippet हे print करतो:

 4 pods at  65% (target 60%) → 4 pods
 4 pods at  67% (target 60%) → 5 pods
 4 pods at  54% (target 60%) → 4 pods
10 pods at  61% (target 60%) → 10 pods
10 pods at 120% (target 60%) → 20 pods
 2 pods at 600% (target 60%) → 20 pods
50 pods at 300% (max 100) → 100
16 pods ×  250m on 2,000m nodes →  2 nodes · on 4,000m nodes →  1
16 pods ×  500m on 2,000m nodes →  4 nodes · on 4,000m nodes →  2
16 pods ×  700m on 2,000m nodes →  8 nodes · on 4,000m nodes →  4
16 pods × 1000m on 2,000m nodes →  8 nodes · on 4,000m nodes →  4
16 pods × 1100m on 2,000m nodes → 16 nodes · on 4,000m nodes →  6

🏁 तुम्ही आत्ताच काय सिद्ध केले

65% आणि 54% हे 10% tolerance च्या आत आहेत — बदल नाही; 67% बाहेर आहे, म्हणून 4 pods चे 5 होतात. सूत्र थेट उत्तरावर उडी मारते (600% वरचे 2 pods → 20) आणि maximum वर थांबते (100). आणि request bill ठरवते: 1,000m आणि 700m ला तेच 8 nodes लागतात, पण 1,100m ला 16 — node च्या अर्ध्यापलीकडे एक millicore जास्त, आणि nodes दुप्पट. मोठा node (4,000m) कमी वाया घालवतो.

⚠️ नेहमीच्या चुका

🏭 प्रत्यक्ष वापरात

On a real account (EKS) — प्रामाणिक requests, default tolerance आणि लिहून ठेवलेल्या 5 मिनिटांच्या scale-down window सह HPA, आणि एक PDB:

apiVersion: apps/v1
kind: Deployment
metadata: { name: results-api }
spec:
  replicas: 4
  selector: { matchLabels: { app: results-api } }
  template:
    metadata: { labels: { app: results-api } }
    spec:
      containers:
        - name: api
          image: 111122223333.dkr.ecr.ap-south-1.amazonaws.com/results-api:1.4.2
          resources:
            requests: { cpu: 500m, memory: 512Mi }
            limits:   { memory: 512Mi }
          readinessProbe: { httpGet: { path: /health, port: 8080 } }
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: { name: results-api }
spec:
  scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: results-api }
  minReplicas: 4
  maxReplicas: 100
  metrics:
    - type: Resource
      resource: { name: cpu, target: { type: Utilization, averageUtilization: 60 } }
  behavior:
    scaleDown: { stabilizationWindowSeconds: 300 }
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: results-api }
spec:
  minAvailable: 3
  selector: { matchLabels: { app: results-api } }

अनेक instance types मधून, On-Demand किंवा Spot, निवडू शकणारा, CPU मर्यादा आणि consolidation असलेला एक Karpenter NodePool:

apiVersion: karpenter.sh/v1
kind: NodePool
metadata: { name: general }
spec:
  template:
    spec:
      nodeClassRef: { group: karpenter.k8s.aws, kind: EC2NodeClass, name: default }
      requirements:
        - { key: karpenter.sh/capacity-type, operator: In, values: ["on-demand", "spot"] }
        - { key: karpenter.k8s.aws/instance-category, operator: In, values: ["c", "m"] }
  limits: { cpu: "400" }
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized
    consolidateAfter: 1m
kubectl get hpa results-api --watch      # TARGETS shows current/target, e.g. 90%/60%
kubectl get pods --field-selector=status.phase=Pending

🏭 प्रत्यक्ष वापरात हे का महत्त्वाचे: load test खाली खरा CPU आणि memory मोजा (धडा 02), त्यावरून requests ठरवा, आणि दर तिमाहीला त्यांचा आढावा घ्या. दोन्ही autoscalers त्या दोन आकड्यांइतकेच चांगले असतात.

⏭️ पुढे

Pods आणि nodes ना अजूनही सेकंदांपासून मिनिटांपर्यंत वेळ लागतो. पालक येताच, प्रत्येक पालकासाठी एक खिडकी उघडली तर? Serverless.

git checkout lesson-08-serverless-scaling

☸️ Lesson 07 — Scaling on Kubernetes: pods and the ground they stand on

📍 You are here: Lesson 07 of 13 · Previous: lesson-06-auto-scaling · Next: lesson-08-serverless-scaling


📦 What's in this branch

Lessons 01–06, plus the two layers of scaling on Kubernetes: the HorizontalPodAutoscaler (HPA) adds pods, and Cluster Autoscaler or Karpenter adds nodes for them to run on. Both depend on honest resource requests. kubernetes() in scale/demo.py runs the HPA formula for a CPU target and packs pods onto nodes.

🧒 Explain like I'm 5

At the fair, each clerk is a pod, and each table is a node. A table has room for a fixed amount of work.

Aishwarya watches the clerks. "Four clerks, each 90% busy, and I want them 60% busy. So I need 4 × 90 ÷ 60 = 6 clerks." She calls two more. That is the HPA.

But the new clerks need space at a table. Each clerk says how much space she needs — her request. A table has 4 chairs. A clerk who asks for one chair sits with three others. A clerk who asks for a little more than two chairs gets a table alone — nobody else fits in what is left, and those chairs stay empty.

When a clerk has no table to sit at, she waits at the side (Pending). Then a second helper — Karpenter or Cluster Autoscaler — brings another table.

🗺️ Diagram

flowchart LR
    m["📊 metrics-server<br/>CPU vs request"] --> hpa["☸️ HPA<br/>ceil(4 × 90 / 60) = 6"]
    hpa -->|"replicas 4 → 6"| dep["📦 Deployment"]
    dep --> p["🟡 new pods Pending<br/>no node has room"]
    p --> k["🚜 Karpenter / Cluster Autoscaler<br/>adds a node"]
    k --> n["🖥️ node 2,000m<br/>4 pods × 500m"]

🗺️ Drawn version + a lab: https://school-edh.pages.dev/scaling/lesson-diagrams.html#l07

❓ What

🤔 Why

Because on Kubernetes a spike needs two things in the right order: more pods, and room for them. The HPA reacts in seconds; a new node takes about a minute or more. And both autoscalers trust the requests: requests too high waste nodes; requests too low put too many pods on one node, and they fight for CPU when the crowd comes.

🔧 How (in this repo)

hpa_desired(replicas, current, target, tolerance=0.1, lo=1, hi=100) in scale/sim.py is the formula for a CPU-utilization target, with its tolerance and a min / max (current and target are percentages of the request). It has no stabilization window — the real HPA would wait up to 5 minutes before going from 6 pods to 3. pack(pod_requests_m, node_cpu_m) places pods with first-fit decreasing (biggest first, on the first node with room) and returns the node count — a simple model of the scheduler plus a node autoscaler.

🧪 Try it

python3 scale/demo.py kubernetes
python3 - <<'EOF'
import sys; sys.path.insert(0, "scale"); from sim import hpa_desired, pack
for reps, cur in ((4, 65), (4, 67), (4, 54), (10, 61), (10, 120), (2, 600)):
    print(f"{reps:>2} pods at {cur:>3}% (target 60%) → {hpa_desired(reps, cur, 60)} pods")
print("50 pods at 300% (max 100) →", hpa_desired(50, 300, 60))
for req in (250, 500, 700, 1000, 1100):
    print(f"16 pods × {req:>4}m on 2,000m nodes → {pack([req] * 16, 2000):>2} nodes · on 4,000m nodes → {pack([req] * 16, 4000):>2}")
EOF

✅ Verify — what you should see

kubernetes prints:

── HorizontalPodAutoscaler on CPU utilisation: desired = ceil(replicas × current % / target %), 10% tolerance
   4 pods at 90% CPU, target 60% → 6 pods
   6 pods at 63% CPU, target 60% → 6 pods
   6 pods at 30% CPU, target 60% → 3 pods
   3 pods at 240% CPU, target 60% → 12 pods
── 16 pods requesting 500m CPU on 2,000m nodes → 4 nodes (4 pods fit per node)
   the same pods requesting 1,100m → 16 nodes (1 fits, 900m wasted per node)

Your snippet prints:

 4 pods at  65% (target 60%) → 4 pods
 4 pods at  67% (target 60%) → 5 pods
 4 pods at  54% (target 60%) → 4 pods
10 pods at  61% (target 60%) → 10 pods
10 pods at 120% (target 60%) → 20 pods
 2 pods at 600% (target 60%) → 20 pods
50 pods at 300% (max 100) → 100
16 pods ×  250m on 2,000m nodes →  2 nodes · on 4,000m nodes →  1
16 pods ×  500m on 2,000m nodes →  4 nodes · on 4,000m nodes →  2
16 pods ×  700m on 2,000m nodes →  8 nodes · on 4,000m nodes →  4
16 pods × 1000m on 2,000m nodes →  8 nodes · on 4,000m nodes →  4
16 pods × 1100m on 2,000m nodes → 16 nodes · on 4,000m nodes →  6

🏁 What you just proved

65% and 54% are inside the 10% tolerance — no change; 67% is outside, so 4 pods become 5. The formula jumps straight to the answer (2 pods at 600% → 20) and stops at the maximum (100). And the request decides the bill: 1,000m and 700m need the same 8 nodes, but 1,100m needs 16 — one more millicore past half a node doubles the nodes. A bigger node (4,000m) wastes less.

⚠️ Common mistakes

🏭 In production

On a real account (EKS) — honest requests, an HPA with the default tolerance and a 5-minute scale-down window written out, and a PDB:

apiVersion: apps/v1
kind: Deployment
metadata: { name: results-api }
spec:
  replicas: 4
  selector: { matchLabels: { app: results-api } }
  template:
    metadata: { labels: { app: results-api } }
    spec:
      containers:
        - name: api
          image: 111122223333.dkr.ecr.ap-south-1.amazonaws.com/results-api:1.4.2
          resources:
            requests: { cpu: 500m, memory: 512Mi }
            limits:   { memory: 512Mi }
          readinessProbe: { httpGet: { path: /health, port: 8080 } }
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: { name: results-api }
spec:
  scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: results-api }
  minReplicas: 4
  maxReplicas: 100
  metrics:
    - type: Resource
      resource: { name: cpu, target: { type: Utilization, averageUtilization: 60 } }
  behavior:
    scaleDown: { stabilizationWindowSeconds: 300 }
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: results-api }
spec:
  minAvailable: 3
  selector: { matchLabels: { app: results-api } }

A Karpenter NodePool that may choose from several instance types, On-Demand or Spot, with a CPU ceiling and consolidation:

apiVersion: karpenter.sh/v1
kind: NodePool
metadata: { name: general }
spec:
  template:
    spec:
      nodeClassRef: { group: karpenter.k8s.aws, kind: EC2NodeClass, name: default }
      requirements:
        - { key: karpenter.sh/capacity-type, operator: In, values: ["on-demand", "spot"] }
        - { key: karpenter.k8s.aws/instance-category, operator: In, values: ["c", "m"] }
  limits: { cpu: "400" }
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized
    consolidateAfter: 1m
kubectl get hpa results-api --watch      # TARGETS shows current/target, e.g. 90%/60%
kubectl get pods --field-selector=status.phase=Pending

🏭 Why this matters in production: measure real CPU and memory under a load test (lesson 02), set requests from that, and review them each quarter. Both autoscalers are only as good as those two numbers.

⏭️ Next

Pods and nodes still take seconds to minutes. What if a counter appeared for every parent, the moment she arrives? Serverless.

git checkout lesson-08-serverless-scaling
← Previousauto scalingNext →serverless scaling

This page is the lesson's README from the lesson-07-kubernetes-scaling branch, shown here so the whole School stays on one site. Code files open on GitHub at the same branch.