← कोर्सच्या मुख्य पानाकडे परत

⏮️ आधी काय होते & फायदे-तोटे

प्रत्येक साधनाने काहीतरी वाईट गोष्ट बदलली — आणि तेच साधन कुठेतरी चुकीचे ठरते. प्रत्येक मोठ्या कल्पनेसाठी: आधी जीवन कसे होते, प्रामाणिक फायदे ✅ / तोटे ❌, आणि कुठे वापरावी 👍 विरुद्ध कुठे नाही 👎.

📜 An SLO alone vs an SLO with a signed error budget policy — lesson 02

⏮️ Before budgets

Reliability vs features was argued in every meeting, and the loudest person won.

✅ फायदे

  • policy: a spent budget triggers agreed actions — no argument in the middle of an outage
  • policy: while budget is left, nobody blocks a release 'to be safe'

❌ तोटे

  • policy: needs product, dev and SRE to sign and keep to it
  • a freeze slows features, sometimes for weeks

👍 वापरा जेव्हा

  • any user-facing service with an SLO
  • levels and actions written in advance, with a named escalation

👎 दोनदा विचार करा जेव्हा

  • an SLO with no actions attached (a decoration)
  • one policy copied to every service whatever its users need

🪣 Carrying buckets vs automating the toil — lessons 01, 03

⏮️ Before SRE

An ops team that grew with every new server and never had time to fix the cause.

✅ फायदे

  • automation: toil falls (24.2 → 6.2 h/week in the lab) and stays down
  • self-service removes the ticket altogether

❌ तोटे

  • automation costs build time and upkeep
  • a fast script with no guard rails does damage fast

👍 वापरा जेव्हा

  • frequent, repetitive, well-understood tasks with a short payback
  • remove the cause first (the stuck print queue)

👎 दोनदा विचार करा जेव्हा

  • rare tasks where the build never pays back — unless the risk is high
  • automating a broken process instead of fixing it

📟 One site 24/7 vs follow-the-sun — lesson 04

⏮️ Before rotations

One hero held the pager all the time, until they left.

✅ फायदे

  • follow-the-sun: nobody is paged at night (17 → 0 in the lab)
  • one site: one team, one language, simple handoffs

❌ तोटे

  • follow-the-sun: needs a second site and clean handoffs twice a day
  • one site: night pages and a bigger rotation (8 people)

👍 वापरा जेव्हा

  • follow-the-sun for large, busy services with two strong sites
  • one site with 8+ people for smaller services

👎 दोनदा विचार करा जेव्हा

  • a rotation of 2–3 people 'for now'
  • follow-the-sun with no shared runbooks

🎚️ Big-bang deploys vs progressive delivery with flags — lessons 05, 06

⏮️ Before canaries

Everyone got the new version at once, and rollback meant a new build.

✅ फायदे

  • progressive: a bad change hurts far fewer users (6,000 → 33 failed requests in the lab)
  • flags: turn a feature off in seconds without a deploy

❌ तोटे

  • progressive: a full rollout takes hours, not minutes
  • flags: every flag is an extra code path to test and later remove

👍 वापरा जेव्हा

  • every change to a user-facing service, with an automatic judge
  • risky features behind a kill switch

👎 दोनदा विचार करा जेव्हा

  • a canary too small or short to see the errors you care about
  • flags that live forever

🚦 First come, first served vs shedding by priority — lesson 10

⏮️ Before criticality

Every request waited in the same line; under overload, logins failed as often as prefetches.

✅ फायदे

  • priority: critical work stays whole while sheddable work is refused
  • degradation serves more users with less

❌ तोटे

  • priority labels must be agreed and carried on every request
  • degraded paths must be tested or they fail when needed

👍 वापरा जेव्हा

  • any service that can be overloaded (all of them)
  • retry budgets in every client

👎 दोनदा विचार करा जेव्हा

  • retries without a budget at every layer
  • returning 500 for overload

🧯 Waiting for outages vs chaos experiments — lesson 11

⏮️ Before chaos engineering

Designs were trusted until a real outage tested them — at 3 a.m., on everyone.

✅ फायदे

  • experiments find weaknesses at a chosen time, with a small blast radius
  • game days test people and runbooks too

❌ तोटे

  • experiments cause some real failures (18,000 requests in one cell in the lab)
  • need a steady-state metric, abort conditions and agreement

👍 वापरा जेव्हा

  • checking design claims: zone loss, failover, load shedding
  • start in staging, then small cells in production, in working hours

👎 दोनदा विचार करा जेव्हा

  • no hypothesis, no abort condition
  • the whole of production as the first blast radius