← कोर्सच्या मुख्य पानाकडे परत

⏮️ आधी काय होते & फायदे-तोटे

प्रत्येक साधनाने काहीतरी वाईट गोष्ट बदलली — आणि तेच साधन कुठेतरी चुकीचे ठरते. प्रत्येक मोठ्या कल्पनेसाठी: आधी जीवन कसे होते, प्रामाणिक फायदे ✅ / तोटे ❌, आणि कुठे वापरावी 👍 विरुद्ध कुठे नाही 👎.

🩺 Monitoring vs observability — lesson 01

⏮️ Before observability

Dashboards for the failures someone predicted; a new kind of failure meant adding code and waiting for it to happen again.

✅ फायदे

  • monitoring: cheap, clear, great for known failure modes
  • observability: rich, correlated signals let you ask questions you did not plan

❌ तोटे

  • monitoring: blind to the unknown
  • observability: more data to collect, store and pay for

👍 वापरा जेव्हा

  • monitoring for the known (disk full, 5xx rate)
  • observability for distributed systems where failures are new each time

👎 दोनदा विचार करा जेव्हा

  • dashboards nobody looks at
  • collecting everything without a question in mind

📓 Logs vs metrics vs traces — lessons 02–05

⏮️ Before the three pillars

One huge text log per server, read with grep during an outage.

✅ फायदे

  • metrics: cheap, fast, perfect for alerts and trends
  • traces: where time goes across services
  • logs: the full story of one event

❌ तोटे

  • metrics: no detail about one request
  • traces: sampling and instrumentation work
  • logs: expensive at volume, hard to aggregate

👍 वापरा जेव्हा

  • alert on metrics · find the slow hop with traces · read the why in logs
  • join them with a trace id

👎 दोनदा विचार करा जेव्हा

  • alerting on log text
  • metrics with user ids as labels

🚨 Cause-based vs symptom-based alerts — lesson 08

⏮️ Before SLOs

Pages for CPU > 80%, disk > 70%, one pod restarting — at 3 a.m., with users perfectly happy.

✅ फायदे

  • symptom (burn-rate) alerts fire when users are actually hurt
  • two windows: fast to fire, fast to clear

❌ तोटे

  • needs an SLI you trust
  • slow leaks need a separate ticket-level alert

👍 वापरा जेव्हा

  • page on fast burn · ticket on slow burn · dashboards for causes
  • keep cause alerts for things that WILL hurt soon (disk nearly full)

👎 दोनदा विचार करा जेव्हा

  • paging on CPU
  • one alert per pod

🔭 Head sampling vs tail sampling — lesson 06

⏮️ Before sampling

Every trace kept — the bill grew faster than the traffic.

✅ फायदे

  • head: decide at the start, cheap and simple
  • tail: decide after the trace ends — keep every error and every slow one

❌ तोटे

  • head: may drop the one trace you needed
  • tail: the Collector must hold traces in memory until they finish

👍 वापरा जेव्हा

  • tail sampling in the Collector for production
  • head sampling for very high volume with low value

👎 दोनदा विचार करा जेव्हा

  • sampling errors away
  • 100% tracing on a high-traffic path without a budget

📈 Pull (Prometheus scrape) vs push (agents, OTLP) — lessons 06, 07

⏮️ दोन्हीपैकी काहीही येण्याआधी

Each server wrote its own graphs; nobody could compare them.

✅ फायदे

  • pull: the server knows who is up; a missing target is itself a signal
  • push: works through firewalls and for short-lived jobs

❌ तोटे

  • pull: needs service discovery and reachable targets
  • push: a silent client looks the same as a dead one

👍 वापरा जेव्हा

  • pull for long-running services in Kubernetes
  • push (Pushgateway / OTLP) for batch jobs and serverless

👎 दोनदा विचार करा जेव्हा

  • pushing from every pod without back-pressure
  • scraping short jobs that finish before the scrape

📝 Blame vs blameless postmortems — lesson 12

⏮️ Before blameless culture

Find the person who pressed the button; people hide mistakes; the same outage returns.

✅ फायदे

  • blameless: people share what really happened
  • fixes go to the system: tests, guards, automation

❌ तोटे

  • takes discipline and follow-up
  • action items can pile up without owners and dates

👍 वापरा जेव्हा

  • every SEV1 and SEV2, and interesting near-misses
  • track action items like product work

👎 दोनदा विचार करा जेव्हा

  • 'human error' as the root cause
  • postmortems with no owners or dates