M8 โ Observability & SRE¶
Core question: It is 2 a.m. and something is slow. How do you see inside a running system โ without SSHing into every pod and grepping log files one by one?
โฑ๏ธ Time: ~55 min padho ยท ๐๏ธ Level: Advanced ยท ๐ Pehle chahiye: M4, M7
Is module ke baad tum kar paoge: - Prometheus metrics, Loki logs, aur OpenTelemetry traces ka fark explain karo aur ek production incident mein teeno use karo - SLI/SLO/error budget define karo aur yeh bolo ke error budget exhaust hone pe team kya karti hai - Cardinality explosion diagnose aur fix karo (high-cardinality label โ Prometheus OOM)
โฉ๏ธ Recall gate โ shuru karne se pehle¶
Pichhle modules se 3 sawaal. Pehle memory se jawab do, phir kholo. (Yeh retrieve karna hi lifetime yaad rakhta hai โ dobara padhna nahi.)
- (M4) K8s mein readiness probe aur liveness probe mein kya fark hai? Readiness probe fail hone par Service ka kya hota hai?
- (M7)
selfHeal: trueset hai aur kisi nekubectl edit deploymentse live image tag change kar diya. Argo kya karega โ Cause A ya Cause B, aur kitne time mein?- (M6) CI mein
github.shatag kyun use karte hainlatestki jagah? Ek concrete problem batao jolatestse production mein aata hai.
Jawab
- Readiness fail โ pod Service ke EndpointSlice se hata diya jaata hai (traffic nahi milta, restart nahi hota). Liveness fail โ pod restart hota hai. Dono alag cheezein โ alag problems ke liye. 2. Cause B (cluster drifted from Git). selfHeal ~30sโ3min mein Git wala version wapas apply karta hai. 3.
latestmutable hai โ alag nodes pe alag images pull ho sakti hain, rollback mushkil.github.shaimmutable hai โ exactly wohi build deploy hota hai jo test se guzri.
(Formerly tracked as M10 in the Principal Track; renumbered M8 here as the first Operate module.)
MODULE MAP
00-INDEX ยท 01-M0-foundations ยท 02-M1-terraform ยท 03-M2-ansible ยท 04-M3-docker ยท 05-M4-kubernetes-core ยท 06-M5-sizing-and-cost ยท 07-M6-cicd ยท 08-M7-gitops ยท 09-connected-system ยท 10-M8-observability-sre ยท 11-M9-advanced-k8s-internals ยท 12-capstone-url-shortener ยท 13-capstone-microshop ยท 14-interview-bank ยท 15-roadmap-M11-M18 ยท 16-reference-appendix
The 60-second version¶
Kubernetes tells you a pod is Running. It does not tell you that pod is answering in 3 seconds instead of 80 milliseconds, or silently returning errors on 2% of requests, or about to run out of memory in 20 minutes. That gap โ between "process is alive" and "system is healthy" โ is what observability fills.
Observability has three pillars, each answering a different question:
- Metrics (Prometheus): numbers over time โ "how much, how many, how fast?"
- Logs (Loki / ELK): discrete events โ "what exactly happened on this request?"
- Traces (OpenTelemetry + Jaeger): a request's journey across services โ "where did the time go?"
The business layer on top is SRE โ Site Reliability Engineering. It turns raw telemetry into a governance contract: an error budget that answers the question "are we reliable enough to ship this risky change right now?"
You alert on metrics. You debug with logs and traces. You make deploy/freeze decisions with error budgets. Each tool has one job.
Why this exists / what it replaced¶
The before state: flying blind¶
Before structured observability, the typical production investigation looked like this:
- Customer complains โ or worse, nobody notices until a spike in support tickets.
- Engineer SSHes into a server and runs
tail -f /var/log/app.log. - Searches for
ERRORwith grep. - Finds something, guesses a cause, restarts the service, hopes it goes away.
- Redeploys and watches manually for ten minutes.
This approach has two fundamental problems:
You can only ask questions you thought of in advance. If you did not add that specific log line before the incident, the evidence you need does not exist. You are flying blind on every incident you did not predict.
Instinct is not evidence. "I think it's the database" is not scoped, not time-bound, and not reproducible. It is a guess. Guesses lead to wrong fixes and repeated incidents.
Monitoring vs observability โ an important distinction¶
These terms are often used interchangeably but they describe different capabilities:
| Monitoring | Observability | |
|---|---|---|
| What it handles | Known-unknowns โ problems you anticipated | Unknown-unknowns โ questions you did not think to ask before the incident |
| Mental model | Set up dashboards and alerts for things you know can break | Instrument the system so you can ask new questions of live data without redeploying |
| Example | "Alert if CPU > 80%" | "What was the p99 latency of /shorten for users in the IN region at 14:32 yesterday?" |
| Limitation | Only catches what you predicted | Requires upfront instrumentation investment |
A well-instrumented system gives you both: pre-built dashboards for the things you expect, and raw queryable data for the things you did not.
๐ฎ๐ณ Hinglish intuition: M4 ka reconciliation loop ek chowkidar hai jo sirf ek sawaal poochta hai โ "kitne pod chahiye, kitne hain?" Observability alag cheez hai: CCTV + ek logbook + ek GPS tracker. CCTV (metrics) counts everything always. Logbook (logs) records what each event was. GPS tracker (traces) follows one customer's complete journey end-to-end. Chowkidar ko in teeno ki zaroorat hai โ sirf headcount kaafi nahi.
The three pillars¶
Every production system must answer three structurally different questions. Each question needs a different data shape โ which is why you run three separate tools, not one.
flowchart LR
APP["App Pod<br/>/metrics"]:::run
PROM[("Prometheus<br/>TSDB")]:::obs
GRAFANA{{"Grafana"}}:::obs
ALERT{{"Alertmanager"}}:::warn
PD(["PagerDuty<br/>Slack"]):::warn
APP -->|"scrape every 15s"| PROM
PROM --> GRAFANA
PROM -->|"alert rules"| ALERT
ALERT -->|"route"| PD
classDef run fill:#e0f2f1,stroke:#00897b,color:#004d40;
classDef obs fill:#f1f8e9,stroke:#689f38,color:#33691e;
classDef warn fill:#fdeeee,stroke:#d64545,color:#b23030;
Metrics path โ alerting spine: App exposes numbers โ Prometheus stores + evaluates โ Grafana shows, Alertmanager pages.
flowchart LR
APP2["App Pod"]:::run
PTAIL["Promtail<br/>DaemonSet"]:::run
OTEL["OTel Collector"]:::shared
LOKI[("Loki<br/>log store")]:::obs
JAEGER[("Jaeger<br/>trace store")]:::obs
GRAFANA2{{"Grafana"}}:::obs
APP2 -->|"stdout logs"| PTAIL
APP2 -->|"spans"| OTEL
PTAIL -->|"ship"| LOKI
OTEL -->|"forward"| JAEGER
LOKI --> GRAFANA2
JAEGER --> GRAFANA2
classDef shared fill:#fff9c4,stroke:#f9a825,color:#4a3800;
classDef run fill:#e0f2f1,stroke:#00897b,color:#004d40;
classDef obs fill:#f1f8e9,stroke:#689f38,color:#33691e;
Logs + traces path โ debugging spine: App streams events โ Promtail to Loki; App emits spans โ OTel to Jaeger; both queryable in Grafana.
Text version (ASCII)
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ THE THREE PILLARS โ
โ โ
โ QUESTION PILLAR TOOL DATA SHAPE โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ "How much / how many METRICS Prometheus Numeric โ
โ over time?" time-series โ
โ โ
โ "What exactly LOGS Loki / ELK Structured โ
โ happened here?" text events โ
โ โ
โ "Where did the time TRACES OTel + Timed span โ
โ go across services?" Jaeger trees โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Metrics say SOMETHING is wrong.
Logs say WHAT it was.
Traces say WHERE (which hop) it happened.
| Pillar | Question it answers | Data shape | Tool | Cost at scale |
|---|---|---|---|---|
| Metrics | How much / how many, over time? | Numeric time-series, cheap to store for years | Prometheus | Low |
| Logs | What exactly happened on this one request? | Unstructured / structured text | Loki (or ELK) | Mediumโhigh |
| Traces | Where did the time go, across services? | Tree of timed spans per request | OpenTelemetry + Jaeger | Medium |
The rule of thumb that saves hours at 2 a.m.: alert on metrics, debug with logs and traces. Alerting on logs directly (grep for ERROR every minute) does not scale past a handful of services โ that is exactly what metrics exist to replace.
Metrics โ Prometheus¶
Model. Prometheus is a fitness tracker that only counts numbers, never records why. It polls (scrapes) your app every N seconds โ a PULL model, the same pattern as Argo CD in M7, where the system reaches out rather than waiting to be pushed to. Your app exposes a /metrics HTTP endpoint in plain text. Prometheus scrapes it on a configurable interval, stores it in a local time-series database (TSDB โ Time Series DataBase), and makes it queryable via PromQL (Prometheus Query Language).
The four metric types:
| Type | Behavior | Use for | PromQL note |
|---|---|---|---|
| Counter | Only goes up; resets on restart | Total requests, total errors | Always query rate(), never raw value |
| Gauge | Goes up and down | Current memory, current queue depth | Query raw value is fine |
| Histogram | Buckets of observations (e.g. <100ms, <500ms, <1s) |
Request latency, response size | Enables p50/p95/p99 queries |
| Summary | Pre-computed quantiles, client-side | Legacy; rarely preferred now | Less flexible than histogram |
Numbers to know:
- Typical scrape interval: 15s (Prometheus binary default is 1m; kube-prometheus-stack and most production deployments configure 15s)
- Default local retention: 15d
- Prometheus is stateful (local TSDB) โ in production, teams run Thanos or Mimir for high availability and long-term storage beyond 15 days
๐ฎ Predict pehle (socho, phir aage padho): Tum ek metric mein
user_idlabel add kar dete ho (lakhon unique values). Prometheus ka kya hota hai?
The number-one real incident in this space โ cardinality explosion:
A Counter or Histogram with a label like user_id or request_path (when the path includes a dynamic ID) creates one new time-series per unique label value. Prometheus RAM usage scales with the count of unique time-series, not with data volume. 10,000 users ร 5 metrics = 50,000 series. Add one high-cardinality label and that becomes millions.
Teams have taken down their entire monitoring stack โ Prometheus OOMing and getting killed by K8s โ by adding a label with good intentions. The full war story is in the Real Production Example section below.
๐ง War story: Prometheus pod baar baar OOMKilled ho raha tha โ RAM double karne ke baad bhi; ek developer ne debugging ke liye
user_idlabel add kiya tha jisse 2M+ unique time-series ban gaye.curl .../api/v1/status/tsdbse root cause mila, label hatate hi fix hua. Poori kahani + lesson โ Interview Bank.
The label rule โ break it once and Prometheus dies in under 24 hours
Labels must have a small, bounded, predictable set of values:
| โ Safe (bounded) | โ Fatal (unbounded) |
|---|---|
method (GET/POST/โฆ) |
user_id (millions) |
status_code (~20 values) |
order_id, session_id |
route_template (/user/{id}) |
raw path (/user/48291, /user/48292, โฆ) |
RAM scales with unique time-series count, not traffic volume. More RAM does NOT fix it โ
the fix is deleting the label. Diagnose with curl .../api/v1/status/tsdb (top series counts).
๐ฎ๐ณ Hinglish intuition: metrics = har 15 second mein CCTV ki ek photo. Prometheus poochta hai "abhi kya count hai?" โ wo history nahi samajhta, sirf current number likhta. Cardinality explosion = har user ke liye alag CCTV channel kholna. Ek building mein 50,000 channels ka CCTV system koi run nahi kar sakta โ RAM khatam ho jaati hai.
Logs โ Loki / ELK¶
Model. Logs are the discrete event record โ what happened on each specific request. Loki is "Prometheus but for logs": same PromQL-style query language called LogQL, same label-based indexing, deliberately does not index full text (unlike Elasticsearch / ELK) to stay cheap at scale.
Mechanism. Your app writes logs to stdout โ never to a file inside the container (containers are cattle, per M0; when the pod dies, so does any file written inside it). A node-level log-shipper agent runs as a DaemonSet: Promtail (Loki's native shipper) or Fluent Bit (more universal). It tails container stdout, attaches Kubernetes labels (pod, namespace, app), and ships to Loki. You then query:
Structured logging vs print():
This is one of the highest-leverage habits a senior engineer has. Consider:
# Beginner: free-text print
print(f"shortened {code} for user {uid}")
# Senior: structured JSON log
logger.info("url_shortened", extra={"code": code, "user_id": uid,
"trace_id": trace_id, "duration_ms": elapsed})
The first produces a string you have to regex-parse at 2 a.m. The second produces a JSON object with named fields you can filter, aggregate, and join to traces by trace_id โ all in the query interface, without writing code.
Senior engineers standardize a JSON log schema across every service on day one. This is not premature optimisation; it is baseline infrastructure.
Log levels โ small but important. Every log line carries a level so you can separate signal from noise: DEBUG (dev detail, off in prod), INFO (normal events โ "request served"), WARN (odd but handled โ "retrying DB"), ERROR (a request failed), FATAL/CRITICAL (the process is dying). In production you run at INFO and alert on ERROR+ rates โ not on raw log volume. Make the level a field in the JSON ("level":"error"), not just free text, so LogQL/queries can filter on it. A flood of ERROR right after a deploy is often your fastest incident signal โ quicker than a metric alert's for: 5m window.
๐ฎ๐ณ Hinglish intuition: Log level = volume knob.
DEBUG= sab kuch (dev me),INFO= normal,WARN= "dekh lena",ERROR= "ab dekho",FATAL= "mar gaya". Prod meINFOpe chalao, page sirfERROR+ pe.
The stdout buffering pitfall. Some language runtimes buffer stdout when not attached to a TTY (a terminal). In production containers, there is no TTY. Logs appear in batches or not at all until the buffer flushes on process exit. Classic symptom: "the pod crashed but there are no logs." Fix: force unbuffered output โ in Python, PYTHONUNBUFFERED=1; in Node, logs go to process.stdout which is synchronous; in Go, the standard library does not buffer stdout.
Traces โ OpenTelemetry + Jaeger¶
Model. A trace is a relay race with a stopwatch at every handoff. One user request gets a trace_id โ a unique identifier that travels with the request across every service it touches. Each service adds one or more timed spans โ units of work with a start time, end time, and metadata. The result is a waterfall diagram showing exactly where the time went, down to the individual database query.
OpenTelemetry (OTel) is the vendor-neutral SDK and protocol for generating traces (and metrics and logs). You instrument once with OTel; you can send to Jaeger, Tempo, Zipkin, or a commercial backend without changing app code. Jaeger is the open-source backend for storing and visualising traces.
Mechanism:
- OTel SDK auto-instruments your framework (FastAPI, Express, Spring Boot) with minimal config โ often 3โ5 lines.
- On each inbound request, the SDK checks for a
traceparentHTTP header (W3C Trace Context standard). If present, it joins the existing trace. If absent, it creates a newtrace_id. - On each outbound call (HTTP, database, queue), the SDK injects
traceparentinto the request headers, propagating the trace downstream. - Spans are collected by an OTel Collector sidecar or DaemonSet, then forwarded to Jaeger.
- In Jaeger's UI, you click one
trace_idand see a full waterfall: which service took how long, where the tail latency lives.
The sampling cost. 100% trace sampling at high traffic is expensive โ each span is stored, indexed, and queried. Production systems use sampling strategies: - Head-based sampling: decide at the start of the request (e.g. "sample 1% of all requests"). - Tail-based sampling: collect all spans, but only persist the trace if it met a condition (was slow, had an error). More expensive to run but more useful โ you keep the interesting traces, drop the boring ones.
Where traces earn their pay. Tracing a single-service app is low-value โ a simple log gives you the same information. Tracing earns its keep the moment you have two or more services calling each other. When your metrics show p99 is 3 seconds but you cannot tell if the slowness is in your app code, the downstream Postgres query, or the Redis call โ a trace for one of those slow requests shows the span breakdown immediately.
๐ฎ๐ณ Hinglish intuition: trace = ek grahak ka poora safar โ restaurant mein ghusa, order diya, chef ne banaya, waiter ne laya. Har step pe stopwatch. Agar 20-minute delay hai to trace seedha batata: "chef ke paas 18 minute lage" โ order lene ya lany nahi. Metrics sirf bolta "20 minute lage." Trace batata "kahan."
SLIs, SLOs, SLAs, and error budgets¶
This section separates "I set up Grafana" from "I run production." It is how observability becomes a decision-making tool, not just a dashboard.
flowchart TD
SLI["SLI โ what you MEASURE<br/>% of /shorten requests<br/>non-5xx AND < 300ms"]:::obs
SLO["SLO โ your INTERNAL target<br/>99.5% over 30 days"]:::obs
EB{{"ERROR BUDGET = 100% โ SLO<br/>0.5% โ 3.6 hours/month of allowed pain"}}:::warn
SHIP["budget remaining > 0<br/>โ ship risky changes ๐"]:::run
FREEZE["budget exhausted<br/>โ reliability work ONLY ๐ง"]:::warn
SLA["SLA โ EXTERNAL contract<br/>99.0% or customer gets credit<br/>(always looser than SLO = your buffer)"]:::shared
SLI --> SLO --> EB
EB -->|"every deploy decision"| SHIP
EB -->|"every deploy decision"| FREEZE
SLO -.->|"breach warning zone"| SLA
classDef obs fill:#f1f8e9,stroke:#689f38,color:#33691e;
classDef warn fill:#fdeeee,stroke:#d64545,color:#b23030;
classDef run fill:#e0f2f1,stroke:#00897b,color:#004d40;
classDef shared fill:#fff9c4,stroke:#f9a825,color:#4a3800;
Measure (SLI) โ target (SLO) โ the leftover is your error budget, and THE BUDGET decides whether you ship or freeze โ not opinions. The SLA sits outside, looser, as the contractual last line.
Text version (ASCII)
| Term | Definition | Capstone example |
|---|---|---|
| SLI โ Service Level Indicator | The actual measured signal | % of /shorten requests returning non-5xx in <300ms |
| SLO โ Service Level Objective | Your internal target for that SLI | 99.5% of requests meet that bar over 30 days |
| SLA โ Service Level Agreement | The external, often contractual promise โ with financial penalty if missed | 99.0% uptime or customer gets credited |
| Error budget | 100% โ SLO โ how much failure you are allowed |
0.5% = ~3.6 hours/month of allowed degradation |
Why the SLO is always stricter than the SLA. You need a buffer between what you promise yourself and what you promise customers. If your SLO is 99.5% and you burn through it, your team goes into reliability-only mode โ no risky deploys. By the time you hit your SLA (99.0%), you should have already fixed the problem. If your SLO and SLA were the same number, you would have no warning before breaching the customer contract.
The senior insight most training skips โ the error budget as governance. The error budget turns "should we ship this risky change?" from a political argument into a number. If budget remains: ship. If budget is exhausted: only reliability work until the window resets. This is how Google SRE resolves the permanent tension between feature velocity and reliability โ without either side winning by politics. The number decides.
p99 vs. average โ why this matters for SLIs. Average latency hides the tail. If 95% of requests take 50ms and 5% take 5,000ms, the average is around 300ms โ looks fine. 1 in 20 users has a terrible experience. A histogram-based SLI on p99 would catch this; an average-based SLI would not. Always define SLIs using percentile latency, not mean.
๐ฎ๐ณ Hinglish intuition: error budget = kitni galti allowed hai. Agar 99.5% uptime ka SLO hai, to 0.5% โ yaani mahine mein roughly 3.6 ghante โ toot-phoot ka allowance hai. Jab tak budget hai, naya feature ship karo. Budget khatam? Rukjao โ pehle system theek karo. Yeh rule politics ko hataata hai โ number decide karta hai.
The Four Golden Signals¶
Google SRE defined four signals that together describe the health of almost any service. If you can only instrument four things, instrument these:
| Signal | What it measures | Example metric | Why it matters |
|---|---|---|---|
| Latency | How long requests take โ split by success and error | p99 of http_request_duration_seconds |
Slow is often worse than down โ users tolerate brief downtime but abandon slow apps |
| Traffic | How much demand is hitting the system | rate(http_requests_total[5m]) |
Baseline for capacity decisions and anomaly detection |
| Errors | The rate of requests that fail | rate(http_requests_total{status=~"5.."}[5m]) |
The most direct signal that users are experiencing failures |
| Saturation | How "full" the system is โ the resource closest to its limit | CPU throttling, memory usage %, queue depth | Predicts future failure before it becomes current failure |
Note on latency and errors: always track latency separately for successful and failed requests. A request that fails in 1ms inflates the "fast" bucket and masks a real problem if you are averaging across success and failure together.
PromQL โ 5 queries you must be able to write¶
SRE interviews regularly ask you to write PromQL on a whiteboard. Not "explain what Prometheus does" โ write the query. These five cover every query category that shows up.
Mental model: instant vector vs range vector.
- Instant vector โ one current sample per matching series:
http_requests_total. Good for gauges (current memory, queue depth). Can be filtered, aggregated, or compared directly. - Range vector โ a sliding window of samples per series:
http_requests_total[5m]. Required byrate(),irate(), andincrease(). You cannot pass an instant vector torate()โ it errors. The[5m]tells Prometheus "give me all scrape samples from the last 5 minutes" so the function can compute a slope.
Why rate() and not raw counter? A counter's cumulative total (e.g. 50,000) is meaningless by itself โ it accumulates since last restart. rate(counter[5m]) returns the per-second rate of increase over the last 5 minutes and handles counter resets (pod restarts) transparently. Always use rate() on counters, never the raw value.
1 ยท Request rate โ the Traffic signal¶
Returns the per-second request rate averaged over the last 5 minutes, broken out per label combination (pod, route, method, status).
When: baseline traffic check, capacity planning, spotting unexpected spikes or drops after a deploy. This is the "Traffic" signal in the Four Golden Signals above.
2 ยท Error rate %¶
Returns the percentage of all requests returning a 5xx status. status=~"5.." is a regex selector (=~ for match, !~ to exclude). sum() aggregates across all pods โ you want the service-wide error rate, not per-pod.
When: SLO burn rate evaluation, post-deploy health gate, incident triage โ "are users receiving errors right now?" This is your most direct SLI query.
3 ยท p99 latency โ histogram_quantile¶
Returns the 99th-percentile request duration in seconds over the last 5 minutes. le is the histogram's built-in "less than or equal" bucket boundary label โ each bucket counts requests that completed within that many seconds. sum by (le) aggregates bucket counts across pods while preserving bucket boundaries; histogram_quantile then interpolates the p99 from those buckets.
When: SLI dashboard panel, latency SLO alert, latency regression investigation. Replace 0.99 with 0.95 or 0.50 for p95/p50.
4 ยท Memory pressure โ OOM risk¶
Returns containers where working-set memory exceeds 85% of their configured limit. Use working_set_bytes, not usage_bytes โ working set excludes reclaimable file-cache pages, so it represents the RSS the kernel will not reclaim under pressure. That is the number that triggers OOMKill.
When: capacity review (M5), before a rollout that increases replica count or adds memory-heavy features, any time pods are OOMKilled and you want to see which containers were at risk.
5 ยท CPU throttling โ ties to M5¶
Returns the per-second rate of CPU time throttled by the Linux CFS scheduler. A non-zero value means the container hit its cpu.limits ceiling โ the kernel queued its CPU work โ even if the node had idle cores available. This is invisible to the app, but the app runs slower.
When: mysterious p99 spikes after a cpu.limits change, diagnosing latency you cannot explain from application code. High throttle rate + high p99 latency = strong signal to raise the limit or profile the hot path.
๐ฎ๐ณ Hinglish ek-liner:
rate()ek speedometer hai โ cumulative count nahi, "abhi per-second kitni baar ho raha hai" bolta hai.[5m]range vector ek sliding window hai. Bina[5m]kerate()error deta hai โ instant vector pe rate() chalti hi nahi.
Alertmanager routing โ the tree¶
Alertmanager ka routing ek labeled decision tree hai. Incoming alert ke labels check karo, sahi receiver ko bhejo. Yeh YAML structure interviews mein poochha jaata hai โ ek baar samjho, hamesha yaad rahega.
route:
receiver: slack-warnings # default fallback โ unmatched alerts yahan jaate hain
group_by: [alertname, service] # in labels pe group karo โ ek alert flood nahi, ek page
group_wait: 30s # 30s ruko taaki ek hi event ke sab alerts ek saath aayein
group_interval: 5m # already-notified group ke liye next notification interval
repeat_interval: 4h # still-firing alert ko 4 ghante baad dobara page karo
routes:
- match:
severity: critical
receiver: pagerduty-oncall # raat ko uthana zaroori โ PagerDuty
continue: false # match hone ke baad ruko; default receiver pe mat ja
- match:
severity: warning
receiver: slack-warnings # subah dekh lena kaafi โ Slack
receivers:
- name: pagerduty-oncall
pagerduty_configs:
- routing_key: <PD_INTEGRATION_KEY>
- name: slack-warnings
slack_configs:
- api_url: 'https://hooks.slack.com/services/YOUR/HOOK'
channel: '#alerts-warning'
inhibit_rules:
- source_match:
severity: critical
target_match:
severity: warning
equal: [service] # agar service=payments ka critical fire ho raha hai,
# usi service ka warning suppress karo โ actionable nahi hai
group_wait: 30s kyun? Ek hi node failure se 15 pods ke 15 alerts ek saath aate hain. 30 second ruko โ sab ek notification mein bundle ho jaate hain; bina wait ke 15 alag pages aa jaate aur on-call flood ho jaata.
inhibit_rules ka role: source_match (critical) fire ho raha ho aur equal: [service] match kare, to target_match (warning) suppress ho jaata hai. Agar service=payments ka critical already pata hai, to usi service ka disk-space warning bhejne ka koi matlab nahi โ pehle bada theek karo.
๐ฎ๐ณ Hinglish ek-liner: Alertmanager = smart courier. Critical ko PagerDuty (seedha call, raat ko uthao), warning ko Slack (subah dekh lena). Aur agar bada problem pehle se known hai, chhote related alerts mat bhejo โ noise sirf distract karta hai.
Alerting that does not page you for nothing¶
How Alertmanager works¶
Prometheus evaluates alert rules โ PromQL expressions that should evaluate to false during normal operation. When an expression evaluates to true (e.g. error rate exceeds threshold), Prometheus fires an alert to Alertmanager (Alert Manager).
Alertmanager's job is not to forward every alert. It does four things:
- Deduplication: if 50 pods all fire the same alert at the same time, Alertmanager sends one page, not 50.
- Grouping: related alerts (same service, same time window) arrive as one notification with context, not as a flood.
- Routing: based on labels (
team="payments",severity="critical"), Alertmanager routes the alert to the right channel โ Slack, PagerDuty, email, webhook. - Silences and inhibitions: a silence suppresses specific alerts during a known maintenance window. An inhibition rule suppresses downstream alerts when a higher-priority alert is already firing โ do not page for "disk full" on a node when "node down" is already firing for the same node. The disk problem is not actionable while the node is unreachable.
flowchart TD
RAW["50 raw firings"]:::warn
DEDUP["Deduplication<br/>same fingerprint โ 1 alert"]:::obs
GROUP["Grouping<br/>3 groups by alertname + service"]:::obs
ROUTE{"Route by label"}:::obs
PD["PagerDuty"]:::warn
SLACK["Slack"]:::shared
INHIB{"Critical parent<br/>already firing?"}:::warn
SUPP["Suppressed<br/>not actionable"]:::net
PAGE["1 actionable page"]:::run
RAW --> DEDUP
DEDUP --> GROUP
GROUP --> ROUTE
ROUTE -->|"team=backend"| PD
ROUTE -->|"severity=warning"| SLACK
PD --> INHIB
INHIB -->|"yes"| SUPP
INHIB -->|"no"| PAGE
classDef shared fill:#fff9c4,stroke:#f9a825,color:#4a3800;
classDef run fill:#e0f2f1,stroke:#00897b,color:#004d40;
classDef obs fill:#f1f8e9,stroke:#689f38,color:#33691e;
classDef net fill:#e3f2fd,stroke:#1976d2,color:#0d47a1;
classDef warn fill:#fdeeee,stroke:#d64545,color:#b23030;
Alertmanager ki value = 50 noisy firings โ 1 actionable page.
Alert on symptoms, not causes¶
This is the single most common mistake in on-call setups:
| Pattern | Example | Problem |
|---|---|---|
| Cause-based (bad) | CPU > 80% |
CPU at 80% might be totally fine โ high CPU during a batch job is expected. Users notice nothing. |
| Symptom-based (good) | error_rate > 1% for 5m |
Users are receiving errors right now. This is unambiguously actionable. |
Alerting on causes produces alert fatigue โ too many pages that require human judgement to determine whether they matter. Engineers start ignoring pages. Real incidents get lost in the noise.
The rule: page on what users feel (symptoms). Put everything else on a dashboard panel โ visible but silent.
The for: duration rule¶
# BAD: a single scrape blip pages someone at 3 a.m.
- alert: HighErrorRate
expr: rate(http_requests_total{status=~"5.."}[5m]) > 0.01
# GOOD: transient spikes self-resolve before anyone is paged
- alert: HighErrorRate
expr: rate(http_requests_total{status=~"5.."}[5m]) > 0.01
for: 5m
labels:
severity: critical
annotations:
summary: "Error rate above 1% for 5 minutes on {{ $labels.service }}"
The for: 5m field tells Prometheus to hold the alert in a pending state for 5 minutes before it fires. A transient network blip that lasts 30 seconds resolves itself; a real problem persists. In practice, the duration threshold matters more than the numeric threshold โ it is the difference between waking someone up for nothing and waking them up for a real incident.
Request lifecycle, instrumented¶
This extends the request lifecycle from 09-connected-system.md โ same request, now with every telemetry emission point annotated.
REQUEST LIFECYCLE WITH TELEMETRY EMISSION POINTS
1. User โโโบ POST :30080/shorten
โ
2. NodePort โ kube-proxy โ Service โ EndpointSlice โ Pod
โ
โ [TRACE] span started: trace_id=abc123, service=urlshort
โ
3. Pod (FastAPI) handles the request
โ
โ [METRIC] http_requests_total{route="/shorten", method="POST"} += 1
โ [METRIC] http_request_duration_seconds histogram observes start
โ [LOG] {"event":"request_received","route":"/shorten",
โ "trace_id":"abc123","ts":"2025-07-02T02:14:03Z"}
โ
4. Pod โโโโโโโบ Postgres :5432 INSERT INTO urls ...
โ
โ [TRACE] child span: db_insert
โ trace_id=abc123, parent=span1, start=+002ms
โ
5. Postgres responds
โ
โ [TRACE] child span closed: db_insert duration=031ms
โ [LOG] {"event":"db_insert_ok","code":"xyz7k",
โ "duration_ms":31,"trace_id":"abc123"}
โ
6. Response โโโบ User HTTP 201
โ
โ [TRACE] root span closed: total_duration=043ms
โ [METRIC] http_request_duration_seconds observes 0.043
โ [LOG] {"event":"request_complete","status":201,
โ "duration_ms":43,"trace_id":"abc123"}
โ
7. (async, every 15s)
Prometheus scrapes Pod /metrics endpoint
โ TSDB records new counter and histogram values
8. (continuous, via DaemonSet)
Promtail/Fluent Bit tails pod stdout
โ ships structured log lines to Loki
9. (every 15โ30s, rule evaluation)
Prometheus evaluates: rate(http_requests_total{status=~"5.."}[5m]) > 0.01 for 5m?
โ FALSE โ nothing happens โ this is the normal, silent case
โ TRUE โ Alertmanager fires โ routes to Slack/PagerDuty โ M9 incident begins
sequenceDiagram
participant User
participant App as App Pod
participant PG as Postgres
participant Prom as Prometheus
participant PT as Promtail
participant OTel as OTel Collector
User->>App: POST /shorten
Note right of App: LOG written inline during handler<br/>TRACE root span started
App->>PG: INSERT INTO urls
PG-->>App: rows
Note right of App: LOG db_insert_ok<br/>TRACE db_insert child span closed
App-->>User: HTTP 201
Note right of App: TRACE root span exported on completion<br/>METRIC histogram observes 43ms
App-->>PT: stdout log lines
Note over PT: ships log batch to Loki every ~5s
Prom->>App: GET /metrics
App-->>Prom: metrics payload
Note over Prom: scrapes every 15s โ incident ke pehle<br/>15s mein koi metric data nahi hota
Har signal ALAG time pe fire hota โ Prometheus ke paas incident ke pehle 15s ka data nahi hota.
The โคท emissions at steps 3, 4, 5, 6 happen on every single request, always โ they are not conditional debugging code turned on during incidents. This is the mental shift: instrumentation is baseline infrastructure you build in before anything breaks, so that when it does break you already have the evidence. Adding instrumentation after an incident to understand the last incident is too late.
Real production example โ the cardinality explosion¶
This is the most common class of observability incident across the industry. The pattern is documented widely and most teams encounter it eventually.
The setup. A team wants to debug a slow endpoint. They add a new metric:
In staging, there are a handful of routes. Looks fine. In production, the path includes a dynamic resource ID:
Each unique path value creates a new time-series. With thousands of users per hour, the series count goes from ~50,000 (healthy) to several million within 24 hours.
What happens next. Prometheus's in-memory series count explodes. RAM usage climbs. Kubernetes OOM-kills the Prometheus pod. The pod restarts, loads the on-disk TSDB, and OOMs again. Dashboards go blank during an unrelated ongoing incident โ the observability layer became the outage, blinding the team at the worst possible moment.
The fix โ two layers.
Layer 1 โ fix the label immediately:
# WRONG: raw path with dynamic segment
request_duration.labels(path=request.url.path).observe(elapsed)
# CORRECT: route template โ bounded set of values
request_duration.labels(route="/users/{id}/orders/{order_id}").observe(elapsed)
Layer 2 โ add a cardinality guardrail in the Prometheus scrape config so no single bad label can repeat this:
# prometheus.yml scrape config
scrape_configs:
- job_name: urlshort
metric_relabel_configs:
- source_labels: [__name__]
regex: 'http_request_duration_seconds'
target_label: path
replacement: '' # drop the 'path' label entirely on this metric
The lesson that generalises. Observability infrastructure is itself production infrastructure. It needs the same sizing and reliability discipline as the apps it watches (M5). A bug in the monitoring layer can blind you during the exact moment you need it most. Senior engineers ask "can my monitoring survive the incident it is supposed to help me see?" as a real design question โ not an afterthought.
Commands, explained¶
# Port-forward Prometheus to localhost:9090
# Why: access the UI and HTTP API without a LoadBalancer; works in any cluster
# Note: the exact Service name depends on the Helm release/chart used;
# kube-prometheus-stack (used in the Hands-on Lab) names it as shown below.
kubectl port-forward svc/monitoring-kube-prometheus-prometheus 9090:9090 -n monitoring
# Port-forward Grafana to localhost:3000
# Why: same reason โ dev and debugging access without exposing a public endpoint
kubectl port-forward svc/grafana 3000:80 -n monitoring
# Check what Prometheus is currently scraping and each target's last-scrape status
# Why: most "I don't see my app's metrics" problems are a misconfigured ServiceMonitor
kubectl exec -it <prometheus-pod> -n monitoring -- \
wget -qO- localhost:9090/api/v1/targets | python3 -m json.tool | grep -E '"health"|"scrapeUrl"'
# Query Prometheus via HTTP (no UI needed) โ 5xx rate over last 5 minutes
# Why: useful in CI, in scripts, or when port-forwarding Grafana is not convenient
curl -s 'http://localhost:9090/api/v1/query' \
--data-urlencode 'query=rate(http_requests_total{status=~"5.."}[5m])'
# Query Prometheus โ p99 latency of /shorten
# Why: this is the SLI query; put it in your dashboard and alert rule
curl -s 'http://localhost:9090/api/v1/query' \
--data-urlencode 'query=histogram_quantile(0.99, rate(http_request_duration_seconds_bucket{route="/shorten"}[5m]))'
# Tail logs from your app via Loki CLI
# Why: faster than the Grafana UI for quick label-filtered tailing in a terminal
logcli query '{app="urlshort", namespace="prod"} |= "ERROR"' --limit=50 --tail
# Evaluate what an alert rule currently resolves to
# Why: before adding the rule, verify the PromQL expression returns the values you expect
kubectl exec -it <prometheus-pod> -n monitoring -- \
promtool query instant http://localhost:9090 \
'rate(http_requests_total{status=~"5.."}[5m])'
# Check Prometheus TSDB statistics including top series by metric name
# Why: first diagnostic when suspecting cardinality explosion
curl -s 'http://localhost:9090/api/v1/status/tsdb' | python3 -m json.tool | head -60
Beginner mistakes vs senior insights¶
| Beginner | Senior |
|---|---|
| "I added Grafana, we are observable now." | A dashboard nobody looks at until something is already broken is not observability โ it is decoration. Observability means alerts that page before users complain. |
| Alerts on every metric that could matter: CPU, memory, disk, queue depth, JVM heap... | Alerts only on SLO burn rate (symptom). Everything else is a dashboard panel โ visible, not loud. |
Alert rule with no for: duration โ a single scrape blip pages on-call at 3 a.m. |
Always pair a threshold with a for: duration. Transient noise self-resolves; real problems persist. |
Adds a user_id label to a metric "to make debugging easier." |
Labels must have a bounded set of values. A user_id label in a high-traffic app destroys Prometheus in under 24 hours. |
| "It's slow" (vague, unscoped). | "p99 on /shorten went from 80ms to 1.2s starting at 14:32, correlates with the RDS CPU spike โ likely a slow query. Here is a trace ID for one of the slow requests." Evidence-based, time-bound, actionable. |
Logs with print(f"done {code}") โ free text. |
Structured JSON logs with a trace_id field on every line, so any log line can be joined to its full trace and correlated metric context without guessing. |
| Builds observability in a "monitoring namespace" that has no resource limits, no alerts on Prometheus itself, no redundancy. | Treats observability infra as production infra: resource requests/limits, a watchdog alert on Prometheus's own health, and (in large orgs) Thanos/Mimir for HA. |
Memory shortcuts¶
| Concept | One line to recall it |
|---|---|
| Metrics / logs / traces | How much / what exactly / where did the time go |
| Prometheus pull model | Same PULL pattern as Argo โ scrapes every 15s, never gets pushed to |
| Counter vs Gauge | Counter only goes up (query rate()); Gauge goes up and down (query raw) |
| Histogram | Buckets of observations โ enables p99; this is why you use it for latency |
| Cardinality | Unique label combinations = RAM cost in Prometheus; keep labels bounded |
| p99 > average | Average hides the tail that users feel; histogram quantile catches it |
| SLI / SLO / SLA | Measured / your target / contractual promise (SLA always looser than SLO) |
| Error budget | 100% โ SLO = allowed failure; exhausted โ freeze risky deploys |
| Four Golden Signals | Latency, Traffic, Errors, Saturation |
| Alert on symptoms | Error rate pages you; CPU-at-80% stays on a dashboard |
for: 5m |
Holds alert in pending โ noise self-resolves, real problems do not |
| Structured logging | JSON with trace_id on every line โ queryable, joinable, not regex'd |
| Tail-based sampling | Keep traces that errored or were slow; drop boring ones โ cost control |
Summary¶
Observability is the layer that turns "the pod is running" into "the system is healthy, and I have evidence." It has three pillars answering three structurally different questions: metrics (how much), logs (what exactly), traces (where). Each requires different infrastructure because each produces a different data shape.
On top of the technical pillars sits the SRE model: SLIs define what you measure, SLOs define your target, the error budget (100% โ SLO) defines how much failure you are allowed, and that number governs whether risky changes ship or freeze. This turns reliability from a vague aspiration into a concrete, daily decision input.
Alerting done right pages on symptoms (error rate, SLO burn rate), not causes (CPU, disk). The for: duration is what separates a page that wakes someone up for nothing from one that wakes them up for a real incident.
The cardinality lesson is the one that bites almost every team once: high-cardinality labels destroy Prometheus from the inside. The fix is label discipline upfront, plus a cardinality guardrail in the scrape config.
Observability infrastructure is itself production infrastructure. If it goes down during an incident, you are flying blind exactly when you need it most.
Self-check quiz¶
Pehle memory se jawab do, phir neeche kholo.
Jawab dekho
- Log-based alerting scale nahi karta โ har rule ke liye ek grep/regex job chahiye, high-volume logs pe aggregate queries slow hoti hain, aur volume badhne par load bahut zyada hota hai. Metrics (pre-aggregated counters/histograms) isi liye exist karte hain โ log-grepping replace karne ke liye alerting mein.
{route, method, status, user_id}โuser_idke millions of unique values hain โ millions of unique time-series โ Prometheus RAM explode โ OOMKilled. Bounded labels (route, method, status) safe hain; unbounded labels (user_id, order_id) cardinality explosion laate hain.- Majority requests fast hain lekin ek tail bahut slow hai (4.8s). p99 pe alert karo โ yeh tail ko pakadta hai joh users feel karte hain. Average mein yeh problem chhup jaati hai, alert nahi aata.
- SLO internal target hai (99.5%); SLA external/contractual promise hai (99.0%). SLO hamesha stricter hota hai taake SLO breach hone par team pehle hi act kare โ SLA breach se pehle warning buffer milta hai.
- Deduplication (ek node se 40 alerts โ 1 page) aur inhibition (downstream alerts suppress ho jaate hain jab higher-priority "node down" alert pehle se fire ho raha ho โ disk-full page actionable nahi jab node hi unreachable ho).
- Threshold decide karta hai kab condition trigger ho;
for:decide karta hai kab page jaaye. 30-second transient spike threshold meet karta hai โ binafor:ke bhi 3am page. Real incidents persist karte hain; noise nahi karta. Duration zyada matter karta hai. - Sabse zyada likely: cardinality explosion โ high-cardinality label ne millions of unique time-series bana diye. Pehla diagnostic:
curl -s 'http://localhost:9090/api/v1/status/tsdb' | python3 -m json.tool | head -60se series count aur top metrics dekho. - Application code fast hai; bottleneck ek slow DB insert (1.9s) hai. Next: RDS slow query logs check karo, missing index dhundho, DB CPU/connections aur Prometheus mein connection pool saturation dekho.
- Why can you not simply alert directly on raw log lines at scale? What breaks first?
- A metric has labels
{route, method, status}versus{route, method, status, user_id}. Which risks cardinality explosion, and why specifically? - Average latency is 120ms. p99 latency is 4.8 seconds. What is actually happening? Which number do you alert on and why?
- What is the practical difference between an SLO and an SLA, and why is the SLO always the stricter number?
- Alertmanager fires 40 pages in 90 seconds when one node goes down. Name the two Alertmanager features that would have prevented this.
- Why does the
for: 5mfield on an alert rule matter more than the threshold value in practice? - Your Prometheus pod is OOMing every hour. No application incidents are occurring. What is the most likely root cause and what is the first diagnostic command you run?
- A trace shows that a
/shortenrequest took 2.1 seconds total, with 1.9 seconds spent in a single child span labelleddb_insert. What does this tell you, and what do you do next?
Hands-on lab โ instrument an app on kind¶
โ
Prove it โ bash labs/check-m8-observability.sh
Lab ho gaya? Tick mat lagao โ machine se verify karo (monitoring stack ยท ServiceMonitor ยท kubectl top ยท HPA). โ pe exact fix-hint. โ The Doer's Path
Do raaste โ koi ek chuno (ye lab Capstone I pe nirbhar nahi):
- A ยท Standalone (recommended agar capstone nahi bana): ek chhota sample app kind pe daalo aur usi ko instrument karo โ
(Ya M4 wala
kubectl create deployment sampleapp --image=nginx:1.27 --replicas=2 kubectl expose deployment sampleapp --port=80nginx-appreuse karo.) Aage ke steps meurl-shortenerki jagahsampleapppadho. - B ยท Capstone extension: agar tumne Capstone I bana liya hai, to seedha
url-shortener/ko instrument karo โ koi naya project nahi.
Prerequisites¶
- Ek Kind cluster (standalone ke liye
kind create clusterkaafi; ya capstone ka cluster) helminstalled โ never used Helm? Read M7.5 ยท Helm & Kustomize Primer first (25 min); this lab only needshelm repo add+helm install- A free Slack workspace with an incoming webhook URL
Is your cluster Argo-managed? Pause auto-sync first
If the url-shortener is deployed via Argo CD with selfHeal: true (M7), the manual
kubectl set env breakage in Step 7 will be auto-reverted within seconds โ Argo wins,
your experiment silently un-breaks itself. Before Step 7:
argocd app set url-shortener --sync-policy none (re-enable after). This is not a bug โ
it is M7's whole lesson happening to you in real time.
Step 1 โ expose metrics from the app (3 lines)¶
# In url-shortener/app/main.py โ add after imports
from prometheus_fastapi_instrumentator import Instrumentator
# After `app = FastAPI()`
Instrumentator().instrument(app).expose(app)
This auto-exposes /metrics with request count and latency histogram, labeled by route template (not raw path โ the cardinality lesson, applied immediately).
pip install prometheus-fastapi-instrumentator
# Rebuild your image and push
docker build -t urlshort:v2-obs .
kind load docker-image urlshort:v2-obs
Step 2 โ deploy the Prometheus stack¶
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
helm install monitoring prometheus-community/kube-prometheus-stack \
-n monitoring --create-namespace \
--set grafana.adminPassword=admin123
This single Helm release deploys Prometheus, Grafana, Alertmanager, and a set of default K8s dashboards.
Step 3 โ tell Prometheus about your app (your first CRD)¶
# monitoring/servicemonitor.yaml
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: urlshort
namespace: monitoring
labels:
release: monitoring # must match the Helm release label
spec:
selector:
matchLabels:
app: urlshort
namespaceSelector:
matchNames:
- default
endpoints:
- port: http
path: /metrics
interval: 15s
kubectl apply -f monitoring/servicemonitor.yaml
# Verify: Prometheus UI โ Status โ Targets โ urlshort should appear as UP
kubectl port-forward svc/monitoring-kube-prometheus-prometheus 9090:9090 -n monitoring
A ServiceMonitor is a Custom Resource Definition (CRD) โ your first preview of the CRD/operator pattern covered in M9.
Step 4 โ build one Grafana panel¶
kubectl port-forward svc/monitoring-grafana 3000:80 -n monitoring
# Open http://localhost:3000 โ login admin / admin123
Create a new dashboard. Add a panel with this PromQL:
histogram_quantile(0.50, rate(http_request_duration_seconds_bucket{route="/shorten"}[5m]))
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{route="/shorten"}[5m]))
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket{route="/shorten"}[5m]))
Three lines on one panel: p50, p95, p99 latency of your endpoint. This is your SLI panel.
Step 5 โ write one alert rule¶
# monitoring/alert-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: urlshort-alerts
namespace: monitoring
labels:
release: monitoring
spec:
groups:
- name: urlshort.rules
interval: 30s
rules:
- alert: HighErrorRate
expr: |
rate(http_requests_total{job="urlshort", status=~"5.."}[5m])
/
rate(http_requests_total{job="urlshort"}[5m])
> 0.01
for: 5m
labels:
severity: critical
team: backend
annotations:
summary: "Error rate above 1% on urlshort"
description: "{{ $value | humanizePercentage }} of requests are failing."
Step 6 โ route to Slack¶
In Alertmanager's config (via the Helm values), add a Slack receiver:
alertmanager:
config:
route:
receiver: slack-notifications
receivers:
- name: slack-notifications
slack_configs:
- api_url: 'https://hooks.slack.com/services/YOUR/WEBHOOK/URL'
channel: '#alerts'
title: '{{ .GroupLabels.alertname }}'
text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'
Step 7 โ prove it fires¶
Break the database connection string deliberately in a test branch:
Watch the error rate climb. Within approximately 5โ6 minutes (scrape interval ร for: window), the alert moves from pending to firing in the Prometheus UI and a message arrives in Slack.
Then revert:
kubectl set env deployment/urlshort DATABASE_URL="postgresql://postgres:postgres@postgres-svc/urldb"
Step 8 โ commit to GitOps¶
git add monitoring/servicemonitor.yaml monitoring/alert-rules.yaml
git commit -m "feat(obs): add Prometheus ServiceMonitor and error rate alert for urlshort"
git push
Argo CD (from M7) picks this up and applies the monitoring config. Your alert rules are now declarative and drift-proof โ they live in Git like everything else in the system.
โ
Sahi hua to aisa dikhega: Grafana panel mein /shorten ke p50/p95/p99 latency ki teen alag lines dikh rahi hain; jab kubectl set env se DB URL galat karo, Prometheus UI mein alert PENDING se FIRING mein jaata hai aur lagbhag 5โ6 minute mein Slack channel mein HighErrorRate message aa jaata hai; DB URL revert karne par alert wapas resolve ho jaata hai aur Grafana mein error rate zero pe aa jaata hai.
Interview questions¶
-
"Design the alerting strategy for a payments API. Walk me through what you page on versus what goes on a dashboard, and why." Expected direction: page on SLO burn rate, success rate, and p99 latency crossing a threshold for a sustained duration. Dashboard-only: CPU, memory, queue depth, disk. Explain the symptom-vs-cause distinction and why cause-based alerts cause fatigue.
-
"Explain p99 latency versus average. Why is average a bad SLI for most services?" Expected direction: average hides the tail. If 95% of requests are fast and 5% are very slow, the average looks fine while 1-in-20 users has a bad experience. p99 catches this; it represents the experience of the slowest 1% of users, which is often a canary for broader degradation. Histograms enable percentile queries; averages do not.
-
"Your error budget is exhausted 10 days into a 30-day window. What actually changes in how your team operates?" Expected direction: no new risky feature deployments until the budget recovers. The team shifts to reliability work: fixing the root cause of the failures that burned the budget, improving test coverage, running postmortems. Describe how this is a governance decision, not a political one.
-
"A team adds a
user_idlabel to their request counter metric and within 24 hours Prometheus is OOMing. What happened and how do you fix it, both immediately and structurally?" Expected direction: cardinality explosion โ millions of unique time-series from unique user IDs. Immediate fix: drop the label at the Prometheus scrape relabelling config. Structural fix: replace with a route template label, enforce cardinality limits at the scrape level, add a Prometheus self-monitoring alert onprometheus_tsdb_head_series > threshold. -
"Metrics, logs, and traces all show green, but a customer calls to say the app is down for them. Where do you look?" Expected direction: your instrumentation only covers what it covers. Possible gaps: CDN or DNS issue outside your cluster, a client-side JavaScript error, a mobile network problem, a synthetic monitor that would have caught this vs real-user monitoring. The answer surfaces the limits of server-side telemetry and introduces the concept of synthetic monitoring.
-
"How would you prevent a single engineer's bad metric label from taking down Prometheus for the whole organisation?" Expected direction: metric relabelling rules to drop or cap high-cardinality labels, per-job cardinality limits in Prometheus config, alerting on
prometheus_tsdb_head_series, code review gates for any newlabels()call in application metrics definitions, and in large orgs โ a separate Prometheus per team so one team's explosion cannot affect another's. -
"What is tail-based trace sampling and when would you use it over head-based sampling?" Expected direction: tail-based sampling collects all spans but only persists a trace if it meets a post-hoc condition (slow, errored). More useful because you keep the interesting traces and drop the boring ones. More expensive to run (need to buffer spans before deciding). Use it when you need high fidelity on errors and outliers but cannot afford 100% sampling at production traffic volumes.
Production challenge¶
Your
/shortenendpoint's p99 latency has crept from 100ms to 1.5 seconds over two weeks. No alert fired โ you are only alerting on error rate, not latency. A customer complained. Using only what M8 gives you โ metrics, logs, and traces โ design your investigation path in order and state what each step rules in or out. Then design the SLO and alert rule that would have caught this on day 3, not day 14.
Guidance: start with the metric to scope the time range and whether it is request-wide or endpoint-specific; move to a trace for one of the slow requests to identify which span is slow; move to logs for that trace_id to see any warnings around the slow span. For the SLO design, define the SLI (p99 < X ms), the SLO target, and write the PromQL for an alert with a for: duration that would have fired early enough to be useful.
Next: 11-M9-advanced-k8s-internals.md โ when an alert fires and you need to investigate a running cluster deeply: node pressure, eviction thresholds, priority classes, and why your critical pod got evicted during a node memory spike.
For the incident response flow โ what you actually run the moment an Alertmanager notification fires โ see 23 ยท Production Incident Playbook (Part III, right after M9). Advanced incident-response theory lives in 15 ยท M11โM18 Roadmap.