Observability: Logs, Metrics, Traces
Read a little, play a little. No scary maths, and no rush.
The 3am question
Somebody pages you: "checkout is broken for users in Germany." You cannot debug that with a hunch. You need three answers, in this order: is it broken, where is it broken, and why. Three tools, three jobs. Confusing them is why on-call nights go badly.
Metrics: the number on the wall
A metric is an aggregate over time — requests per second, p99 latency, error rate, queue depth, CPU. Cheap to store, cheap to graph, and the right thing to alert on because one number can page a human. Nobody wants an alert that says "log line count changed". A good alert fires on a symptom users feel ("1% of checkouts failing"), not a cause ("CPU at 80%").
Logs: what one instance did
A log is a timestamped event from one process: "payment attempt 3 timed out after 2s". Rich and specific, and expensive — a busy service writes gigabytes a minute. Logs are for investigating after you know something is wrong, not for watching. Give every request a correlation id so one user's flow can be stitched together across services.
Traces: what one request did
A trace follows a single request across every service it touches, with each span timed. This is the one that answers "why was this request slow?" — because you can see that 800ms of a 900ms request went into one downstream call, and which one. This is what OpenTelemetry exists to standardize.
request #7f3a total 912ms
├─ api-gateway 4ms
├─ auth 11ms
├─ cart-service 38ms
├─ pricing-service 9ms
├─ inventory-service 62ms
└─ payment-service 780ms ← here, and it timed out twice first
The rule that keeps it cheap
Instrument metrics first (cheap, always on), traces second (sample them — 1% of requests is usually enough to find a pattern), and logs always but at a sensible level. Also log at info by default and debug only when switched on; a library that logs by default will be removed from production to save money.
Remember this
- Metrics tell you something is broken. Traces tell you where. Logs tell you why one instance did what it did.
- Alert on user-visible symptoms, not on causes.
- Correlation ids are what turn a wall of logs into a single story.
Check your understanding
2 questions · correct answers earn XP once each
My notes
Saved in this browser. Highlight a line above and save it, or write it in your own words.
Nothing saved yet. Your highlights will live here.
References
Finished reading?
Ticking it here also ticks the chapter in the sidebar, the section count and your streak — it is all one number.
Related chapters
Spotted a mistake or want a topic covered? Report an issue