SproutStackSproutStack home

LLM Observability & Tracing

2 min readBeginner-friendlyPremium
Premium chapter — free for early learners.

Read a little, play a little. No scary maths, and no rush.

Share:

You can't set a breakpoint inside a model

Traditional debugging steps through code line by line. An LLM call is a black box between input and output, so observability has to work from the outside: capture every prompt, every tool call, every retry, and every token/cost/latency number around that black box, then look for patterns across many requests instead of stepping through one.

Traces and spans

A trace is everything that happened for one user request. A span is one step inside it — one model call, one tool invocation, one retrieval. Nest them: a RAG answer might be one trace containing a retrieval span, a rerank span, and a generation span. Without this structure, "the agent gave a weird answer" has no path back to which step went wrong — the lookup, the ranking, or the final generation.

What's worth logging on every call

Prompt and completion text (redact secrets, not substance — you need to read it later), token counts and cost, latency, model/version, and any tool calls with their arguments and results. Model/version matters more than it looks: a silent provider-side model swap can shift behavior with no code change on your side, and without a logged version you can't tell "we broke something" from "they changed something."

The failure mode this catches

Most LLM failures aren't crashes — they're confident wrong answers that return a clean 200. Uptime dashboards miss these entirely. Pairing traces with a small sample of human-reviewed outputs (or an LLM-as-judge pass) is how teams catch quality regressions that no error rate would ever surface.

Remember this

  • Observability replaces breakpoints: capture everything around the black box.
  • Traces nest spans — one user request can be many model/tool calls deep.
  • Log model/version too — the model underneath you can change without your code changing.

Check your understanding

2 questions · correct answers earn XP once each

1. A trace in an LLM app is…
2. Why log prompts/outputs, not just success/failure?

My notes

Saved in this browser. Highlight a line above and save it, or write it in your own words.

Nothing saved yet. Your highlights will live here.

References

Finished reading?

Ticking it here also ticks the chapter in the sidebar, the section count and your streak — it is all one number.

Related chapters

Spotted a mistake or want a topic covered? Report an issue