5 · Production (MLOps + LLMOps) › Phase 6 › Lesson 3 of 5
Observability that survives a 3am page
What this lesson covers
Debugging a system that fails 0.5% of the time and can’t be reproduced locally is the moment observability stops being nice-to-have and becomes the only tool. This lesson is what to instrument so that moment is a thirty-minute page and not a two-week investigation.
The outline
- The trace-span mental model. A trace is one user request; a span is one call within it. Why this is the primitive that makes LLM systems debuggable.
- What to log at each span. Model, prompt, response, latency, tokens, cost. Why redacting the prompt is a lie you’ll regret.
- The two dashboards. “Is it broken” (error rates, p95 latency, cost per request) and “why is it slow” (span breakdown, tail latency).
- LangSmith and Braintrust. What each buys you. When they’re worth it and when you should roll your own.
- Sampling. Why full-fidelity logging costs more than the model itself past a certain scale, and how to sample without losing the debug traces you actually needed.
- Correlating telemetry with users. So a customer report becomes a trace lookup, not a needle-in-haystack search.
Coming soon
In outline.