aiengineering.guideaiengineering.guide

5 · Production (MLOps + LLMOps) › Phase 6 › Lesson 3 of 5

Observability that survives a 3am page

What this lesson covers

Debugging a system that fails 0.5% of the time and can’t be reproduced locally is the moment observability stops being nice-to-have and becomes the only tool. This lesson is what to instrument so that moment is a thirty-minute page and not a two-week investigation.

The outline

  1. The trace-span mental model. A trace is one user request; a span is one call within it. Why this is the primitive that makes LLM systems debuggable.
  2. What to log at each span. Model, prompt, response, latency, tokens, cost. Why redacting the prompt is a lie you’ll regret.
  3. The two dashboards. “Is it broken” (error rates, p95 latency, cost per request) and “why is it slow” (span breakdown, tail latency).
  4. LangSmith and Braintrust. What each buys you. When they’re worth it and when you should roll your own.
  5. Sampling. Why full-fidelity logging costs more than the model itself past a certain scale, and how to sample without losing the debug traces you actually needed.
  6. Correlating telemetry with users. So a customer report becomes a trace lookup, not a needle-in-haystack search.

Coming soon

In outline.

Outline

Enter to go · Esc to close