aiengineering.guideaiengineering.guide

5 · Production (MLOps + LLMOps) › Phase 6 › Lesson 4 of 5

Caching and cost control

What this lesson covers

The gap between a bill you can defend and one you can’t is usually caching, done right. This lesson names the three kinds, when each applies, and the ways they quietly fail.

The outline

  1. Prompt caching, provider-side. How Anthropic and OpenAI’s cache mechanisms work, the cache-key rules that break naive setups, and the discount you actually get.
  2. The cache-friendly prompt shape. Static content first, dynamic content last — and how to refactor a prompt that gets this wrong.
  3. Response caching. When identical requests deserve a cached answer, and how to key the cache without leaking one user’s answer to another.
  4. Semantic caching. Answering “close-enough” requests from cache. The single most misused optimisation in AI systems — when it works and how it goes wrong.
  5. Cost budgets and alerts. Per-tenant, per-feature, per-day. The alerts that catch a bill going sideways.
  6. The chart that proves it worked. Cost per request over time — and how to keep it honest as traffic grows.

Coming soon

In outline.

Outline

Enter to go · Esc to close