5 · Production (MLOps + LLMOps) › Phase 6 › Lesson 4 of 5
Caching and cost control
What this lesson covers
The gap between a bill you can defend and one you can’t is usually caching, done right. This lesson names the three kinds, when each applies, and the ways they quietly fail.
The outline
- Prompt caching, provider-side. How Anthropic and OpenAI’s cache mechanisms work, the cache-key rules that break naive setups, and the discount you actually get.
- The cache-friendly prompt shape. Static content first, dynamic content last — and how to refactor a prompt that gets this wrong.
- Response caching. When identical requests deserve a cached answer, and how to key the cache without leaking one user’s answer to another.
- Semantic caching. Answering “close-enough” requests from cache. The single most misused optimisation in AI systems — when it works and how it goes wrong.
- Cost budgets and alerts. Per-tenant, per-feature, per-day. The alerts that catch a bill going sideways.
- The chart that proves it worked. Cost per request over time — and how to keep it honest as traffic grows.
Coming soon
In outline.