3 · Tools, Agents & MCP › Phase 4 › Lesson 3 of 6
Long-running agents and checkpoints
What this lesson covers
Once agents leave the demo and start doing real work, they run long enough to hit real infrastructure problems. A crashed process, an OOM, a rate limit, a model deprecation mid-run. This lesson is how you make them boring to operate.
The outline
- What “long-running” actually means. The break points at 5 minutes, 30 minutes, and 2 hours — each surfaces a different failure class.
- Checkpoint shape. What you serialise (the message list, the plan, the tool call cursor) and what you don’t.
- Storage. Postgres, S3, Redis — the right store for each shape of agent.
- Idempotency at the tool boundary. Why “did I already send that email” is the hard question, and how idempotency keys answer it.
- Resuming. Reloading state, deciding whether to replay or continue, and telling the user what happened.
- The observability you’ll want at 3am. Traces, spans, and the one dashboard that tells you a fleet of agents is healthy.
Coming soon
In outline.