aiengineering.guideaiengineering.guide

3 · Tools, Agents & MCP › Phase 4 › Lesson 3 of 6

Long-running agents and checkpoints

What this lesson covers

Once agents leave the demo and start doing real work, they run long enough to hit real infrastructure problems. A crashed process, an OOM, a rate limit, a model deprecation mid-run. This lesson is how you make them boring to operate.

The outline

  1. What “long-running” actually means. The break points at 5 minutes, 30 minutes, and 2 hours — each surfaces a different failure class.
  2. Checkpoint shape. What you serialise (the message list, the plan, the tool call cursor) and what you don’t.
  3. Storage. Postgres, S3, Redis — the right store for each shape of agent.
  4. Idempotency at the tool boundary. Why “did I already send that email” is the hard question, and how idempotency keys answer it.
  5. Resuming. Reloading state, deciding whether to replay or continue, and telling the user what happened.
  6. The observability you’ll want at 3am. Traces, spans, and the one dashboard that tells you a fleet of agents is healthy.

Coming soon

In outline.

Outline

Enter to go · Esc to close