aiengineering.guideaiengineering.guide

2 · Talking to Models › Phase 3 › Lesson 2 of 4

Streaming responses, honestly

What this lesson covers

Streaming is not free. It cuts perceived latency by showing the first token faster; it moves complexity into your error handling and your rendering path. When the tradeoff is worth it depends on what you’re building.

See also the companion note.

The outline

  1. What streaming actually is. Server-sent events, chunk shape, the difference between “streaming” and “faster”.
  2. Consuming a stream in Python. The generator loop, why you have to accumulate deltas, what the finish event contains.
  3. The mid-stream error class. A stream that dies at token 200 of 500 — what your client shows, what you retry, what you tell the user.
  4. Streaming with structured output. Why streaming JSON is a trap, and what to do if you need both.
  5. Cost accounting. How usage is reported in streams and why your invoicing code needs to know.
  6. When not to stream. Batch jobs, tool-using loops, systems where the client isn’t a human.

Coming soon

In outline.

Outline

Enter to go · Esc to close