aiengineering.guideaiengineering.guide

Field Note

Streaming responses aren't free

Published

Streaming a model’s response token by token is the first thing everyone reaches for when the demo feels slow. The perceived latency drops because the reader sees the first word in a few hundred milliseconds instead of waiting for the whole completion. That part is real. What is less obvious is that streaming does not make the request faster — it makes the first byte faster and moves a surprising amount of complexity downstream.

Start with error handling. In a non-streaming call, a failure is a single event: the request either returns a completion or an error, and your retry logic wraps one boundary. Once you stream, the request can fail after you have already sent usable tokens to the client. Now you have a half-written answer on the screen and an error in your hand. Do you retry from scratch and discard what the reader already saw? Do you show the partial answer and a quiet “generation interrupted” notice? Do you resume — and if so, from which token, given the model has no memory of the bytes you already streamed? None of these are wrong, but all of them are decisions you did not have to make before, and the default framework behavior usually picks the worst one for you: it throws away the partial and shows a spinner that never resolves.

Then there is the retry budget. A streamed request holds a connection open for the full generation. If you put streaming behind the same aggressive retry policy you use for short JSON calls, a slow generation that trips your timeout gets retried — and now you are paying for the tokens twice and holding two connections open for one answer. Streaming and retries interact badly unless the timeout is measured against time-to-first-token, not total duration.

The client pays too. A streamed response is a sequence of small writes, and every layer between the model and the reader has to forward them promptly or the benefit evaporates. A buffering proxy, a CDN that waits for a complete response, or a client that re-renders the whole message on every token will quietly erase the latency win and add jank. The number that matters is not throughput; it is whether each token reaches the DOM without a layout thrash.

None of this argues against streaming. For a chat surface it is almost always the right call, and readers genuinely feel the difference. The point is that “turn on streaming” is not a latency optimization you get for free — it is a trade that buys a better first impression with more careful error handling, a different timeout model, and a forwarding path that does not buffer. Budget for those three when you turn it on, and streaming is a clear win. Skip them, and you have shipped a faster way to show a broken answer.

Enter to go · Esc to close