Interview cram · 2 weeks out · Stop 4 of 4 lessons
Design an LLM gateway
The problem
Every service that calls a model directly reinvents rate limiting, retries, and timeouts — inconsistently. The first traffic spike or provider outage then takes all of them down at once, because nothing coordinates the blast.
The intuition
A gateway is one door. Every model request passes through it, so the policies that protect a shared, metered, occasionally-failing backend live in exactly one place instead of being copied into every service and drifting apart.
The mechanism
The gateway keeps a little state per tenant and decides admit or reject before any upstream call happens. The classic admission policy is the token bucket: a counter that refills at a fixed rate up to a cap, where each admitted request spends one token and an empty bucket rejects.
The bucket has two knobs. The refill rate is the sustained requests-per-second you are willing to serve a tenant; the cap is how large a burst you tolerate before the sustained rate takes over. Because admission is a function of elapsed time and the current token count only, the limiter needs no request log — it is memoryless past the last refill, which is what makes it cheap enough to run on every request.
Everything else the gateway does hangs off that admit decision: a retry budget so a failing dependency is not hammered, a circuit breaker so a dead provider is skipped during its cooldown, and a fallback path — a smaller model or a cached answer — for requests the primary can no longer serve.
Build it
The limiter is small enough to read in one sitting and to put a print statement anywhere. Stdlib only, no framework.
Build it · stdlib
class TokenBucket:
def __init__(self, rate, cap):
self.rate, self.cap = rate, cap
self.tokens, self.t = cap, 0.0
def allow(self, now, cost=1):
self.tokens += (now - self.t) * self.rate
self.tokens = min(self.cap, self.tokens)
self.t = now
if self.tokens < cost:
return False
self.tokens -= cost
return True11 lines · in-process
Use it · a shared store
allowed = limiter.hit("tenant-42", cost=1)1 call · shared state
Differs in: whether the count is shared across replicas
Use it
The from-scratch bucket is correct for one process. In production the same policy moves behind a shared store so every replica counts against one tenant budget — the one property that differs, highlighted above.
bucket = TokenBucket(rate=5, cap=5)
print(bucket.allow(now=0.0)) # True
print(bucket.allow(now=0.0)) # TrueIn production
The moment there is more than one gateway replica, the in-process bucket is wrong: a tenant simply spreads load across replicas and exceeds its global budget. Production gateways push the counter into a shared store. Envoy’s rate-limit service keeps token buckets in Redis and is consulted per request; Kong runs the same pattern as a plugin; Cloudflare’s AI Gateway and cloud API gateways expose it as managed configuration. The trade is one network hop on the hot path against a globally correct count, and that hop is itself on the latency budget, so the limiter call is given a tight timeout and fails open or closed deliberately rather than by accident.
Retries are the next trap. A naive “retry three times” turns one provider hiccup into a three-times-larger stampede exactly when the provider is least able to absorb it. Real systems put retries on a budget — a small fraction of total requests may be retries — and pair them with a circuit breaker that stops calling a provider that has crossed a failure threshold, giving it a cooldown before a single trial request probes whether it recovered. When the primary is down or over budget, the gateway serves a fallback: a cheaper model, a cached response, or a queued job with backpressure signalled to the caller, so load sheds in a way the caller can see.
Cost is the other budget the gateway owns. Because every request already passes through it, the gateway is the natural place to meter tokens per tenant, attribute spend, and enforce a hard concurrency limit so one tenant’s batch job cannot starve interactive traffic. Without that chokepoint, spend is discovered on the monthly invoice rather than controlled in the request path, and there is no single number to alert on when a bug starts looping requests.
Every mutating request carries an idempotency key so a client retry after a timeout is de-duplicated to one effect instead of double-charging a tenant or double-writing a result. And the number that actually matters is tail latency, not the mean: a gateway whose p50 is 40 ms but whose p99 is 2 s under load is failing the requests that decide whether the product feels reliable. Anthropic’s latency guidance is blunt that p99 is the SLO users feel, not the average.1 The gateway is where you measure and defend that percentile, because it is the one place every request already passes through.
Failure modes
| Symptom | Trigger | Scale | Detection |
|---|---|---|---|
| 429 storms | Clients retry-all on error | Any | Error-rate + retry-ratio alert |
| Budget overrun | Per-replica bucket, no shared store | 2+ replicas | Per-tenant admit audit |
| Retry stampede | Unbounded retries on a slow provider | Provider degraded | Retry-budget dashboard |
| Silent degradation | Fallback with no signal | Primary outage | Fallback-rate alert |
| p99 blowup | Limiter hop on the hot path | Peak load | Tail-latency histogram |
Keep this
The token bucket is the smallest correct limiter, and the gateway is the one place a shared count, a retry budget, a breaker, and a fallback can be enforced once for every service.
Why does the token bucket need no request history?
Admission depends only on elapsed time and the current token count, so the bucket is memoryless past its last refill — no per-request log is required.
What breaks when each replica keeps its own bucket?
A tenant spreads traffic across replicas and exceeds its global budget, because no replica sees the others’ admissions. The count must live in a shared store.
Why measure p99 rather than the mean latency?
The mean hides the tail. A healthy average can coexist with a p99 that violates the SLO, and the slow requests are the ones that make the product feel unreliable.
Footnotes
-
Percentile latency and SLOs, provider engineering guidance (author’s paraphrase). ↩