Interview Q&A · 2 · Talking to Models
Temperature and top-p both control randomness. What's the difference, and when would you reach for each?
Reveal the answer
Both reshape the next-token probability distribution before sampling.
Temperature is a global spread control: T=0 makes decoding deterministic
(always the argmax), T>1 flattens the distribution so unlikely tokens get
picked more often. Top-p (nucleus sampling) caps the sampling pool: it
keeps only the smallest set of tokens whose cumulative probability sums to
p, then samples from that.
In practice: use low T (0 or 0.2) when you need reliability — extraction,
classification, structured output. Use moderate T (0.7) with top-p ~0.9
for open-ended generation — writing, brainstorming. Setting both aggressively
compounds the effect and usually hurts more than it helps; pick one lever.
Common variants
- What's the effect of top-k, and why do most APIs default to top-p?
- How would you get deterministic behaviour from an API that samples?
- Why can two calls with T=0 still return different answers?
Track this card
Stored in this browser only — no signup, no sync. Clearing site data clears it.