aiengineering.guideaiengineering.guide

Interview Q&A · 2 · Talking to Models

Temperature and top-p both control randomness. What's the difference, and when would you reach for each?

mediumsamplingapisdecodingasked at OpenAICoherePerplexity· 2026source: Foundations · What is an LLM

Reveal the answer
Both reshape the next-token probability distribution before sampling. Temperature is a global spread control: T=0 makes decoding deterministic (always the argmax), T>1 flattens the distribution so unlikely tokens get picked more often. Top-p (nucleus sampling) caps the sampling pool: it keeps only the smallest set of tokens whose cumulative probability sums to p, then samples from that. In practice: use low T (0 or 0.2) when you need reliability — extraction, classification, structured output. Use moderate T (0.7) with top-p ~0.9 for open-ended generation — writing, brainstorming. Setting both aggressively compounds the effect and usually hurts more than it helps; pick one lever.

Common variants

  • What's the effect of top-k, and why do most APIs default to top-p?
  • How would you get deterministic behaviour from an API that samples?
  • Why can two calls with T=0 still return different answers?

Track this card

Stored in this browser only — no signup, no sync. Clearing site data clears it.

Verified · Sept 2026

Enter to go · Esc to close