aiengineering.guideaiengineering.guide

Interview cram · 2 weeks out · Stop 1 of 4 lessons

What is an LLM, really?

The problem

Most engineers treat a language model as an oracle: text in, answer out. Then it invents a function that doesn’t exist, costs triple what you budgeted, and returns a different answer to the same prompt twice in a row — and you have no mental model to explain any of it. You don’t need the maths. You need the mechanism.

The intuition

Strip away the chat interface and an LLM is a single, boring function: a sequence of tokens goes in, and a probability distribution over the next token comes out. The model reads everything so far and answers one question — “what token is likely to come next?” — as a score for every token it knows.1 It does not plan a sentence, look anything up, or decide to be truthful.

An “answer” is that function called in a loop: pick a next token, glue it onto the end of the input, run the function again, repeat until a stop signal. This is what autoregressive means, and it’s why the GPT and Claude families generate one token at a time.1 Chat, “reasoning,” and “memory” are all this one loop with text arranged around it — nothing more. Hold onto that picture; almost every surprising behaviour later in the roadmap falls out of it.

The mechanism

Four moving parts turn “predict the next token” into the thing you type at.

Tokens, not words. The model never sees characters or words. A tokenizer first chops your text into tokens — common chunks of bytes, usually via byte-pair encoding, so frequent sequences become one token and rare ones split into several.2 A rough rule of thumb for English is one token ≈ 4 characters ≈ three-quarters of a word, so ~100 tokens is about 75 words.3 Tokens are case- and space-sensitive: GPT and GPT (with a leading space) are different tokens, because most tokenizers fold the leading space into the token.2 The same text tokenizes to different counts on different models — that, plus different per-token prices, is why the same prompt costs different amounts.3

One forward pass → a distribution. The token sequence runs through the network — a transformer, the architecture from Attention Is All You Need4 — which emits one raw score, a logit, for every token in the vocabulary. A softmax turns those scores into a probability distribution over the whole vocabulary: the model’s answer to “what’s next?”1

Sampling — turning the distribution into one token. How you pick is the part you control.5 Greedy takes the single highest-probability token (most APIs treat temperature=0 as this). Temperature scales the logits before the softmax: T = 1 leaves the distribution unchanged, T < 1 sharpens it toward the top candidates, T > 1 flattens it and gives long-shots a chance — it never changes the ranking, only how confident the model acts.5 Top-k keeps the k highest tokens; top-p (nucleus) keeps the smallest set whose probabilities sum to p, so the pool grows when the model is unsure and shrinks when it’s confident — introduced by Holtzman et al.6

The context window, and no memory. Between two API calls the model remembers nothing — it is stateless; each request stands alone.7 The illusion of a conversation exists because the client re-sends the whole transcript every turn. The context window is the most tokens (prompt + output) the model can attend to at once — from ~2K in early GPT-style models to 128K in Llama 3.1 and beyond in current systems.8 The “KV cache” you’ll read about is a speed-up within one call, not memory across calls.7

Build it

You don’t need a GPU to feel the mechanism — the sampling step is a few lines of pure Python. Given the model’s raw logits, this is the exact transformation the runtime does before it picks a token. The one thing that changes between “deterministic” and “creative” output is a single argument, highlighted.

Show

Build it · stdlib

import math, random

def sample_next_token(logits, temperature=1.0, top_p=1.0):
  if temperature <= 0:            # greedy: argmax
      return max(range(len(logits)),
                 key=lambda i: logits[i])
  scaled = [z / temperature for z in logits]
  m = max(scaled)                 # softmax, max-shifted
  exps = [math.exp(z - m) for z in scaled]
  total = sum(exps)
  probs = [e / total for e in exps]
  order = sorted(range(len(probs)),
                 key=lambda i: probs[i],
                 reverse=True)
  kept, cum = [], 0.0
  for i in order:                 # nucleus (top-p)
      kept.append(i)
      cum += probs[i]
      if cum >= top_p:
          break
  z = sum(probs[i] for i in kept) # renormalise
  r, acc = random.random() * z, 0.0
  for i in kept:
      acc += probs[i]
      if acc >= r:
          return i
  return kept[-1]

temperature → top-p → sample

Use it · the model runtime

tok = sample_next_token(logits, temperature=0.7)

you pass the knobs

Differs in: temperature 0 (greedy) vs > 0 (sampled from the nucleus)

The whole “answer” is this called in a loop — forward → sample → append — until the model emits its stop token or you hit max_tokens. That loop is the shape of every generation endpoint you’ll ever call.

Use it

Run the sampler on one fixed distribution at two temperatures and watch the mechanism: temperature=0 returns the same token every time; a higher temperature spreads the picks across the plausible candidates.

python
import math, random

def sample(logits, temperature):
  if temperature <= 0:
      return max(range(len(logits)), key=lambda i: logits[i])
  e = [math.exp(z / temperature) for z in logits]
  s = sum(e); p = [x / s for x in e]
  r, acc = random.random(), 0.0
  for i, pi in enumerate(p):
      acc += pi
      if acc >= r: return i
  return len(p) - 1

logits = [2.0, 1.0, 0.2, -1.0]          # 4 candidate tokens
print("greedy :", [sample(logits, 0) for _ in range(8)])
print("t=1.0  :", [sample(logits, 1.0) for _ in range(8)])

In real use you never touch logits — you turn four knobs, in rough order of leverage: model choice (the ceiling on everything), the prompt/context (the model’s entire world for that call), sampling (temperature + top_p + max_tokens), and context management (what you keep, summarise, or drop as the window fills). Reach for temperature≈0 on extraction, classification, and structured output; raise it toward 0.7–1.0 for drafting and brainstorming. Move one of temperature or top-p, not both — they interact.5

In production

“Deterministic” is a lie you’ll tell yourself. Even at temperature=0, identical prompts can return different outputs. Sampling is only one source of variation; floating-point non-associativity, GPU batching, and mixture-of-experts routing mean the same input isn’t guaranteed the same logits from one run to the next. Design evals and caches to tolerate near-duplicates — compare on meaning or a tolerance, never demand a byte-identical reply, or your regression suite will flake on the model’s nature rather than on real bugs.

Cost is a token bill, and output dominates. You pay per input token and per output token, and providers price output tokens higher than input. With the rule of thumb of ~4 characters per token,3 a 2,000-word document is roughly 2,600 tokens of input you pay for on every call that includes it — so a long system prompt or a big retrieved context is a recurring cost, not a one-off. Because generation is one token at a time, latency also scales with output length: a 20-token answer streams far faster than a 2,000-token one. Cap max_tokens deliberately, and don’t ask for a paragraph when you need a word. The one input-cost lever worth knowing early is prompt caching: providers can bill a repeated, stable prefix — a big system prompt or a fixed context — at a discount on later calls, so order your prompt with the unchanging part first and the request last.

The window is a hard budget. Prompt plus output must fit inside the context window — 128K tokens on a current mid-size model, but still finite.8 Blow past it and you get a truncation, an error, or silently dropped history, none of them fun to debug at 2 a.m. Track your token count the way you track memory in a tight loop: measure it, cap it, and decide on purpose what gets evicted when you approach the ceiling — the oldest turns, or a running summary of them.

Knowledge is frozen at training time. A model’s parametric knowledge stops at its training cutoff and does not update as the world changes.9 Anything fresh, private, or after the cutoff has to arrive in the context on the call — which is the entire reason retrieval-augmented generation exists, and the whole of the RAG track downstream. If a fact matters and it’s newer than the model, you must hand it to the model; hoping it “knows” is how stale answers ship.

Failure modes

How the mechanism bites in practice — and the first place you notice.
SymptomTriggerScaleDetection
Confident hallucinationThin/absent training signal; answer past the cutoff — the model predicts a plausible token, not a true one9AnyGrounded eval; cite-or-refuse checks
Retrieved fact ignoredContext contradicts parametric memory; the model sides with what it “knows”10RAG, fresh dataAnswer vs. source diff
Character-level errorsCounting letters, reversing strings, digit arithmetic — a word/number is one opaque token2AnyUnit tests on exact-string tasks
Context overflow / lost-in-the-middleTranscript exceeds the window; mid-context facts down-weightedLong sessionsToken-budget alarm; position probes
Same prompt, different answerNon-zero temperature or run-to-run non-determinismAny sampled callSemantic (not exact-match) evals

Keep this

An LLM is one function — tokens in, a distribution over the next token out — called in a loop until it stops. You control four things: which model, what’s in the context, how you sample, and how you manage the window. It is stateless between calls, its knowledge is frozen at training time, and it optimises for plausible, not true.

Reason from that and the rest of this roadmap is mostly detail.

Why does the same prompt sometimes return a different answer?

Generation samples from a probability distribution, so any temperature above 0 can pick a different token — and even at temperature 0, batching, floating-point, and expert routing make outputs non-deterministic. Test on meaning, not exact strings.

A model gives a wrong date for a recent event. Why, mechanically?

Its parametric knowledge is frozen at the training cutoff and it predicts a plausible token rather than verifying truth, so it fills the gap fluently. Fresh facts have to be supplied in the context (retrieval), not assumed.

Why does ' GPT' cost a different number of tokens than 'GPT'?

Tokenizers fold the leading space into the token, so GPT and GPT are different tokens, and token counts (and therefore price) shift with spacing, casing, and the specific model’s tokenizer.

Footnotes

  1. Sebastian Raschka — Understanding next-token prediction. ↩ ↩2 ↩3

  2. Hugging Face — Summary of the tokenizers (byte-pair encoding, leading-space handling). See also Andrej Karpathy — Let’s build the GPT Tokenizer. ↩ ↩2 ↩3

  3. OpenAI — What are tokens and how to count them: ~4 characters or ¾ word per token for English. ↩ ↩2 ↩3

  4. Vaswani et al., 2017 — Attention Is All You Need. Illustrated walkthrough: Jay Alammar — The Illustrated Transformer. ↩

  5. Sebastian Raschka — Temperature, top-k, and top-p sampling. ↩ ↩2 ↩3

  6. Holtzman et al., 2019 — The Curious Case of Neural Text Degeneration (nucleus / top-p sampling). ↩

  7. Hivenet — KV cache, LLM context, and GPU memory (stateless requests; KV cache is within-call). ↩ ↩2

  8. A Survey on KV Cache Management for LLM Acceleration (context-length growth: ~2K early GPT → 128K Llama 3.1 and beyond). ↩ ↩2

  9. Lilian Weng, 2024 — Extrinsic Hallucinations in LLMs (parametric knowledge frozen at cutoff; models predict plausible tokens). ↩ ↩2

  10. Xu et al., 2024 — When Context Leads but Parametric Memory Follows (knowledge conflict). ↩

Verified · Sept 2026

Enter to go · Esc to close