Working SWE → AI engineer · Stop 1 of 15 lessons
What is an LLM, really?
The problem
Most engineers treat a language model as an oracle: text in, answer out. Then it invents a function that doesn’t exist, costs triple what you budgeted, and returns a different answer to the same prompt twice in a row — and you have no mental model to explain any of it. You don’t need the maths. You need the mechanism.
The intuition
Strip away the chat interface and an LLM is a single, boring function: a sequence of tokens goes in, and a probability distribution over the next token comes out. The model reads everything so far and answers one question — “what token is likely to come next?” — as a score for every token it knows.1 It does not plan a sentence, look anything up, or decide to be truthful.
An “answer” is that function called in a loop: pick a next token, glue it onto the end of the input, run the function again, repeat until a stop signal. This is what autoregressive means, and it’s why the GPT and Claude families generate one token at a time.1 Chat, “reasoning,” and “memory” are all this one loop with text arranged around it — nothing more. Hold onto that picture; almost every surprising behaviour later in the roadmap falls out of it.
The mechanism
Four moving parts turn “predict the next token” into the thing you type at.
Tokens, not words. The model never sees characters or words. A tokenizer
first chops your text into tokens — common chunks of bytes, usually via
byte-pair encoding, so frequent sequences become one token and rare ones split
into several.2 A rough rule of thumb for English is one token ≈ 4
characters ≈ three-quarters of a word, so ~100 tokens is about 75 words.3
Tokens are case- and space-sensitive: GPT and GPT (with a leading space)
are different tokens, because most tokenizers fold the leading space into the
token.2 The same text tokenizes to different counts on different models —
that, plus different per-token prices, is why the same prompt costs different
amounts.3
One forward pass → a distribution. The token sequence runs through the network — a transformer, the architecture from Attention Is All You Need4 — which emits one raw score, a logit, for every token in the vocabulary. A softmax turns those scores into a probability distribution over the whole vocabulary: the model’s answer to “what’s next?”1
Sampling — turning the distribution into one token. How you pick is the part
you control.5 Greedy takes the single highest-probability
token (most APIs treat temperature=0 as this). Temperature scales the logits
before the softmax: T = 1 leaves the distribution unchanged, T < 1 sharpens
it toward the top candidates, T > 1 flattens it and gives long-shots a chance —
it never changes the ranking, only how confident the model acts.5
Top-k keeps the k highest tokens; top-p (nucleus) keeps the smallest set
whose probabilities sum to p, so the pool grows when the model is unsure and
shrinks when it’s confident — introduced by Holtzman et al.6
The context window, and no memory. Between two API calls the model remembers nothing — it is stateless; each request stands alone.7 The illusion of a conversation exists because the client re-sends the whole transcript every turn. The context window is the most tokens (prompt + output) the model can attend to at once — from ~2K in early GPT-style models to 128K in Llama 3.1 and beyond in current systems.8 The “KV cache” you’ll read about is a speed-up within one call, not memory across calls.7
Build it
You don’t need a GPU to feel the mechanism — the sampling step is a few lines of pure Python. Given the model’s raw logits, this is the exact transformation the runtime does before it picks a token. The one thing that changes between “deterministic” and “creative” output is a single argument, highlighted.
Build it · stdlib
import math, random
def sample_next_token(logits, temperature=1.0, top_p=1.0):
if temperature <= 0: # greedy: argmax
return max(range(len(logits)),
key=lambda i: logits[i])
scaled = [z / temperature for z in logits]
m = max(scaled) # softmax, max-shifted
exps = [math.exp(z - m) for z in scaled]
total = sum(exps)
probs = [e / total for e in exps]
order = sorted(range(len(probs)),
key=lambda i: probs[i],
reverse=True)
kept, cum = [], 0.0
for i in order: # nucleus (top-p)
kept.append(i)
cum += probs[i]
if cum >= top_p:
break
z = sum(probs[i] for i in kept) # renormalise
r, acc = random.random() * z, 0.0
for i in kept:
acc += probs[i]
if acc >= r:
return i
return kept[-1]temperature → top-p → sample
Use it · the model runtime
tok = sample_next_token(logits, temperature=0.7)you pass the knobs
Differs in: temperature 0 (greedy) vs > 0 (sampled from the nucleus)
The whole “answer” is this called in a loop — forward → sample → append — until
the model emits its stop token or you hit max_tokens. That loop is the shape of
every generation endpoint you’ll ever call.
Use it
Run the sampler on one fixed distribution at two temperatures and watch the
mechanism: temperature=0 returns the same token every time; a higher
temperature spreads the picks across the plausible candidates.
import math, random
def sample(logits, temperature):
if temperature <= 0:
return max(range(len(logits)), key=lambda i: logits[i])
e = [math.exp(z / temperature) for z in logits]
s = sum(e); p = [x / s for x in e]
r, acc = random.random(), 0.0
for i, pi in enumerate(p):
acc += pi
if acc >= r: return i
return len(p) - 1
logits = [2.0, 1.0, 0.2, -1.0] # 4 candidate tokens
print("greedy :", [sample(logits, 0) for _ in range(8)])
print("t=1.0 :", [sample(logits, 1.0) for _ in range(8)])In real use you never touch logits — you turn four knobs, in rough order of
leverage: model choice (the ceiling on everything), the prompt/context
(the model’s entire world for that call), sampling (temperature + top_p +
max_tokens), and context management (what you keep, summarise, or drop as
the window fills). Reach for temperature≈0 on extraction, classification, and
structured output; raise it toward 0.7–1.0 for drafting and brainstorming. Move
one of temperature or top-p, not both — they interact.5
In production
“Deterministic” is a lie you’ll tell yourself. Even at temperature=0,
identical prompts can return different outputs. Sampling is only one source of
variation; floating-point non-associativity, GPU batching, and mixture-of-experts
routing mean the same input isn’t guaranteed the same logits from one run to the
next. Design evals and caches to tolerate near-duplicates — compare on meaning or
a tolerance, never demand a byte-identical reply, or your regression suite will
flake on the model’s nature rather than on real bugs.
Cost is a token bill, and output dominates. You pay per input token and per
output token, and providers price output tokens higher than input. With the
rule of thumb of ~4 characters per token,3 a 2,000-word document is
roughly 2,600 tokens of input you pay for on every call that includes it — so a
long system prompt or a big retrieved context is a recurring cost, not a one-off.
Because generation is one token at a time, latency also scales with output
length: a 20-token answer streams far faster than a 2,000-token one. Cap
max_tokens deliberately, and don’t ask for a paragraph when you need a word.
The one input-cost lever worth knowing early is prompt caching: providers can
bill a repeated, stable prefix — a big system prompt or a fixed context — at a
discount on later calls, so order your prompt with the unchanging part first and
the request last.
The window is a hard budget. Prompt plus output must fit inside the context window — 128K tokens on a current mid-size model, but still finite.8 Blow past it and you get a truncation, an error, or silently dropped history, none of them fun to debug at 2 a.m. Track your token count the way you track memory in a tight loop: measure it, cap it, and decide on purpose what gets evicted when you approach the ceiling — the oldest turns, or a running summary of them.
Knowledge is frozen at training time. A model’s parametric knowledge stops at its training cutoff and does not update as the world changes.9 Anything fresh, private, or after the cutoff has to arrive in the context on the call — which is the entire reason retrieval-augmented generation exists, and the whole of the RAG track downstream. If a fact matters and it’s newer than the model, you must hand it to the model; hoping it “knows” is how stale answers ship.
Failure modes
| Symptom | Trigger | Scale | Detection |
|---|---|---|---|
| Confident hallucination | Thin/absent training signal; answer past the cutoff — the model predicts a plausible token, not a true one9 | Any | Grounded eval; cite-or-refuse checks |
| Retrieved fact ignored | Context contradicts parametric memory; the model sides with what it “knows”10 | RAG, fresh data | Answer vs. source diff |
| Character-level errors | Counting letters, reversing strings, digit arithmetic — a word/number is one opaque token2 | Any | Unit tests on exact-string tasks |
| Context overflow / lost-in-the-middle | Transcript exceeds the window; mid-context facts down-weighted | Long sessions | Token-budget alarm; position probes |
| Same prompt, different answer | Non-zero temperature or run-to-run non-determinism | Any sampled call | Semantic (not exact-match) evals |
Keep this
An LLM is one function — tokens in, a distribution over the next token out — called in a loop until it stops. You control four things: which model, what’s in the context, how you sample, and how you manage the window. It is stateless between calls, its knowledge is frozen at training time, and it optimises for plausible, not true. Reason from that and the rest of this roadmap is mostly detail.
Why does the same prompt sometimes return a different answer?
Generation samples from a probability distribution, so any temperature above 0 can pick a different token — and even at temperature 0, batching, floating-point, and expert routing make outputs non-deterministic. Test on meaning, not exact strings.
A model gives a wrong date for a recent event. Why, mechanically?
Its parametric knowledge is frozen at the training cutoff and it predicts a plausible token rather than verifying truth, so it fills the gap fluently. Fresh facts have to be supplied in the context (retrieval), not assumed.
Why does ' GPT' cost a different number of tokens than 'GPT'?
Tokenizers fold the leading space into the token, so GPT and GPT are
different tokens, and token counts (and therefore price) shift with spacing,
casing, and the specific model’s tokenizer.
Footnotes
-
Sebastian Raschka — Understanding next-token prediction. ↩ ↩2 ↩3
-
Hugging Face — Summary of the tokenizers (byte-pair encoding, leading-space handling). See also Andrej Karpathy — Let’s build the GPT Tokenizer. ↩ ↩2 ↩3
-
OpenAI — What are tokens and how to count them: ~4 characters or ¾ word per token for English. ↩ ↩2 ↩3
-
Vaswani et al., 2017 — Attention Is All You Need. Illustrated walkthrough: Jay Alammar — The Illustrated Transformer. ↩
-
Sebastian Raschka — Temperature, top-k, and top-p sampling. ↩ ↩2 ↩3
-
Holtzman et al., 2019 — The Curious Case of Neural Text Degeneration (nucleus / top-p sampling). ↩
-
Hivenet — KV cache, LLM context, and GPU memory (stateless requests; KV cache is within-call). ↩ ↩2
-
A Survey on KV Cache Management for LLM Acceleration (context-length growth: ~2K early GPT → 128K Llama 3.1 and beyond). ↩ ↩2
-
Lilian Weng, 2024 — Extrinsic Hallucinations in LLMs (parametric knowledge frozen at cutoff; models predict plausible tokens). ↩ ↩2
-
Xu et al., 2024 — When Context Leads but Parametric Memory Follows (knowledge conflict). ↩