aiengineering.guideaiengineering.guide

ROADMAP · ~24 hours over 4 weekends

DS/MLE → LLM engineer

Data scientists or ML engineers with production experience on classical ML, new to LLMs.

Start: What is an LLM, really?13 stops · ~24 hours over 4 weekends

Who this is for

You have 2+ years shipping ML models to production. You know train/eval/deploy, you’ve dealt with drift and mislabeled data, and you can explain why a random forest is a bad idea in half the cases you’ve been given.

You have not built an LLM-based system that a team relies on.

What’s actually different

Three shifts you’ll notice:

  • You don’t own the model. No hyperparameter tuning, no feature engineering. Your competitive lever is the system around the model: prompts, retrieval, eval, cost control.
  • Non-determinism is a feature and a bug. Same input, different output is the default. Your test strategy changes.
  • Cost dominates. Classical ML cost is inference compute. LLM cost is tokens × requests × redo rate. Cost engineering is now an engineering discipline.

What accelerates fast for you

  • Eval instinct. If you’ve built a well-designed eval loop for a classical model, you’re 3 months ahead of most people learning LLMs.
  • Production instinct. You know why serving is different from training. Half the LLM production war stories will feel familiar.
  • Cost math. You’ve calculated inference cost per prediction. Now you’ll do it per token.

THE PATH

13 stops, in order

Phase 1What's different4 stops

  1. 1

    Lesson

    What is an LLM, really?

    Read once for framing. Focus on how LLMs differ from your models, not on how they're similar.

  2. 2

    Lesson

    Tokens, context, and what a call actually costs

    Cost accounting is different. Batch inference doesn't save you money the same way.

  3. 3

    Lesson

    Reading a model card and a pricing sheet

    The equivalent of reading a paper's results table for someone else's frozen model.

  4. 4

    Lesson

    Your first real API call

    Skim if you've done this. The rest of the track assumes you can.

Phase 2LLM-native patterns3 stops

  1. 5

    Lesson

    Structured output you can trust

    JSON mode + schema-constrained decoding is the pattern for anything you'd have used a downstream classifier for.

  2. 6

    Lesson

    Streaming responses, honestly

    Latency budget is different. Streaming is the pattern; understand its trade-offs.

  3. 7

    Lesson

    Tool use, from three lines of Python

    Same shape as classical retrieval-augmented systems but LLM-native.

Phase 3Retrieval at scale4 stops

  1. 8

    Lesson

    Embeddings, without the maths

    You've used embeddings. Focus on what changes at 1B+ vectors.

  2. 9

    Lesson

    Vector stores worth using in 2026

    Choice depends on scale + query pattern. Same rules as classical vector search you've done.

  3. 10

    Lesson

    Evaluating a RAG system without lying to yourself

    The hardest transferable skill from classical ML — you have to build the eval or you don't have a system.

  4. 11

    Interview card

    What's the difference between an eval and a benchmark, and why would a team run both?

    Sharp distinction most LLM tutorials get wrong. You already know it from classical ML.

Phase 4Production2 stops

  1. 12

    Lesson

    Design an LLM gateway

    The FLAGSHIP. Production LLM systems need this layer. Classical ML rarely does.

  2. 13

    Case file

    The eval pipeline that lied for a month

    Failure mode you've probably lived through with classical models but the LLM version is subtler.

NOT COVERED HERE

What this roadmap skips

  • Foundations. You have them.
  • Careful math. If you want the math you already know where to find it.
  • "Just add RAG" tutorials. This roadmap assumes you can reason about retrieval systems.

← All roadmaps

Enter to go · Esc to close