ROADMAP · ~24 hours over 4 weekends
DS/MLE → LLM engineer
Data scientists or ML engineers with production experience on classical ML, new to LLMs.
Who this is for
You have 2+ years shipping ML models to production. You know train/eval/deploy, you’ve dealt with drift and mislabeled data, and you can explain why a random forest is a bad idea in half the cases you’ve been given.
You have not built an LLM-based system that a team relies on.
What’s actually different
Three shifts you’ll notice:
- You don’t own the model. No hyperparameter tuning, no feature engineering. Your competitive lever is the system around the model: prompts, retrieval, eval, cost control.
- Non-determinism is a feature and a bug. Same input, different output is the default. Your test strategy changes.
- Cost dominates. Classical ML cost is inference compute. LLM cost is tokens × requests × redo rate. Cost engineering is now an engineering discipline.
What accelerates fast for you
- Eval instinct. If you’ve built a well-designed eval loop for a classical model, you’re 3 months ahead of most people learning LLMs.
- Production instinct. You know why serving is different from training. Half the LLM production war stories will feel familiar.
- Cost math. You’ve calculated inference cost per prediction. Now you’ll do it per token.
THE PATH
13 stops, in order
Phase 1What's different4 stops
- 1
Lesson
What is an LLM, really?Read once for framing. Focus on how LLMs differ from your models, not on how they're similar.
- 2
Lesson
Tokens, context, and what a call actually costsCost accounting is different. Batch inference doesn't save you money the same way.
- 3
Lesson
Reading a model card and a pricing sheetThe equivalent of reading a paper's results table for someone else's frozen model.
- 4
Phase 2LLM-native patterns3 stops
- 5
Lesson
Structured output you can trustJSON mode + schema-constrained decoding is the pattern for anything you'd have used a downstream classifier for.
- 6
Lesson
Streaming responses, honestlyLatency budget is different. Streaming is the pattern; understand its trade-offs.
- 7
Lesson
Tool use, from three lines of PythonSame shape as classical retrieval-augmented systems but LLM-native.
Phase 3Retrieval at scale4 stops
- 8
- 9
Lesson
Vector stores worth using in 2026Choice depends on scale + query pattern. Same rules as classical vector search you've done.
- 10
Lesson
Evaluating a RAG system without lying to yourselfThe hardest transferable skill from classical ML — you have to build the eval or you don't have a system.
- 11
Interview card
What's the difference between an eval and a benchmark, and why would a team run both?Sharp distinction most LLM tutorials get wrong. You already know it from classical ML.
Phase 4Production2 stops
- 12
Lesson
Design an LLM gatewayThe FLAGSHIP. Production LLM systems need this layer. Classical ML rarely does.
- 13
Case file
The eval pipeline that lied for a monthFailure mode you've probably lived through with classical models but the LLM version is subtler.
NOT COVERED HERE
What this roadmap skips
- Foundations. You have them.
- Careful math. If you want the math you already know where to find it.
- "Just add RAG" tutorials. This roadmap assumes you can reason about retrieval systems.