Working SWE → AI engineer · Stop 12 of 15 lessons
Chunking is the decision
What this lesson covers
Nobody writes a blog post about chunking. Everyone with a production RAG system has spent a week on it. This lesson names the strategies and the tradeoffs so you spend that week deliberately.
The outline
- Why chunking exists. The context-window and embedding-quality pressures that make “one document, one embedding” a bad idea for almost any real corpus.
- Fixed-size chunking. The baseline. Why 512 tokens with 64 of overlap is a defensible default and when it isn’t.
- Structural chunking. Splitting on headings, paragraphs, sections — when the document has structure worth using.
- Semantic chunking. Splitting on meaning boundaries with a smaller model. What it buys you and what it costs.
- Contextual chunking (2024–2026). Prefixing each chunk with a generated summary of its position — why this quietly became the state of the art.
- The eval you have to run. How to measure whether your chunking change actually helped, instead of guessing.
Coming soon
In outline.