aiengineering.guideaiengineering.guide

4 · Retrieval & Long Context › Phase 5 › Lesson 2 of 4

Chunking is the decision

What this lesson covers

Nobody writes a blog post about chunking. Everyone with a production RAG system has spent a week on it. This lesson names the strategies and the tradeoffs so you spend that week deliberately.

The outline

  1. Why chunking exists. The context-window and embedding-quality pressures that make “one document, one embedding” a bad idea for almost any real corpus.
  2. Fixed-size chunking. The baseline. Why 512 tokens with 64 of overlap is a defensible default and when it isn’t.
  3. Structural chunking. Splitting on headings, paragraphs, sections — when the document has structure worth using.
  4. Semantic chunking. Splitting on meaning boundaries with a smaller model. What it buys you and what it costs.
  5. Contextual chunking (2024–2026). Prefixing each chunk with a generated summary of its position — why this quietly became the state of the art.
  6. The eval you have to run. How to measure whether your chunking change actually helped, instead of guessing.

Coming soon

In outline.

Outline

Enter to go · Esc to close