aiengineering.guideaiengineering.guide

Working SWE → AI engineer · Stop 4 of 15 lessons

Tokens, context, and what a call actually costs

What this lesson covers

Cost is the variable that decides whether an AI product is a business. This lesson makes the cost model concrete, in the shape you’ll actually reason about.

The outline

  1. What a token is. The tokenizer, why “GPT” and ” GPT” cost different amounts, and how to count them from Python.
  2. Input vs. output tokens. Why they’re priced differently and why output almost always dominates the bill.
  3. The context window as a budget. Why treating the window as free memory is the single most expensive mistake beginners make.
  4. Predicting your bill. Requests × average call size × unit price — with worked examples for a chatbot, a summariser, and an agent loop.
  5. Prompt patterns that waste tokens. Restating the system prompt, verbose few-shot, unbounded histories, and how to fix each.
  6. Caching as a cost lever. A preview of prompt caching, covered in depth in Production.

Coming soon

In outline.

Outline

Enter to go · Esc to close