2 · Talking to Models › Phase 3 › Lesson 4 of 4
Tokens, context, and what a call actually costs
What this lesson covers
Cost is the variable that decides whether an AI product is a business. This lesson makes the cost model concrete, in the shape you’ll actually reason about.
The outline
- What a token is. The tokenizer, why “GPT” and ” GPT” cost different amounts, and how to count them from Python.
- Input vs. output tokens. Why they’re priced differently and why output almost always dominates the bill.
- The context window as a budget. Why treating the window as free memory is the single most expensive mistake beginners make.
- Predicting your bill. Requests × average call size × unit price — with worked examples for a chatbot, a summariser, and an agent loop.
- Prompt patterns that waste tokens. Restating the system prompt, verbose few-shot, unbounded histories, and how to fix each.
- Caching as a cost lever. A preview of prompt caching, covered in depth in Production.
Coming soon
In outline.