5 · Production (MLOps + LLMOps) › Phase 6 › Lesson 2 of 5
Prompt injection is a real vulnerability
What this lesson covers
Prompt injection is treated as a research curiosity by teams that haven’t been hit yet and as a Sev-1 by teams that have. Which side you’re on is mostly a question of when your product got popular enough to attract attention.
The outline
- The threat model. Attacker-controlled text reaches the model’s context. What they can do — data exfiltration, tool abuse, output manipulation.
- Defences that don’t work. Regex filters, “ignore previous instructions” detectors, adding “you are a helpful assistant” to the system prompt.
- The one thing that does work. Architectural separation — treat any untrusted text as inert data and design so the model can’t act on it with privilege.
- Tool budgets and confirmation gates. How to make an exfiltration attempt cost the attacker more than it costs you to detect.
- PII and secret handling. Redaction at the boundary and why doing it in the prompt is too late.
- Detection. What to log so you notice injection attempts, before an incident forces you to.
Coming soon
In outline.