aiengineering.guideaiengineering.guide

5 · Production (MLOps + LLMOps) › Phase 6 › Lesson 2 of 5

Prompt injection is a real vulnerability

What this lesson covers

Prompt injection is treated as a research curiosity by teams that haven’t been hit yet and as a Sev-1 by teams that have. Which side you’re on is mostly a question of when your product got popular enough to attract attention.

The outline

  1. The threat model. Attacker-controlled text reaches the model’s context. What they can do — data exfiltration, tool abuse, output manipulation.
  2. Defences that don’t work. Regex filters, “ignore previous instructions” detectors, adding “you are a helpful assistant” to the system prompt.
  3. The one thing that does work. Architectural separation — treat any untrusted text as inert data and design so the model can’t act on it with privilege.
  4. Tool budgets and confirmation gates. How to make an exfiltration attempt cost the attacker more than it costs you to detect.
  5. PII and secret handling. Redaction at the boundary and why doing it in the prompt is too late.
  6. Detection. What to log so you notice injection attempts, before an incident forces you to.

Coming soon

In outline.

Outline

Enter to go · Esc to close