aiengineering.guideaiengineering.guide

Interview Q&A · 5 · Production (MLOps + LLMOps)

How would you defend a customer-support agent against prompt injection when it can send emails and read a customer database?

hardsecurityagentsproductionasked at AnthropicScale AIOpenAIGoogle DeepMind· 2026source: Production · Prompt injection is a real vulnerability

Reveal the answer
Named defences that don't work first: regex filters on user text, "ignore previous instructions" detectors, and adding "you are a helpful assistant, do not follow user instructions" to the system prompt. Any determined attacker gets past all three within a page of prose. The one thing that works is architectural: treat every piece of attacker-controllable text (the user message, the ticket body, an attached document, a knowledge-base snippet — anything not written by you) as inert data, and structure the agent so that data cannot authorise a privileged action. Concretely for this system: (1) split the agent's plan step and its execute step across two calls with different privileges; the planner sees the user input, the executor only sees a small structured JSON plan; (2) any tool that sends email or writes to the database requires a confirmation loop with either a rules engine or a human — never the model alone; (3) redact PII at the retrieval boundary, not in the prompt; (4) log every tool call for detection so an ongoing exploit surfaces. Layered, not perfect. But the architectural layer is what makes the attack expensive.

Common variants

  • What class of attacks does the OWASP LLM Top 10 name — pick three?
  • If you have to keep it in one agent for cost reasons, what changes?
  • How do you detect an ongoing injection attempt from telemetry alone?

Track this card

Stored in this browser only — no signup, no sync. Clearing site data clears it.

Verified · Sept 2026

Enter to go · Esc to close