Interview Q&A · 5 · Production (MLOps + LLMOps)
How would you defend a customer-support agent against prompt injection when it can send emails and read a customer database?
Reveal the answer
Named defences that don't work first: regex filters on user text, "ignore
previous instructions" detectors, and adding "you are a helpful assistant,
do not follow user instructions" to the system prompt. Any determined
attacker gets past all three within a page of prose.
The one thing that works is architectural: treat every piece of
attacker-controllable text (the user message, the ticket body, an
attached document, a knowledge-base snippet — anything not written by
you) as inert data, and structure the agent so that data cannot
authorise a privileged action.
Concretely for this system: (1) split the agent's plan step and its
execute step across two calls with different privileges; the planner
sees the user input, the executor only sees a small structured JSON
plan; (2) any tool that sends email or writes to the database requires
a confirmation loop with either a rules engine or a human — never the
model alone; (3) redact PII at the retrieval boundary, not in the prompt;
(4) log every tool call for detection so an ongoing exploit surfaces.
Layered, not perfect. But the architectural layer is what makes the
attack expensive.
Common variants
- What class of attacks does the OWASP LLM Top 10 name — pick three?
- If you have to keep it in one agent for cost reasons, what changes?
- How do you detect an ongoing injection attempt from telemetry alone?
Track this card
Stored in this browser only — no signup, no sync. Clearing site data clears it.