Case 1DC165 · AI security · L5 Expert
Injection-Resistant Agent Architectures
Practise as: Deep dive · Interview
Interview questionPrompt injection detection is probabilistic. How would you architect an agent so injected content cannot cause harmful tool calls?
- 01 Context
- 02 Mechanism
- 03 Lessons
What a strong answer covers
Try it out loud first. Then check yourself:
- Dual LLM pattern: a privileged LLM plans and calls tools but never sees untrusted text; a quarantined LLM reads it with no tools.
- Untrusted results are passed by reference as variables, so their contents can never instruct the privileged model.
- CaMeL (Google DeepMind, 2025) turns the trusted query into code; its interpreter tracks data provenance and checks tool policies.
- Plan-then-execute fixes the control flow up front, but injected data can still change argument values like recipient or amount.
- Cost: lower utility on tasks needing reasoning over untrusted text, and per-tool policies to write. It complements authz.
If the interviewer pushes back
- Which attacks still work against a CaMeL-style design, e.g. a manipulated summary shown to a human who then acts?
- How would you apply this pattern to an email agent that has to reply to incoming messages?
cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.