Open cardosec

Case 1DC165 · AI security · L5 Expert

Injection-Resistant Agent Architectures

Practise as: Deep dive · Interview

Interview questionPrompt injection detection is probabilistic. How would you architect an agent so injected content cannot cause harmful tool calls?

  1. 01 Context
  2. 02 Mechanism
  3. 03 Lessons
What a strong answer covers

Try it out loud first. Then check yourself:

  1. Dual LLM pattern: a privileged LLM plans and calls tools but never sees untrusted text; a quarantined LLM reads it with no tools.
  2. Untrusted results are passed by reference as variables, so their contents can never instruct the privileged model.
  3. CaMeL (Google DeepMind, 2025) turns the trusted query into code; its interpreter tracks data provenance and checks tool policies.
  4. Plan-then-execute fixes the control flow up front, but injected data can still change argument values like recipient or amount.
  5. Cost: lower utility on tasks needing reasoning over untrusted text, and per-tool policies to write. It complements authz.

If the interviewer pushes back

  • Which attacks still work against a CaMeL-style design, e.g. a manipulated summary shown to a human who then acts?
  • How would you apply this pattern to an email agent that has to reply to incoming messages?

cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.