Open cardosec

Case 8B766F · AI security · L4 Advanced

Threat Modeling AI Agents

Practise as: Deep dive
  1. 01 Context
  2. 02 Mechanism
  3. 03 Lessons
What a strong answer covers

Try it out loud first. Then check yourself:

  1. Map trust boundaries: user, model, tools, data sources and outbound channels; mark every untrusted content input.
  2. Flag the dangerous combo: access to private data + exposure to untrusted content + ability to communicate out.
  3. Use MITRE ATLAS for adversary techniques and OWASP LLM Top 10 for app-level weaknesses as checklists.
  4. Controls by layer: scoped tool permissions, sandboxed code exec, egress allowlists, approvals, full action logging.
  5. Assume the model will be manipulated; design so a fully hijacked model still cannot exceed the user's authority.

If the interviewer pushes back

  • How does threat modeling change when agents call other agents (multi-agent or MCP-style tool ecosystems)?
  • What telemetry would you require to reconstruct an agent incident end to end?

Go deeper

cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.