Case 8B766F · AI security · L4 Advanced
Threat Modeling AI Agents
Practise as: Deep dive
- 01 Context
- 02 Mechanism
- 03 Lessons
What a strong answer covers
Try it out loud first. Then check yourself:
- Map trust boundaries: user, model, tools, data sources and outbound channels; mark every untrusted content input.
- Flag the dangerous combo: access to private data + exposure to untrusted content + ability to communicate out.
- Use MITRE ATLAS for adversary techniques and OWASP LLM Top 10 for app-level weaknesses as checklists.
- Controls by layer: scoped tool permissions, sandboxed code exec, egress allowlists, approvals, full action logging.
- Assume the model will be manipulated; design so a fully hijacked model still cannot exceed the user's authority.
If the interviewer pushes back
- How does threat modeling change when agents call other agents (multi-agent or MCP-style tool ecosystems)?
- What telemetry would you require to reconstruct an agent incident end to end?
Go deeper
cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.