Open cardosec

Case D31F3E · AI security · L3 Applied

Indirect Prompt Injection

Practise as: Explain it · Interview · Deep dive

Interview questionExplain indirect prompt injection and how you would defend an LLM assistant that reads email and browses the web.

  1. 01 What is it?
  2. 02 How is it abused?
  3. 03 How do you stop it?
What a strong answer covers

Try it out loud first. Then check yourself:

  1. Malicious instructions arrive via content the model ingests (web pages, emails, PDFs, RAG docs), not from the user.
  2. Payloads hide in white text, HTML comments, alt text, metadata or Unicode tag characters invisible to humans.
  3. Classic exfil: injected text makes the model render a markdown image URL with stolen data in the query string.
  4. Defenses: tag untrusted content, block auto-rendered external URLs, restrict tools after reading untrusted data.
  5. Highest risk when an agent has private data, untrusted input and an outbound channel at the same time.

If the interviewer pushes back

  • Design a policy that lets an agent summarize untrusted web pages but never exfiltrate data. What does it lose?
  • How would you detect indirect injection attempts in production logs?

Go deeper

cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.