Open cardosec

Case 9A84E1 · AI security · L3 Applied

Data and Model Poisoning

Practise as: Explain it · Deep dive
  1. 01 What is it?
  2. 02 How is it abused?
  3. 03 How do you stop it?
What a strong answer covers

Try it out loud first. Then check yourself:

  1. Attacker manipulates training, fine-tuning or RAG data to change model behavior. OWASP LLM04 in the 2025 list.
  2. Backdoor poisoning: a trigger phrase causes a chosen output while normal accuracy stays intact, evading evals.
  3. Web-scale sets are exposed: expired domains in URL-indexed datasets can be bought and their content swapped.
  4. Controls: dataset provenance and hashing, dedup and anomaly filtering, curated fine-tune data, red-team evals.
  5. RAG poisoning is cheaper: one planted doc in the index can steer answers without touching model weights.

If the interviewer pushes back

  • How would you test a third-party fine-tuned model for a backdoor you do not know the trigger for?
  • Why can a small number of poisoned samples be effective even in very large datasets?

Go deeper

cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.