Case 9A84E1 · AI security · L3 Applied
Data and Model Poisoning
Practise as: Explain it · Deep dive
- 01 What is it?
- 02 How is it abused?
- 03 How do you stop it?
What a strong answer covers
Try it out loud first. Then check yourself:
- Attacker manipulates training, fine-tuning or RAG data to change model behavior. OWASP LLM04 in the 2025 list.
- Backdoor poisoning: a trigger phrase causes a chosen output while normal accuracy stays intact, evading evals.
- Web-scale sets are exposed: expired domains in URL-indexed datasets can be bought and their content swapped.
- Controls: dataset provenance and hashing, dedup and anomaly filtering, curated fine-tune data, red-team evals.
- RAG poisoning is cheaper: one planted doc in the index can steer answers without touching model weights.
If the interviewer pushes back
- How would you test a third-party fine-tuned model for a backdoor you do not know the trigger for?
- Why can a small number of poisoned samples be effective even in very large datasets?
Go deeper
cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.