Case 8D7042 · AI security · L3 Applied
Model Extraction Attacks
Practise as: Explain it
- 01 What is it?
- 02 How is it abused?
- 03 How do you stop it?
What a strong answer covers
Try it out loud first. Then check yourself:
- Attacker queries a model API at scale and trains a surrogate on the outputs, cloning capability (distillation).
- Returned logprobs/confidence scores leak far more signal than top-1 text and speed up extraction.
- Stolen surrogate enables IP theft and offline crafting of adversarial or jailbreak inputs that transfer back.
- Defenses: per-key rate limits and quotas, query-pattern anomaly detection, limit or round logprobs, ToS enforcement.
- Output watermarking or canary responses can give evidence (not proof) that a model was trained on your outputs.
If the interviewer pushes back
- How would you distinguish a heavy legitimate customer from an extraction campaign using only API telemetry?
cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.