Open cardosec

Case 8D7042 · AI security · L3 Applied

Model Extraction Attacks

Practise as: Explain it
  1. 01 What is it?
  2. 02 How is it abused?
  3. 03 How do you stop it?
What a strong answer covers

Try it out loud first. Then check yourself:

  1. Attacker queries a model API at scale and trains a surrogate on the outputs, cloning capability (distillation).
  2. Returned logprobs/confidence scores leak far more signal than top-1 text and speed up extraction.
  3. Stolen surrogate enables IP theft and offline crafting of adversarial or jailbreak inputs that transfer back.
  4. Defenses: per-key rate limits and quotas, query-pattern anomaly detection, limit or round logprobs, ToS enforcement.
  5. Output watermarking or canary responses can give evidence (not proof) that a model was trained on your outputs.

If the interviewer pushes back

  • How would you distinguish a heavy legitimate customer from an extraction campaign using only API telemetry?

cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.