Case 7866C1 · AI security · L3 Applied
Training Data Leakage
Practise as: Explain it
- 01 What is it?
- 02 How is it abused?
- 03 How do you stop it?
What a strong answer covers
Try it out loud first. Then check yourself:
- Models memorize rare or duplicated sequences; extraction attacks recover verbatim PII, keys or code from them.
- Membership inference: decide whether a specific record was in the training set from loss or confidence signals.
- Fine-tuning on support tickets or internal docs can leak one customer's data to another via completions.
- Mitigate: scrub PII/secrets before training, deduplicate, differential privacy (DP-SGD), output PII filters.
- Maps to OWASP LLM02 Sensitive Information Disclosure; also a GDPR issue since deletion from weights is hard.
If the interviewer pushes back
- A user invokes their right to erasure. What are your realistic options if their data was in a fine-tuning set?
cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.