Case D248DF · AI security · L2 Practitioner
Jailbreaks
Practise as: Explain it · Interview
Interview questionWhat is the difference between a jailbreak and a prompt injection? Give examples of jailbreak techniques.
- 01 What is it?
- 02 How is it abused?
- 03 How do you stop it?
What a strong answer covers
Try it out loud first. Then check yourself:
- Jailbreak: bypassing the model's own safety training. Injection: overriding the application's instructions.
- Techniques: role-play personas, hypothetical framing, encoding (base64, leetspeak), low-resource languages.
- Many-shot jailbreaking fills a long context with fake compliant Q&A pairs to shift behavior.
- Optimized adversarial suffixes (GCG) are gibberish token strings found by gradient search that can transfer to other models.
- Defenses: safety fine-tuning, input/output classifiers, conversation-level monitoring, and limiting tool impact.
If the interviewer pushes back
- Why do output classifiers often catch jailbreaks that input filters miss?
cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.