Open cardosec

Case D248DF · AI security · L2 Practitioner

Jailbreaks

Practise as: Explain it · Interview

Interview questionWhat is the difference between a jailbreak and a prompt injection? Give examples of jailbreak techniques.

  1. 01 What is it?
  2. 02 How is it abused?
  3. 03 How do you stop it?
What a strong answer covers

Try it out loud first. Then check yourself:

  1. Jailbreak: bypassing the model's own safety training. Injection: overriding the application's instructions.
  2. Techniques: role-play personas, hypothetical framing, encoding (base64, leetspeak), low-resource languages.
  3. Many-shot jailbreaking fills a long context with fake compliant Q&A pairs to shift behavior.
  4. Optimized adversarial suffixes (GCG) are gibberish token strings found by gradient search that can transfer to other models.
  5. Defenses: safety fine-tuning, input/output classifiers, conversation-level monitoring, and limiting tool impact.

If the interviewer pushes back

  • Why do output classifiers often catch jailbreaks that input filters miss?

cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.