Open cardosec

Case 1B410D · AI security · L2 Practitioner

Chatbot Promises a Refund Policy

Practise as: Incident drill · Interview

Live alertA screenshot goes viral showing your customer-service chatbot, after a role-play jailbreak, promising a full refund on any purchase and insulting the company.

Interview questionA jailbroken support chatbot output is trending on social media. What are your first moves?

  1. 01 What happened?
  2. 02 What is the impact?
  3. 03 What do you do?
What a strong answer covers

Try it out loud first. Then check yourself:

  1. Contain fast: disable or degrade the bot to a scripted flow via a kill switch or feature flag.
  2. Verify with conversation logs that the screenshot is real, find the prompt chain, and check for copycat sessions.
  3. Loop in legal and comms: a tribunal held Air Canada liable for its chatbot (Moffatt, 2024), so clarify the actual policy.
  4. Harden: output classifier for commitments and toxicity, narrower system scope, retrieval-grounded policy answers.
  5. Add jailbreak prompts from the incident to a regression red-team suite before re-enabling.

If the interviewer pushes back

  • How do you decide whether to honor commitments the bot made to real customers during the incident window?

cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.