Case 1B410D · AI security · L2 Practitioner
Chatbot Promises a Refund Policy
Practise as: Incident drill · Interview
Live alertA screenshot goes viral showing your customer-service chatbot, after a role-play jailbreak, promising a full refund on any purchase and insulting the company.
Interview questionA jailbroken support chatbot output is trending on social media. What are your first moves?
- 01 What happened?
- 02 What is the impact?
- 03 What do you do?
What a strong answer covers
Try it out loud first. Then check yourself:
- Contain fast: disable or degrade the bot to a scripted flow via a kill switch or feature flag.
- Verify with conversation logs that the screenshot is real, find the prompt chain, and check for copycat sessions.
- Loop in legal and comms: a tribunal held Air Canada liable for its chatbot (Moffatt, 2024), so clarify the actual policy.
- Harden: output classifier for commitments and toxicity, narrower system scope, retrieval-grounded policy answers.
- Add jailbreak prompts from the incident to a regression red-team suite before re-enabling.
If the interviewer pushes back
- How do you decide whether to honor commitments the bot made to real customers during the incident window?
cardosec draws a security topic and gives you a clock: explain it out loud with no notes, then see what you covered and what you missed. Free during early access.