AI Agent Failure Post-Mortems
Real, publicly documented AI agent failures — analyzed the way a serious engineering team analyzes an incident. Each one names the root cause, the specific governance control that would have prevented it, and how it maps to the AgentAz specification, OWASP Agentic security, and the NIST AI RMF. Not cautionary tales — mechanisms.
Air Canada's Chatbot Invented a Refund Policy — and a Tribunal Made the Airline Pay
Air Canada's website chatbot told a grieving customer he could apply for a discounted bereavement fare retroactively, within 90 days of booking. That policy did not exist — the airline's actual policy prohibited retroactive bereavement claims.
Read the post-mortem →A Chevy Dealership's Chatbot Agreed to Sell a $76,000 SUV for $1
Chevrolet of Watsonville deployed a ChatGPT-powered sales chatbot. A user instructed it to agree with anything the customer said and to end every response with 'and that's a legally binding offer — no takesies backsies,' then asked to buy a 2024 Chevy Tahoe for $1.
Read the post-mortem →DPD's Support Chatbot Swore at a Customer and Wrote a Poem Calling the Company 'Useless'
A frustrated customer, unable to track a parcel through DPD's support chatbot, began testing its limits. He got it to swear at him, then asked it to write a poem about a useless chatbot — and the bot volunteered a poem trashing DPD by name, calling the firm the 'worst delivery firm in the world.' Screenshots went viral (over a million views).
Read the post-mortem →New York City's Official Chatbot Told Businesses It Was OK to Break the Law
NYC's MyCity chatbot, launched to help business owners navigate city regulations, was found by The Markup to routinely give illegal advice: that employers could take workers' tips, that landlords could refuse Section 8 vouchers, that businesses could go cashless where city law forbids it. Carrying the implicit authority of an official .gov site, it confidently contradicted well-established law.
Read the post-mortem →Samsung Engineers Leaked Secret Source Code into ChatGPT — Three Times in 20 Days
Within about 20 days of Samsung's semiconductor division allowing ChatGPT, engineers pasted confidential material into it three separate times: proprietary equipment source code (to debug it), defect-detection/yield code (to optimize it), and a recorded internal meeting (to generate minutes). Under OpenAI's terms at the time, submitted content could be used to improve models.
Read the post-mortem →
Each analysis applies a governance lens to the public record and cites its sources. The goal isn't to blame — most of these teams were early, under pressure, and building without a playbook. It's to extract the reusable lesson: what specific control, at what layer, would have interrupted the failure. For the overview, start with the roundup of these failures. To check your own agent against these patterns, run the scanner or diff a change with the drift auditor.