Fortune 500 Companies Let AI Make Fatal Mistakes
Updated: Sep 11

A drive-thru speaker adds bacon to a customer's ice cream order. A consulting report cites a book that was never written. A grieving passenger has to sue an airline because its chatbot lied to him. These aren't hypothetical warnings about some future AI dystopia. They already happened, at McDonald's, Deloitte, Air Canada, and UnitedHealth. Four organizations with the budgets, engineers, and leadership that should know better.
AI failures are costly, embarrassing, and almost always avoidable. This article breaks down four real AI blunders from some of the biggest companies on the planet, the single thread that ties them together, and what your company can do to avoid joining the list.
Generative AI is a wonderful tool when it's used properly. It's brilliant and a huge time-saver once you know where it fits. For a rundown of use cases that make sense (and ones that don't), see my earlier piece, 95% of AI Pilots Fail. Here's What the Other 5% Got Right.. In short, AI works very well on problems that involve complex pattern matching, and fails at problems that require complex reasoning. That one distinction is the thread running through every failure below.
McDonald's Failed Drive-Thru Assistant

McDonald's partnered with IBM to build an automated drive-thru assistant. The Automated Order Taker (AOT) took roughly three years to develop and was piloted at more than 100 U.S. locations before being abruptly pulled in June 2024. The system was caught on video taking an order for 260 Chicken McNuggets and adding bacon to a customer's ice cream. Other customers reported cancelling a single item and getting eight or nine different ones back. The experiment wasn't just a demonstration of terrible customer service. It was an expensive, multi-year failed tech project.
The failure appears to trace back to how the system was tested. Accents, background noise, and slurred speech rarely show up in a quiet office environment. But a customer ordering at 3 a.m., trying to sober up after last call, in the middle of a thunderstorm, is a realistic scenario, not an edge case. Anomalies happen with any technology, but an AI system is only as good as the data and conditions it's trained on. Had a human been monitoring the system in real time, that 260-nugget order likely would have been caught and corrected before it ever reached the window.
Deloitte's Error-Filled Report for the Australian Government

Deloitte was contracted to help Australia's Department of Employment and Workplace Relations produce a welfare compliance report. The 237-page document was riddled with AI-generated hallucinations, including fabricated footnotes, fake quotes, and a citation for a book that doesn't exist. Chris Rudge, a University of Sydney researcher, spotted the fabrications and flagged them publicly.
The fallout cost Deloitte A$97,000 in refunds, a fraction of the A$440,000 contract, but a real stain on a firm whose entire business depends on being a trusted source of truth. A round of human review, or even a quick cross-check with a second model, likely would have caught most of these errors. Rushing a high-stakes report through a single AI pass, unchecked, is a risky way to do business.
Air Canada Disowns Its Own Chatbot

Booking a flight on short notice after a death in the family is stressful enough without an airline actively misleading you. That's what happened to Jake Moffatt, who needed a last-minute ticket from Vancouver to Toronto for his grandmother's funeral. When he asked about bereavement fares, Air Canada's chatbot told him he could pay full price and then apply for the discount retroactively, within 90 days. That wasn't true. Air Canada's actual policy requires the discount to be requested before travel, not after. When Moffatt tried to claim the refund the bot had promised, a human representative denied it and offered him a $200 travel voucher instead. He declined and took the airline to Canada's Civil Resolution Tribunal.
Air Canada argued it shouldn't be held responsible for what its own chatbot told a customer. The tribunal disagreed, ruling that a company is responsible for the information its virtual agents provide, and ordered Air Canada to pay Moffatt $812.02.
Humans make mistakes too, but this one was preventable with better guardrails. Air Canada could have restricted the bot from discussing refund policy at all, or required it to route anything involving money to a human agent. A chatbot that's free to invent financial promises on a company's behalf is a liability waiting to happen.
UnitedHealth Denies Care to Elderly Patients

UnitedHealth, the largest health insurer in the U.S., built an AI tool to help make coverage determinations for Medicare Advantage patients. According to CBS News, a lawsuit alleges the company knowingly deployed a model with a 90% error rate. Nonetheless, UnitedHealth used it to override physicians' treatment recommendations. That's more than an ethics problem. It could easily become a fatal one. The suit, filed in federal court in Minnesota, claims elderly patients were wrongly denied medically necessary care they were entitled to under their Medicare Advantage plans.
This is the most egregious example in this article. If even a fraction of the reporting holds up, no responsible organization should have deployed a model that unreliable in a domain where the cost of being wrong is a person's health. When an algorithm is less accurate than a coin flip, the fix isn't better QA testing. It's holding the people who approved its use accountable.
NCSBN Uses AI to Protect the Integrity of the NCLEX Exam
Four disasters in a row can make it feel like AI itself is the problem. It isn't. Before wrapping up, it's worth looking at one organization that used the exact same technology and got it right, not to let the four companies above off the hook, but to show that none of this was inevitable.
Every year, more than 300,000 people sit for the NCLEX, the exam that decides who's allowed to practice nursing in the United States. Let a single flawed question slip through and you could fail a competent nurse or pass someone who shouldn't be treating patients. The organization behind that exam, the National Council of State Boards of Nursing (NCSBN), a non-profit, has started using AI in parts of its exam-development process, including AI-assisted item banking. But it hasn't handed the pen to a machine.
Every NCLEX question is still drafted by practicing nurses and subject-matter experts, then passed through sensitivity panels and bias-review committees before it's used on a real exam. NCSBN holds every item to the standard it always has, valid, reliable, psychometrically sound, legally defensible, and fair, and treats AI as a tool that supports that process rather than a shortcut around it. At NCSBN's own 2025 conference on AI in nursing regulation, technology consultant Jack Shaw told the room to keep humans in the loop and let people make the final call. A state regulator on the same panel went further, warning that if AI ever became the final word in a licensing decision, it would open the door to legal trouble.
The bigger difference might be incentives. NCSBN isn't chasing a quarterly earnings target or racing a competitor to market. As a non-profit responsible for an exam that determines who's allowed to practice nursing, it has little reason to rush AI into production faster than its testing can support. That's the exact incentive that was missing at McDonald's, Deloitte, Air Canada, and UnitedHealth. AI didn't fail at those four companies because the technology is inherently unreliable. It failed because it was pushed out faster than anyone bothered to check it.
The Common Thread
Every one of the failed projects above went wrong during, or because of, a skipped testing phase. That's true of most software projects that flop after launch, whether AI is involved or not. Whether it's a rushed deadline that cut QA short, poor planning at the executive level, or a rush to claim an early lead in the AI space, the result is the same. In every case here, the failure was costly and unnecessary.
My bigger concern is how little management, in general, understands about how generative AI actually works. These are machine learning models, statistical systems trained on data, and they will never be perfect. That's exactly why a human needs to stay in the loop. Fears that AI will wipe out every job and eventually turn on humanity, Terminator-style, are overblown. But executives looking for a shortcut to cut headcount and inflate margins should be far more cautious than they currently are. Few people outside AI research can explain exactly how a transformer works or where the vectors in a RAG pipeline come from, and that's fine. What matters is understanding, at a basic level, that these are pattern-matching systems, not reasoning engines. They're extraordinarily good at finding structure in data. They are not good at knowing when they're wrong.
Companies that use AI responsibly will win in the long run, and it comes down to three rules that aren't complicated.
● Set the guardrails before you scale. Decide what the AI is and isn't allowed to say or decide, especially anything touching money, medical care, or legal claims.
● Keep a human with real authority in the loop. Not someone who rubber-stamps what the model says, but someone with the power to override it similar to the way NCSBN's expert panels do on every single exam question.
● Test with real people in real conditions. A quiet office and a QA checklist won't surface the 3 a.m. drive-thru customer, the pranking teenager, or the grieving passenger who takes you to a tribunal. Real users will.
Not every failure can be caught in advance. But most of the ones in this article could have been.






Comments