5 AI Projects That Failed to Deliver ROI (and Why)

The five failures below span fifteen years, four industries, and somewhere north of half a billion dollars in documented losses. They were run by well-resourced organisations with access to the best available technology.
What's striking is that not one of them failed because the AI wasn't capable enough. Each failed on something adjacent: data that couldn't be reached, an error tolerance nobody defined, deployment conditions the pilot never tested. Studying AI project failure is useful precisely because the causes repeat, and they repeat in places most project plans don't look.
1. IBM Watson for Oncology at MD Anderson
What happened: MD Anderson Cancer Center built an Oncology Expert Advisor on IBM Watson, intended to recommend treatment paths by ingesting oncology guidelines, trial records, and patient charts. A University of Texas audit found the centre had spent roughly $62 million before the project was cancelled in 2016. It never entered clinical use.
Why it failed: The audit surfaced a detail that should unsettle anyone planning an AI project: the system could not sync with MD Anderson's Epic electronic health record, and ran on outdated data as a result. The most advanced medical AI of its era could not read the patient records sitting in the same building. Separately, the wider Watson for Oncology programme drew criticism for being trained substantially on hypothetical cases rather than real patient outcomes.
The lesson: integration is not a late-stage implementation detail. A model that cannot access live production data is a demonstration, however sophisticated its reasoning, and no amount of model quality compensates for a connection that doesn't exist.
Hire Edge Computer Vision
2. Zillow Offers
What happened: Zillow used its valuation models to buy homes directly, renovate, and resell. In late 2021 the company wound the programme down, wrote down homes it had overpaid for, and cut roughly 25% of its workforce, about 2,000 jobs, with reported losses on the programme exceeding $400 million.
Why it failed: The valuation model wasn't absurd. Automated valuation models are genuinely accurate within a few percent on standard properties in stable markets. The problem was what the model was wired to. An estimate with a few percentage points of error is perfectly useful as guidance to a human buyer and dangerous as the trigger for an irreversible purchase at scale, because the errors compound across thousands of transactions and cannot be unwound when the market moves.
The lesson: match the model's error profile to the consequence of being wrong. The failure here was an error tolerance mismatch, not a modelling one: a probabilistic output was given authority to make deterministic financial commitments with no human checkpoint between estimate and purchase.
3. Amazon's Recruiting Engine

What happened: Amazon built an experimental tool to score job applicants, developed from 2014. It was scrapped in 2018 after the team found it systematically disadvantaged women, penalising resumes that contained the word "women's" and downgrading graduates of women's colleges. Amazon has said the tool was never the sole basis for hiring decisions.
Why it failed: The model did exactly what it was trained to do. It learned from a decade of submitted resumes and hiring outcomes in an industry where those outcomes skewed heavily male, so it correctly identified the historical pattern and reproduced it. The objective it was optimising, match the profile of people we hired before, was not the objective anyone actually wanted.
The lesson: a model trained on historical decisions inherits the logic of those decisions, including the parts nobody would defend out loud. Where the training signal is past human judgement, the question to ask before building is whether that judgement is something you want to scale.
4. McDonald's and IBM's Drive-Thru Voice Ordering
What happened: A partnership formed in 2021 deployed automated voice ordering across roughly 100 restaurants. McDonald's ended the test in 2024. Analyst reporting during the trial put order accuracy in the low-to-mid 80% range, and franchisee feedback was consistently underwhelming.
Why it failed: Sources familiar with the technology cited difficulty interpreting different accents and dialects, which directly affected order accuracy. A drive-thru is an unusually hostile environment for speech recognition: engine noise, variable microphone distance, background conversation, regional speech variation, and customers who change their minds mid-sentence. None of that shows up in a controlled demonstration.
The lesson: pilot conditions are not deployment conditions, and accuracy measured in one is not predictive of the other. A scoped pilot run on genuinely representative real-world input is the cheapest way to find this out, and building an AI PoC covers structuring one so it tests the hard cases rather than the convenient ones.
Hire Edge Computer Vision
5. Air Canada's Customer Service Chatbot

What happened: A passenger asked Air Canada's website chatbot about bereavement fares. The chatbot described a policy allowing retroactive application for the discount. No such policy existed. In February 2024 a Canadian tribunal held the airline liable for the misinformation and ordered it to compensate the passenger.
Why it failed: The airline's defence included the argument that the chatbot was a separate entity responsible for its own actions. The tribunal rejected it, which established the thing every business deploying a customer-facing assistant needs to understand: you are accountable for what it says. The technical cause was an ungrounded generative system permitted to state policy without retrieval against the actual policy documents or human review.
The lesson: the financial loss here was trivial; the precedent was not. Deciding what an AI system is allowed to assert, and on what authority, is an architecture decision made before launch. AI agent security covers scoping that authority deliberately rather than discovering its limits through a ruling.
What the Five Have in Common
Placed side by side, the failure modes are distinct but the meta-pattern is consistent.
|
Project |
Failure mode |
Where it should have been caught |
|---|---|---|
|
Watson at MD Anderson |
Couldn't reach the production data |
Integration feasibility, before build |
|
Zillow Offers |
Error tolerance mismatch |
Scoping: what does being wrong cost? |
|
Amazon recruiting |
Objective mismatch in training signal |
Problem definition: is the historical pattern the goal? |
|
McDonald's drive-thru |
Pilot conditions unlike deployment |
Pilot design, on representative input |
|
Air Canada chatbot |
Unbounded authority, no grounding |
Architecture: what may it assert? |
Every one of these was decidable before significant money was committed, and none required better models to avoid. They required questions asked earlier: can we actually reach this data, what does an error cost, is the training signal the outcome we want, will the pilot resemble reality, and what is this system permitted to do unsupervised.
That's also a reasonable checklist for evaluating anyone proposing to build for you. A partner who raises these unprompted has seen projects fail; one who doesn't may be about to find out. choosing an AI development partner covers the wider evaluation.
Hire Edge Computer Vision
