Follow Me

© 2026 Shreyans Padmani. All rights reserved.
5 AI Projects That Failed to Deliver ROI (and Why)
Artificial Intelligence

5 AI Projects That Failed to Deliver ROI (and Why)

Five documented AI project failures, from Watson at MD Anderson to Zillow Offers, and the specific failure mode behind each one.

5 AI Projects That Failed to Deliver ROI (and Why)
Share

5 AI Projects That Failed to Deliver ROI (and Why)

AI Generated Image

The five failures below span fifteen years, four industries, and somewhere north of half a billion dollars in documented losses. They were run by well-resourced organisations with access to the best available technology.

What's striking is that not one of them failed because the AI wasn't capable enough. Each failed on something adjacent: data that couldn't be reached, an error tolerance nobody defined, deployment conditions the pilot never tested. Studying AI project failure is useful precisely because the causes repeat, and they repeat in places most project plans don't look.

1. IBM Watson for Oncology at MD Anderson

What happened: MD Anderson Cancer Center built an Oncology Expert Advisor on IBM Watson, intended to recommend treatment paths by ingesting oncology guidelines, trial records, and patient charts. A University of Texas audit found the centre had spent roughly $62 million before the project was cancelled in 2016. It never entered clinical use.

Why it failed: The audit surfaced a detail that should unsettle anyone planning an AI project: the system could not sync with MD Anderson's Epic electronic health record, and ran on outdated data as a result. The most advanced medical AI of its era could not read the patient records sitting in the same building. Separately, the wider Watson for Oncology programme drew criticism for being trained substantially on hypothetical cases rather than real patient outcomes.

The lesson: integration is not a late-stage implementation detail. A model that cannot access live production data is a demonstration, however sophisticated its reasoning, and no amount of model quality compensates for a connection that doesn't exist.

Hire Edge Computer Vision

Expertly deploy AI at the edge, schedule a consultation

Get Free POC Scoping

 

2. Zillow Offers

What happened: Zillow used its valuation models to buy homes directly, renovate, and resell. In late 2021 the company wound the programme down, wrote down homes it had overpaid for, and cut roughly 25% of its workforce, about 2,000 jobs, with reported losses on the programme exceeding $400 million.

Why it failed: The valuation model wasn't absurd. Automated valuation models are genuinely accurate within a few percent on standard properties in stable markets. The problem was what the model was wired to. An estimate with a few percentage points of error is perfectly useful as guidance to a human buyer and dangerous as the trigger for an irreversible purchase at scale, because the errors compound across thousands of transactions and cannot be unwound when the market moves.

The lesson: match the model's error profile to the consequence of being wrong. The failure here was an error tolerance mismatch, not a modelling one: a probabilistic output was given authority to make deterministic financial commitments with no human checkpoint between estimate and purchase.

3. Amazon's Recruiting Engine

AI Generated Image

What happened: Amazon built an experimental tool to score job applicants, developed from 2014. It was scrapped in 2018 after the team found it systematically disadvantaged women, penalising resumes that contained the word "women's" and downgrading graduates of women's colleges. Amazon has said the tool was never the sole basis for hiring decisions.

Why it failed: The model did exactly what it was trained to do. It learned from a decade of submitted resumes and hiring outcomes in an industry where those outcomes skewed heavily male, so it correctly identified the historical pattern and reproduced it. The objective it was optimising, match the profile of people we hired before, was not the objective anyone actually wanted.

The lesson: a model trained on historical decisions inherits the logic of those decisions, including the parts nobody would defend out loud. Where the training signal is past human judgement, the question to ask before building is whether that judgement is something you want to scale.

4. McDonald's and IBM's Drive-Thru Voice Ordering

What happened: A partnership formed in 2021 deployed automated voice ordering across roughly 100 restaurants. McDonald's ended the test in 2024. Analyst reporting during the trial put order accuracy in the low-to-mid 80% range, and franchisee feedback was consistently underwhelming.

Why it failed: Sources familiar with the technology cited difficulty interpreting different accents and dialects, which directly affected order accuracy. A drive-thru is an unusually hostile environment for speech recognition: engine noise, variable microphone distance, background conversation, regional speech variation, and customers who change their minds mid-sentence. None of that shows up in a controlled demonstration.

The lesson: pilot conditions are not deployment conditions, and accuracy measured in one is not predictive of the other. A scoped pilot run on genuinely representative real-world input is the cheapest way to find this out, and building an AI PoC covers structuring one so it tests the hard cases rather than the convenient ones.

Hire Edge Computer Vision

Expertly deploy AI at the edge, schedule a consultation

Get Free POC Scoping

 

5. Air Canada's Customer Service Chatbot

AI Generated Image

What happened: A passenger asked Air Canada's website chatbot about bereavement fares. The chatbot described a policy allowing retroactive application for the discount. No such policy existed. In February 2024 a Canadian tribunal held the airline liable for the misinformation and ordered it to compensate the passenger.

Why it failed: The airline's defence included the argument that the chatbot was a separate entity responsible for its own actions. The tribunal rejected it, which established the thing every business deploying a customer-facing assistant needs to understand: you are accountable for what it says. The technical cause was an ungrounded generative system permitted to state policy without retrieval against the actual policy documents or human review.

The lesson: the financial loss here was trivial; the precedent was not. Deciding what an AI system is allowed to assert, and on what authority, is an architecture decision made before launch. AI agent security covers scoping that authority deliberately rather than discovering its limits through a ruling.

What the Five Have in Common

Placed side by side, the failure modes are distinct but the meta-pattern is consistent.

Project

Failure mode

Where it should have been caught

Watson at MD Anderson

Couldn't reach the production data

Integration feasibility, before build

Zillow Offers

Error tolerance mismatch

Scoping: what does being wrong cost?

Amazon recruiting

Objective mismatch in training signal

Problem definition: is the historical pattern the goal?

McDonald's drive-thru

Pilot conditions unlike deployment

Pilot design, on representative input

Air Canada chatbot

Unbounded authority, no grounding

Architecture: what may it assert?

Every one of these was decidable before significant money was committed, and none required better models to avoid. They required questions asked earlier: can we actually reach this data, what does an error cost, is the training signal the outcome we want, will the pilot resemble reality, and what is this system permitted to do unsupervised.

That's also a reasonable checklist for evaluating anyone proposing to build for you. A partner who raises these unprompted has seen projects fail; one who doesn't may be about to find out. choosing an AI development partner covers the wider evaluation.

Hire Edge Computer Vision

Expertly deploy AI at the edge, schedule a consultation

Get Free POC Scoping

 

 

Frequently asked questions

Why do most AI projects fail to deliver ROI?
Rarely because the model isn't capable. The recurring causes are data that can't actually be accessed in production, a mismatch between the model's error rate and the cost of being wrong, training objectives that don't match the real goal, pilot conditions that don't resemble deployment, and systems given authority to act or assert without grounding or review. All five are scoping problems, identifiable before major spend.
What happened with IBM Watson at MD Anderson?
MD Anderson spent roughly $62 million building an oncology treatment advisor on IBM Watson before cancelling the project in 2016. A University of Texas audit found the system couldn't sync with the centre's Epic electronic health record and ran on outdated data as a result. It never entered clinical use. The failure was largely one of integration rather than of the model's reasoning ability.
Why did Zillow Offers fail if the valuation model was accurate?
Because accuracy sufficient for guidance is not accuracy sufficient for irreversible commitment. Automated valuation models are typically within a few percent on standard properties in stable markets, which is useful advice to a human buyer. Wiring that output directly to purchase decisions at scale meant small percentage errors compounded across thousands of homes with no way to unwind them when the market moved.
What does the Amazon recruiting AI case teach about training data?
That a model trained on historical decisions inherits the logic of those decisions. Amazon's tool learned from a decade of resumes and hiring outcomes in a male-skewed industry, so it accurately identified and reproduced that pattern, penalising resumes containing "women's" and downgrading women's college graduates. The model performed its stated objective correctly; the objective itself was wrong.
Is a company legally responsible for what its AI chatbot says?
A Canadian tribunal ruled in February 2024 that Air Canada was liable after its website chatbot described a bereavement fare policy that did not exist. The airline argued the chatbot was a separate entity responsible for its own actions, and the tribunal rejected that. The practical implication is that what a customer-facing AI system is permitted to assert should be treated as an architecture decision made before launch.
How can a business avoid these failure patterns?
Ask five questions before committing significant budget: can we genuinely access this data in production, what does an error cost in each direction, is the training signal the outcome we actually want, will the pilot use genuinely representative conditions, and what is this system allowed to do or say without review. Each of the five failures above was avoidable at the scoping stage rather than requiring better technology.
Summarise this article with AI Open it in your assistant of choice.
ChatGPT Perplexity You AI Claude Groq
ai project failure AI ROI failure Watson MD Anderson Zillow Offers AI post-mortem Cost and ROI algorithmic bias deployment conditions AI lessons learned AI risk
Shreyans Padmani
Written by

Shreyans Padmani

100% Upwork JSSMicrosoft AI Certified12 case studies5+ years

Shreyans Padmani has 5+ years of experience leading innovative software solutions, specializing in AI, LLMs, RAG, and strategic application development. He transforms emerging technologies into scalable, high-performance systems, combining strong technical expertise with business-focused execution to deliver impactful digital solutions.

Where to go from here

Let's talk about your project

Bring the problem you're solving, the metric you want to move, and where the data lives. You leave the call with a scoped project and a realistic timeline.

AI Summarizer