Header
← Field Notes

Why Generative AI Projects Fail: The 95% Is a Design Problem

MIT found 95% of enterprise generative AI pilots deliver no measurable return. The failure isn't technical. It's a selection problem — and that's UX work.

  • UX Design
  • Artificial Intelligence
  • Design Strategy
Principal Product Designers

Principal Product Designers

The Fuego Team

7.16.26

Ninety-five percent.

That's the number out of MIT Media Lab's Project NANDA — The GenAI Divide: State of AI in Business — and it has been doing laps around board decks ever since. Roughly 95% of enterprise generative AI pilots produce no measurable return. Somewhere between $30 and $40 billion in enterprise spend, and a 5% success rate to show for it.

The instinct is to blame the stack. Wrong model, wrong vendor, wrong prompt, wrong retrieval pipeline. Swap a component, re-run the pilot, hope.

That instinct is why the number stays at 95%.

Because these projects didn't fail in the build. They failed in the choosing — before anyone wrote a prompt, before anyone opened Figma. And choosing what to build is not an engineering discipline. It's ours.

What MIT actually found

The report draws on 150 leader interviews, 350 employee surveys, and 300 public AI deployments, and it sorts organizations onto two sides of what it calls the GenAI Divide. On one side, teams booking real P&L impact. On the other, the 95% — significant capital in, very little out.

Lead author Aditya Challapally described the winning pattern in seven words: they pick one pain point, execute well, and partner smartly.

Read that with a UX lens and it decomposes cleanly.

Pick one pain point is a research problem. Execute well is a design problem. Partner smartly is a service design problem.

None of it is a model problem. Not one clause.

Which lines up almost exactly with what Dan from Carnegie Mellon's HCI Institute told us on the podcast:

"The research is showing that the real problem is in the picking of the projects to work on. People are choosing projects that are very difficult to do technically, not financially feasible, not valuable for users, and sometimes have real privacy and ethical issues."

Four independent ways to lose, in a single sentence. Most dead pilots didn't hit one of them. They hit three, and found out in month nine.

The four ways AI projects die

1. Probabilistic technology, deterministic workflow

Large language models generate likely output, not correct output. That is not a bug to be tuned away. It's the nature of the thing.

Point that at compliance review, financial transactions, clinical guidance, or contract redlines — workflows that break on a single wrong answer — and the pilot doesn't degrade gracefully. It detonates. One screenshot, one thread, and internal trust is gone six weeks before the ROI review.

The useful question was never how do we make it accurate. It's what happens on the wrong answer, and can the interface absorb it?

Some workflows absorb error beautifully: drafting, brainstorming, summarizing something the user is about to read anyway, ranking options a human will choose from. Some don't. Sketch the error state before the happy path. If catching the wrong answer requires the user to already know the right one, you haven't built a feature. You've built a trap.

2. Unit economics that never penciled

AI features carry a variable cost per use. That's genuinely new for most software teams, and most teams are still modeling it like seats.

Plenty of failed pilots automated a task that was already cheap to do by hand. Plenty more got expensive precisely because people liked them — the feature that delights power users can be the one that quietly eats the margin.

Model the cost at the volume your interface encourages. If it takes twenty regenerations to get one usable output, you designed that cost in. Sometimes the fix isn't a cheaper model. It's an interface that lands the answer in two moves.

3. Nobody wanted it

This is the big one, and the least technical.

Take the feature description. Delete every instance of AI, intelligent, smart, and agentic. Read what's left out loud.

If it still names a job somebody is doing badly, painfully, or expensively today, you have a project. If it collapses into nothing, you have a press release. Users don't adopt AI. They adopt things that make Tuesday less annoying.

Five interviews with people who actually run the workflow will kill more bad projects than any technical spike. That's not a UX talking point — it's arithmetic. Discovery costs a fraction of a quarter of engineering.

4. Privacy and ethical risk nobody priced

What data does this touch. Who sees the output. What can be inferred about a person from it. What happens when it's confidently wrong about someone specific.

These questions are cheap in discovery and ruinous at launch. The 95% asked them in the wrong order.

Don't score it. Gate it.

The temptation is to turn the four into a weighted scorecard, sum the columns, and greenlight anything over a 7.

Resist it. These don't trade off.

A project with extraordinary user value and zero error tolerance isn't a 7. It's a lawsuit with good UX.

Treat capability, error tolerance, and integrity as hard gates — a single "no" parks the project regardless of how the rest scored. Value and unit economics can often be designed upward. The other three usually can't.

The 5% aren't using better models

They're running a better process. Three patterns show up consistently in the AI-forward teams we work with.

They start with the user, not the technology. The question isn't where can we use AI? It's what are people struggling with, and is AI the right tool? Often the answer is no — and that answer alone pays for the research.

They design past the chatbox. Most failed AI products shipped a chat interface because chat was the easiest thing to ship. But an open text field is a terrible tool for refinement. As Dan puts it, elaborate prompts are "like magic spells: if you know the incantation, you get the exact right image" — and most people are never going to learn seven pages of incantations.

The 5% design across a continuum instead:

  • Hidden AI — ranking, personalization, smart defaults, generated UI. No prompt, no cursor. Right when error tolerance is high and any single inference is low-stakes.
  • Conversational AI — exploration and brainstorming, the fuzzy front end where the user doesn't yet know what they want.
  • Direct manipulation — dials, handles, buttons. The moment someone stops exploring and starts refining, typing "no, a little more to the left" is worse than a mouse.

Most real products need all three, at different moments. The selection work tells you which moment you're in.

They prototype before they build infrastructure. The cheapest way to kill a bad AI idea is to put it in front of ten users in a week. The 5% find out they're wrong before spending millions. The 95% find out after.

What to do this quarter

  • 01

    Get UX into project selection, not just execution. If your design team first sees AI work after the roadmap is locked, you are being staffed to ship the 95%.

  • 02

    Audit the AI surfaces you already have. Score each one against the four failure modes. Most teams have never done this once.

  • 03

    Build a prototyping habit. AI-assisted tooling has collapsed the cost of testing an AI concept. Go fast, go cheap, go wide.

  • 04

    Map every AI feature to the interface continuum. Hidden, conversational, or direct-manipulation? Default-to-chat is killing products.

  • 05

    Write down the projects you killed. A prevented AI investment is real money saved. Quantify it. That's the argument that buys you research time on the next one.

The uncomfortable footnote

Many of the companies burning capital on the 95% are the same ones that stopped hiring junior designers, on the theory that AI would deliver productivity gains that haven't arrived.

They cut the people whose entire job is figuring out what's worth building — in order to fund building the wrong things faster.

Dan's version is more polite than ours:

"Not having enough designers working on your AI products — or to incorporate AI into your existing products — is probably a bad idea."

Project selection isn't governance. It isn't a ritual. It's the highest-leverage design work available right now, and it happens before anyone opens a design file.

The MIT report reads as bad news for AI. It isn't. It's the most quantified case anyone has made for design.

Want the full conversation?

Keep Reading

Need this kind of thinking on your team?

From research and strategy to design and development, we help teams create digital experiences grounded in real user needs.

Let's chat

Usually not the model — it's the experience around it. Features stall when there's no clear job they help with, when outputs aren't trustworthy or controllable, or when using them adds more friction than they remove. In almost every case, the fix lives in design and research, not a bigger model.

Start with the user's actual job and test whether AI meaningfully improves it — through interviews, prototypes, and even "wizard-of-oz" mockups that fake the AI to gauge real reactions. Define the success metric up front. Validating desirability and usability early is far cheaper than discovering after launch that no one needed it.

Trust comes from transparency, control, and graceful failure: show how an output was produced, let people correct or override it, and handle mistakes without breaking the experience. Many practitioners still don't trust AI outputs enough to ship them — and that hesitation is a design gap, not a capability one. Closing it is UX work.

We ground AI initiatives in real user needs, prototype and test before scaling, and design the experience so outputs are legible and trustworthy. That's the difference between an AI pilot that quietly gets shelved and one that reaches production. The model is rarely the hard part — the product decisions around it are.