Insights · AI & the enterprise · 9 min read
The AI pilots are not failing because AI was overstated
MIT found 95% of enterprise GenAI pilots produced no measurable P&L impact. Its own diagnosis — a learning gap, not a technology gap — is the part almost nobody quotes.
In 2025, MIT's Project NANDA published the most widely quoted number in enterprise AI: across an analysis of 300 public deployments, surveys of 153 leaders and 52 executive interviews, 95% of generative AI pilots delivered no measurable impact on the P&L. Somewhere between $30 and $40 billion had gone in.
The number travelled everywhere, usually as evidence that AI was oversold. That is not what the study found, and the difference matters to anyone deciding what to do next — because the two readings lead to opposite conclusions. If the technology was oversold, the rational response is to slow down. If the technology works and the adoption failed, the rational response is to change how you adopt.
What the study actually said
MIT's authors were specific about the cause, and it was not the models. They called it a learning gap: organizations had bought tools without building the capability to fold them into how work actually gets done. Of every hundred companies, sixty had deployed nothing at all. Forty had deployed something — and within those forty, only five had integrated it into a workflow at scale. The other thirty-five were running pilots that touched real work nowhere: proofs of concept that proved the concept and then had nowhere to live.
The supporting findings are just as instructive. Budgets flowed overwhelmingly toward sales and marketing uses — the demonstrations that impress in a board meeting — while the measurable returns clustered in operations, finance and back-office work, where hours are countable and processes repeat. Only two of nine major sectors examined showed material business change from generative AI at all. And deployments bought from specialized vendors or built with partners reached production roughly twice as often as internal builds, which inverts the instinct many technology organizations have to build AI capability from scratch as a learning exercise.
The inversion hiding in the data
Coverage of the study noted something almost comic sitting underneath the headline number. While official corporate pilots stalled, employees at the large majority of the same companies were quietly using personal AI tools for their daily work — drafting, summarizing, analyzing, coding — often against policy and usually without telling anyone.
Read those two facts together and the "AI does not work here" conclusion collapses. AI was working inside these companies every day. What failed was the program: the formally chartered, carefully governed initiative that never connected to the way work actually flows through the business. Adoption from below succeeded while adoption from above stalled — which tells you the constraint was never the model's capability. It was the organization's.
The 95% is not a verdict on AI. It is a census of companies that tried to adopt a practice by buying its tools.
Why committed companies still get nothing
The pattern repeats across industries with remarkable consistency. The commitment is real — boards have approved budgets, and a good number of CEOs are backing initiatives ahead of any formal business case. Gartner reached a similar place from different data, projecting that at least 30% of generative AI projects would be abandoned after proof of concept, citing poor data quality, inadequate risk controls and unclear business value.
What is missing in the stalled majority is rarely enthusiasm and almost never budget. It is experience. Nobody in the building has run this specific change before, so the default plan assembles itself from familiar parts: license the tools, book a training course, nominate a champion, run a pilot. Every step is individually reasonable. The combination reliably produces the 95%, because using AI well is a practice — how work gets broken down, what gets checked, where a human stays accountable, what "done" means when a machine wrote the first draft — and practices are not transferred by courses. They are transferred by working alongside people who already have them, on real work, until the judgement rubs off.
What the 5% do differently
Read across MIT's successful cases and the shape is consistent enough to act on.
They start where hours are measurable rather than where demos are impressive — operations, service, finance, the work that repeats — because a repeating process gives AI something to compound against, and gives the organization a number that either moves or does not.
They put AI inside an existing workflow rather than beside it. The stalled pilots almost all lived in a parallel world: a separate tool, a separate login, a separate habit that nobody had time to form. The integrated deployments changed a step in a process people already ran every day, which means adoption was not a favor employees did the program — it was the path of least resistance.
They bought or partnered where the capability already existed, and reserved internal build for what was genuinely specific to them. And someone experienced was accountable for what the AI produced — not a steering committee, a person, with the standing to say no to output that looked plausible and was not.
The question to take into your next review
The useful question about any AI initiative is no longer "is the technology ready" — the shadow usage inside your own company has already answered that. The question is which side of the divide the initiative is designed for: whether it changes a process someone runs every day, whether the hours it claims to save are counted anywhere, and whether anyone with real experience is accountable for the gap between plausible and correct.
Companies that can answer those three questions tend not to appear in the 95%.