In mid-2025, MIT's NANDA initiative published a study that has since become the most quoted number in this industry: of the enterprise generative-AI pilots it examined, roughly 95% produced no measurable impact on the profit-and-loss statement. The research behind it — interviews with business leaders, employee surveys, and an analysis of some 300 public AI deployments — put a figure on something operators had been saying quietly for two years. The tools demo well. The businesses do not change.

The number went viral because it sounds like an indictment of the technology. Read the study, and it is nothing of the sort. It is an indictment of how the technology is bought, scoped, and integrated. That distinction matters, because a technology problem would be someone else's to fix. A procurement and process problem is yours — which means it is also fixable.

95%

of enterprise AI pilots produced no measurable P&L impact

67%

of partnered builds reached successful deployment, against 22% of internal ones

90 days

from pilot to full implementation, for the fastest movers

What the failed 95% have in common

The study's authors call the gap between adoption and transformation the "GenAI Divide." On one side: near-universal experimentation. On the other: a small minority extracting real value. Four patterns separate the sides, and none of them involves model quality.

  1. The money went to the wrong department.

    More than half of AI budgets in the dataset flowed to sales and marketing pilots — the visible, presentable use cases. The measurable returns, meanwhile, showed up in back-office work: document processing, compliance workflows, internal operations. The places nobody demos at a board meeting are the places the numbers actually move.

  2. The tools could not learn the business.

    MIT describes a "learning gap": systems that perform in a controlled demonstration but cannot retain context, adapt to feedback, or connect to the data they would need in production. A tool that has to be re-taught the business every session is not an employee. It is a party trick with a subscription fee.

  3. Partnered builds succeeded roughly three times as often as internal ones.

    Pilots that combined internal knowledge with external implementation expertise reached successful deployment far more often than IT-only internal builds — the study puts the two rates at roughly 67% versus 22%. Not because internal teams lack skill, but because integration is a discipline of its own, and most organisations exercise it once. An implementation partner exercises it weekly.

  4. Nobody defined what success meant.

    This is the quiet one, and in our view the decisive one. A pilot without a written success criterion cannot fail — and therefore cannot succeed. It simply runs until enthusiasm or budget expires. When MIT's researchers looked at what the successful 5% shared, they found tightly scoped initiatives with a specific pain point, a specific workflow, and a measurable outcome agreed before work began.

The speed finding nobody markets

Buried in the same research is a finding that received a fraction of the headline's attention: the organisations that moved fastest from pilot to full implementation did so in around 90 days, while the slowest took nine months or longer. What separated them was not budget, and it was not headcount. The fast movers had short decision chains, one clearly owned process, and someone who felt the cost of the broken workflow personally. The slow movers had committees, parallel pilots, and no individual whose name was attached to the outcome.

This cuts both ways, whatever the size of the company. A small firm has the fast profile by default — and squanders it the moment it chases three tools at once. A large organisation has the slow profile by default — and escapes it the moment one team is allowed to run one scoped project to completion before anything scales. The research is blunt on this point: the organisations running the most pilots completed the fewest.

What the 5% actually do

Strip the study down to its operational lessons and you get a short list. It reads less like an AI strategy and more like ordinary engineering discipline — which is precisely the point.

  1. They start with the process, not the tool.

    The failed pilots began with a technology looking for an application. The successful ones began with a bottleneck — a specific workflow where hours, errors, or cycle time could be counted — and selected tools last. If the process is not mapped, the automation automates the confusion.

  2. They write the number down first.

    Hours saved per week. Error rate before and after. Days of cycle time removed. A criterion agreed in writing before the build starts does two things: it forces honest scoping, and it makes the result checkable by someone with no stake in the project's reputation. Any partner unwilling to sign a number before building is telling you something about their confidence in the outcome.

  3. They integrate into the systems already running.

    The pilots that survived were embedded in the ERP, the CRM, the accounting stack — the places where work actually happens. The ones that died lived in a separate tab. A tool that requires people to leave their workflow to use it will be abandoned the first busy week, and every week is a busy week.

  4. They measure after, and publish what they find.

    Including the misses. The study notes that user trust collapsed fastest around tools whose promised results were never verified. Measurement is not bureaucracy; it is the mechanism by which the second project gets easier to approve than the first.

The uncomfortable question to ask your next vendor

If there is one practical takeaway, it is a question: "What is the written success criterion, and what happens if we miss it?"

A vendor selling a demo will answer with capabilities. A partner selling an outcome will answer with a number, a measurement method, and a date. The 95% failure rate is, at root, a market full of the first answer being purchased by buyers who needed the second.

The technology works. The MIT study, for all its bleak headline, documents companies extracting substantial value from the same models everyone else has access to. The difference was never the model. It was whether anyone in the room could say, in one sentence, what the system was supposed to change — and prove afterwards that it did.


Sources: MIT NANDA, "The GenAI Divide: State of AI in Business 2025" (July 2025); reporting on the study by Fortune and Forbes (August 2025). Figures cited as reported in the study and its coverage.