Start a conversation Contact
← Research

Most pilots return nothing, and the model is not the reason

The widely quoted finding that almost no enterprise AI pilot produces a measurable financial result is not a verdict on model capability. It is a description of what happens when a tool is bought without changing the workflow it was bought to change.

The number has been in circulation long enough that it now arrives without its source. Somebody says that ninety-five per cent of enterprise AI pilots fail, the room reacts, and the conversation goes to one of two places: this proves the technology is oversold, or this proves everyone else is doing it wrong.

Neither is what the underlying work says, and the actual finding is more useful than either.

The source is The GenAI Divide: State of AI in Business 2025, from MIT’s Project NANDA. It examined enterprise deployments and reported that around 95% of the generative AI pilots it looked at produced no discernible financial result, drawing on interviews, a survey of employees and an analysis of public deployments — Fortune’s account of the study sets out the method, which matters here because the report circulates as a PDF rather than through a journal and the headline has travelled a long way from it. What the authors attribute the gap to is not model quality. It is a learning gap: tools that never entered the workflow they were bought to change.

The claim: a pilot that does not change a workflow cannot produce a financial result, and most pilots are designed in a way that makes changing the workflow somebody else’s job.

What a pilot usually measures

The gap between what a pilot proves and what a business case needs Fig. 01
What the pilot demonstrated What the business case assumed
The model produces good output on sample tasks Staff will use it on real tasks, under time pressure
Users report the tool is helpful Handling time falls, and the saved time is redeployed
The output is accurate on the cases tried The output is trusted enough to act on without re-checking
The integration works The process around it changed — approvals, handoffs, staffing

The third row is where most of the missing value goes. A system that produces a good draft which a person then reads in full, checks against the source and rewrites has not removed the work. It has moved it, and in some cases added to it. Whether that changes depends on whether anyone was allowed to remove the verification step — which is a governance decision, not a model capability, and it is almost never inside the pilot’s scope.

The shape that produces a result

What a pilot has to establish before it can return anything Fig. 02
  1. 01 A named process One workflow, with a measured baseline before anything changes.
  2. 02 A decision to remove Which step goes away, and who is accountable for the risk of removing it.
  3. 03 A trust threshold The accuracy at which the removed step is not needed, agreed in advance.
  4. 04 The measurement Same units as the baseline, on the same population, after the change.

Stage two is the one that separates the pilots that return something from the pilots that return a report. Somebody has to be willing to say that a check is no longer performed, or performed on a sample, or performed only above a threshold. That is an accountability transfer and it cannot be made by the AI team. Where it is not made, the tool sits alongside the existing process, and the honest financial result of a tool that sits alongside an unchanged process is the cost of the tool.

Stage one is the one that is skipped for the most understandable reason: nobody measured the process before. The baseline does not exist, so the improvement cannot be stated, so the pilot is judged on enthusiasm. It is the same failure that makes an expensive model impossible to justify against a statistical one — the absence of a number that was cheap to collect at the start and impossible to reconstruct later.

What to require of the next pilot

The fifth line is the one that changes behaviour most. Pilots without a stated failure condition do not fail. They are extended, rescoped, and eventually absorbed into business as usual with the question unanswered, which is how an organisation accumulates twelve systems nobody can evaluate and a growing sense that none of it is working.

What this does not tell you

The headline figure is one study, with one definition of measurable impact, over one sample of deployments in one period. It is not a law, and treating it as one produces the mirror-image error of the hype it corrects. Plenty of organisations have deployed systems that return value and were never in that sample.

It also does not follow that pilots are the wrong instrument. They are the right instrument for reducing uncertainty. The argument is that a pilot which reduces uncertainty about model quality — a question the published evaluations already answer reasonably well — has spent an organisation’s scarce attention on the least uncertain part of the problem. The uncertain parts are whether the workflow can change, whether the people in it will trust the output, and whether anyone will sign for removing a control.

The person who should read this differently is whoever approves the next round of pilots. Ask which step disappears if it works, and who has agreed to that. If nobody can answer, the pilot is a demonstration, and demonstrations belong in a different budget line with a different expectation attached.

Filed under · Method · Adoption · Method · Operating model Inference Institute · 07 Jul 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.