Request a scoping call Contact
← Research

Give the pilot a measurable operating decision

A pilot needs workflow and economic evidence as well as model output. Define the operating change, baseline and decision rule before interpreting adoption or benchmark scores as financial value.

Method / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

A widely repeated claim about enterprise AI pilots is often read as a verdict on the technology. The cited report concerns a specific sample and definition of financial impact. Read it within that scope before setting expectations for another programme.

The local question is whether the pilot can establish the result its business case assumes. Model output, staff interest and technical integration do not answer it alone.

The source is The GenAI Divide: State of AI in Business 2025, from MIT’s Project NANDA. It examined enterprise deployments and reported that around 95% of the generative AI pilots it looked at produced no discernible financial result, drawing on interviews, a survey of employees and an analysis of public deployments — Fortune’s account of the study sets out the method, which matters here because the report circulates as a PDF rather than through a journal and the headline has travelled a long way from it. What the authors attribute the gap to is not model quality. It is a learning gap: tools that never entered the workflow they were bought to change.

Design the pilot around a named workflow and observable outcome. Specify what changes, who owns that change and what supports further investment. Assess model quality alongside adoption, review effort and operating constraints.

What a pilot usually measures

The gap between what a pilot proves and what a business case needs Fig. 01
What the pilot demonstrated What the business case assumed
The model produces good output on sample tasks Staff will use it on real tasks, under time pressure
Users report the tool is helpful Handling time falls, and the saved time is redeployed
The output is accurate on the cases tried The output is trusted enough to act on without re-checking
The integration works The process around it changed — approvals, handoffs, staffing

A draft may still require checking and repair. Measure that effort before claiming released capacity. Removing verification is only one option and needs evidence and authority. Gains can also come from better service or fewer errors without removing a control.

The shape that produces a result

What a pilot has to establish before it can return anything Fig. 02
  1. 01 A named process One workflow, with a measured baseline before anything changes.
  2. 02 An operating change What changes in the workflow and who can authorise it.
  3. 03 An acceptance rule The outcome, quality and guardrails required for the proposed use.
  4. 04 The measurement Same units as the baseline, on the same population, after the change.

Name the operating change, whether task allocation, improved outcomes or reduced avoidable work. Keep necessary checks and measure their effort. The owner should approve changes in authority or control rather than assume financial value requires deleting a step.

Where an existing-process baseline is missing, identify what can still be reconstructed from reliable operating records. If the comparison remains uncertain, collect it prospectively before expanding the claim. Record the uncertainty and the observation needed to resolve it, rather than substituting positive feedback for a measured improvement.

Include the people who receive the changed work. A drafting improvement may create an approval queue, while a routing improvement may reduce repeated contacts. Record those handoffs and measure their effort so the pilot’s economic claim reaches the complete workflow rather than stopping at the model output.

What to require of the next pilot

Agree what result would stop, narrow or redesign the pilot. Record incomplete findings, including cases where the observation period or comparison cannot support expansion. An extension needs a defined question and observation plan.

What this does not tell you

The headline figure is one study, with one definition of measurable impact, over one sample of deployments in one period. It is not a law, and treating it as one produces the mirror-image error of the hype it corrects. Plenty of organisations have deployed systems that return value and were never in that sample.

Pilots can reduce uncertainty about the model and workflow. Choose questions relevant to the local task and population, since published evaluations may not cover them. Compare with the incumbent process and include operating effort.

The investment owner should establish what changes if the pilot succeeds, how the change will be measured and who can authorise it. Resolve those questions before expanding exposure or treating the pilot as evidence for a financial commitment.

Filed under · Method · Adoption · Method · Operating model Inference Institute · 02 Oct 2026 (updated)

Related engagement

The decision behind this article

A clear design your team or chosen delivery partner can build from.

Explore AI Architecture →

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.