Start a conversation Contact
← Research

Time saved at a desk is not capacity released

Individual AI users can spend less time on a task without changing the organisation's workload or output. Measure where the released time goes before treating a productivity estimate as capacity.

A business unit trials an office assistant. Staff report shorter email sessions and faster drafting. The sponsor multiplies the saved minutes by the number of employees and calls the result new capacity. Finance asks which queue became shorter, which service improved or which work can now be done without another hire. The trial has no answer.

Time saved by one person is a useful observation. Capacity released by a system of work is a different claim. To make the second one, the buyer must show how the organisation redirected time, changed throughput or reduced a bottleneck without creating new review work elsewhere.

What the field experiment measured

In Shifting Work Patterns with Generative AI, researchers randomly assigned access to a generative AI tool across 66 firms and 7,137 knowledge workers. Among treated workers who used it in the latter half of the six-month experiment, they found less time spent on email and less work outside regular hours. They did not detect a change in the quantity or composition of workers’ tasks from individual-level provision.

That is a precise result, not a verdict against the tool. Reducing work outside regular hours may itself matter. A worker may use time for better attention, breaks or quality that the task categories did not capture. What the study does not support is converting a personal time saving directly into an organisation-wide output figure. Its setting also does not establish the result for every role, tool or firm.

The existing article Most pilots return nothing asks whether a pilot has a decision and a credible baseline. This article focuses on the next measurement boundary: once an individual benefit is observed, what evidence shows it became organisational capacity?

Instrument the workflow, not just the user

Select one work queue before the rollout. Record arrivals, completions, backlog, rework and quality under the incumbent process. Note which roles are actually constrained. An assistant that saves an analyst twenty minutes may leave throughput unchanged if approvals remain the bottleneck. It may also shift work to a reviewer, whose extra checking is invisible in the analyst’s time survey.

Specify where the time should go. If the intention is faster case resolution, measure cases completed at consistent quality. If the intention is fewer after-hours emails, measure that and say it is a working-pattern outcome. If the intention is to support new demand, record the work accepted that would otherwise have been delayed. Different outcomes require different evidence.

A capacity claim needs a workflow comparison Fig. 01
Baseline
The existing queue's throughput, backlog, review time and quality before access changes.
Outcome
Cases completed or waiting time at comparable demand and service quality.
Guardrails
Rework, reviewer load, customer complaints and work outside regular hours.
Decision rule
Expand only where the measured workflow improves without shifting unacceptable work or harm to another team.

Keep a comparison that survives the enthusiasm of the launch. A staged rollout, matched team or stable control group can help separate tool effects from a quiet month, a staffing change or revised policy. Record adoption, but do not divide the final outcome only by active users: choosing to use the tool may itself reflect easier cases or different staff. The decision owner should see both the effect of offering the tool and how it was used.

Suppose the assistant cuts drafting time for claims analysts while senior reviewers still approve every case. If each reviewer already has a full queue, cases accumulate at that review step. The sponsor can measure the analysts’ experience as a benefit, but cannot count those minutes as additional settled claims. To make a capacity claim, the firm would need to change the review process or its staffing, then observe more completed cases at comparable quality and demand.

This is why the cost side belongs in the same measurement. Include licences, integration, support, extra review and the time staff spend correcting outputs. If the new workflow moves ten minutes from an analyst to a reviewer, the local gain may be real while the service gain is smaller or absent. If it reduces late work without raising throughput, call that a working-pattern improvement. It may be worth paying for, but it should be valued as such.

Set a review date before launch. At that point, the owner can choose to redesign the bottleneck, continue for a documented employee benefit or stop the rollout. Leaving the trial running because people report that it feels faster turns a measurement into a habit, while the promised capacity remains untested.

What this does not tell you

The field experiment does not prove that AI cannot create capacity. It shows a case in which detectable task-level reallocation did not follow individual access during the observation period. A queue with a clear bottleneck and an explicit redesign may respond differently. The measures above also require care: faster completions can mask lower quality or deferred work.

The COO or service owner should decide which benefit is worth buying before the next rollout. If the desired result is a better day for employees, record that honestly. If the business case promises capacity, follow the work until the extra capacity appears in an operating measure someone owns.

Filed under · Method · Productivity · Generative AI · Measurement Inference Institute · 25 Sept 2026

Related engagement

The decision behind this article

Define the problem, requirements, feasibility, options, costs and delivery approach.

Explore AI Discovery & Scoping →

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.