Request a scoping call Contact
← Research

Do not give every demand forecast a human override

Human input improves some demand forecasts and weakens others. Define what planners may contribute, preserve the original forecast and test the intervention policy against inventory and service outcomes.

Data / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

A forecasting system may allow planners to adjust every product and store prediction before replenishment. That permission is an operating policy. Its value depends on whether planners hold useful additional information and whether the adjustment improves the decisions made from the forecast.

Routine editing creates another forecasting process and consumes planning time. Preserve the model output and the reason for each adjustment so the organisation can distinguish new information from a reaction to an unexpected result. Assess the effect on availability, inventory and the planning queue as well as forecast error.

Treat the intervention policy as a design choice to test. Give planners a way to supply information the model lacks, define when that information may change the forecast and evaluate the combined result. The objective is better replenishment and working-capital decisions at an acceptable review cost.

An override is a second forecast

The distinction matters because a forecast is an input to several decisions. A buyer uses it to place an order, a distribution team uses it to allocate stock and a finance team uses it to plan cash. An unstructured adjustment changes all three without recording which fact justified the change.

Revilla and colleagues tested this problem in a retail field experiment. They randomly assigned 1,888 stock-keeping units to full automation, adjustable automation or augmentation for a 50-week forecasting process. Revilla and colleagues, 2023

The result depended on the context. For innovative products and short forecast horizons, forecasts were more accurate when the system operated without human intervention. For established products and long horizons, more human intervention reduced forecast error. The authors conclude that the useful level of intervention depends on demand uncertainty and the time until the prediction will be used. A blanket override policy therefore applies the wrong kind of attention to at least some of the range.

The same pattern appears when the intervention method changes. Brau and colleagues combined a controlled experiment with a retail field study of more than three million weekly forecasts. Their field dataset contained 219,363 product-store observations. Direct judgmental adjustment was less accurate than methods that let planners identify an event while the model estimated how much that event should change the forecast. Brau and colleagues, 2023

The useful design is therefore not “AI or planner”. It is a boundary between a model’s repeatable pattern detection and a planner’s private information. The planner might know that a store is closing, a promotion has changed or a competitor has run out of stock. The system can estimate the effect of that fact from comparable history. Asking the person to type the final quantity throws away that division of labour.

A demand-planning workflow that makes human information testable Fig. 01
  1. 01 Forecast Generate a baseline with the data and horizon the system can observe.
  2. 02 Explain Ask the planner which event or missing signal could change the baseline.
  3. 03 Weight Estimate the event effect from comparable history rather than accepting an unbounded edit.
  4. 04 Release Send the forecast to replenishment with the reason, version and accountable owner attached.

Put the exception where it can earn its place

The buyer for this decision is a chief supply chain officer, a VP of planning or the operations leader who owns service levels and inventory. Their pilot should start with a map of the current process, including every place a forecast is changed before an order or allocation is released.

Establish a baseline by horizon, product age, store and demand pattern. Record each adjustment, its supporting evidence, the time spent and the downstream action. Consistent reason codes help compare interventions, while explanatory notes can preserve details that do not fit a code. Free text remains evidence when it can be retrieved and reviewed.

Second, segment the intervention policy. New products with little history and short horizons may be a good place for the model to run without a routine human edit. Established products with a long horizon may benefit from a planner flagging a known event. The segmentation is a hypothesis to test, not a rule to copy from a paper.

Third, change what the planner is asked to do. Instead of “set the forecast”, ask “which event is missing, and when will it stop affecting demand?” The model can then learn whether that signal has a repeatable effect. If the event has no precedent, route it to an explicit exception queue with a stated owner and expiry.

Fourth, measure the business result. Forecast error is useful diagnostic evidence, but it is not the outcome the operation buys. Read it alongside stockout rate, excess or obsolete inventory, service level, markdowns, order expedites, planner minutes and the time from forecast to released order. Keep a control group under the existing process long enough to observe the same demand events.

The economic mechanism is capacity as well as accuracy. If planners review every line, adding more products or stores adds review work in proportion to the catalogue. A selective policy spends that attention on cases where private information can change the decision. The gain may appear as fewer shortages, less working capital tied up in the wrong stock, a shorter planning cycle or more coverage from the same team. The organisation should choose which of these it is buying before it chooses a model.

The control is the reason, not the edit box

Supplier demonstrations often show a planner moving a forecast up and down. That proves the interface accepts a number. It does not prove that the number will improve an order. Ask the supplier to preserve the system forecast, the planner’s reason, the evidence shown, the final released value and the later actual demand in one trace.

Ask how the system learns from an intervention. Does a repeated promotion become a feature, or does the planner keep typing the same correction each week? Can an owner see when a reason has stopped improving outcomes? Can the policy prevent a large edit when the supporting event has expired? These questions turn a useful human signal into a process the business can inspect.

The pilot also needs an unhappy path. Hold an intervention for a product whose event is cancelled, release the queue and check that the forecast is recomputed or the case is rejected. Compare a planner’s edit with the actual event outcome. An override that cannot be evaluated after the fact is a preference, not a control.

What the Inference Institute can help decide

The useful engagement begins with the replenishment workflow and its economic constraint. It maps where forecasts are consumed, separates model evidence from planner knowledge and defines the smallest intervention policy that can be tested.

The result is a pilot specification a supply chain leader can fund and review. It names the segments, reason codes, human boundary, control group, observation window, operational measures and stopping rule. It also makes clear whether the intended gain is availability, working-capital efficiency, planning capacity or cycle time.

What this does not tell you

The Revilla study concerns one retail forecasting process and one set of products. Its 50-week experiment measures forecast error, not a universal change in profit, stockouts or service levels. The Brau field study ran during the early COVID-19 period and evaluates forecast integration methods rather than a modern generative AI product. Neither paper establishes a threshold, vendor or intervention rule for your operation.

The recommendation is also not to remove expertise from planning. A planner may hold the only timely evidence about a disruption. The recommendation is to record that evidence as an event the system and the business can evaluate, rather than treating an unexplained edit as an approval right.

Before enabling overrides, the demand-planning owner should explain what additional information a planner can supply, how the system will use it and which later measure will show its value. If these questions remain unanswered, begin with a bounded comparison rather than a universal edit permission.

Filed under · Data · Supply chain · Demand forecasting · Inventory Inference Institute · 02 Oct 2026 (updated)

Related engagement

The decision behind this article

A clear design your team or chosen delivery partner can build from.

Explore AI Architecture →

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.