Start a conversation Contact
← Research

Do not give every demand forecast a human override

STATE THE CLAIM IN ONE OR TWO SENTENCES. Between 60 and 320 characters, and it has to read alone — it is the meta description and the RSS summary, so people will see it who never see the page.

A supply chain director has bought an AI demand forecasting system. The first review meeting is reassuring: the dashboard produces a forecast for every product and store, and planners can adjust any line before the replenishment run. The team keeps the permission because it feels like control.

The permission creates a second forecasting process. Planners spend time rebuilding a number the model already produced, and the business cannot tell whether an adjustment reflects information the model did not have or a reaction to a surprising chart. A lower forecast error on a spreadsheet can still leave the operation with more stock, more stockouts or a longer planning queue.

Do not give every demand forecast a human override. Give planners a defined way to contribute information the model cannot see, then decide where that information should change the forecast and where it should be left alone. The commercial question is not whether a person can edit a number. It is whether the combination produces better availability and working-capital decisions at a manageable review cost.

An override is a second forecast

The distinction matters because a forecast is an input to several decisions. A buyer uses it to place an order, a distribution team uses it to allocate stock and a finance team uses it to plan cash. An unstructured adjustment changes all three without recording which fact justified the change.

Revilla and colleagues tested this problem in a retail field experiment. They randomly assigned 1,888 stock-keeping units to full automation, adjustable automation or augmentation for a 50-week forecasting process. Revilla and colleagues, 2023

The result depended on the context. For innovative products and short forecast horizons, forecasts were more accurate when the system operated without human intervention. For established products and long horizons, more human intervention reduced forecast error. The authors conclude that the useful level of intervention depends on demand uncertainty and the time until the prediction will be used. A blanket override policy therefore applies the wrong kind of attention to at least some of the range.

The same pattern appears when the intervention method changes. Brau and colleagues combined a controlled experiment with a retail field study of more than three million weekly forecasts. Their field dataset contained 219,363 product-store observations. Direct judgmental adjustment was less accurate than methods that let planners identify an event while the model estimated how much that event should change the forecast. Brau and colleagues, 2023

The useful design is therefore not “AI or planner”. It is a boundary between a model’s repeatable pattern detection and a planner’s private information. The planner might know that a store is closing, a promotion has changed or a competitor has run out of stock. The system can estimate the effect of that fact from comparable history. Asking the person to type the final quantity throws away that division of labour.

A demand-planning workflow that makes human information testable Fig. 01
  1. 01 Forecast Generate a baseline with the data and horizon the system can observe.
  2. 02 Explain Ask the planner which event or missing signal could change the baseline.
  3. 03 Weight Estimate the event effect from comparable history rather than accepting an unbounded edit.
  4. 04 Release Send the forecast to replenishment with the reason, version and accountable owner attached.

Put the exception where it can earn its place

The buyer for this decision is a chief supply chain officer, a VP of planning or the operations leader who owns service levels and inventory. Their pilot should start with a map of the current process, including every place a forecast is changed before an order or allocation is released.

First, establish a baseline. Record forecast error by horizon, product age, store and demand pattern. Record the planner’s adjustment, the reason selected, the time spent and the downstream action. The reason must be a small, declared set such as promotion, weather, new store, supply disruption or model bias. A free-text explanation that cannot be grouped later is not an evidence trail.

Second, segment the intervention policy. New products with little history and short horizons may be a good place for the model to run without a routine human edit. Established products with a long horizon may benefit from a planner flagging a known event. The segmentation is a hypothesis to test, not a rule to copy from a paper.

Third, change what the planner is asked to do. Instead of “set the forecast”, ask “which event is missing, and when will it stop affecting demand?” The model can then learn whether that signal has a repeatable effect. If the event has no precedent, route it to an explicit exception queue with a stated owner and expiry.

Fourth, measure the business result. Forecast error is useful diagnostic evidence, but it is not the outcome the operation buys. Read it alongside stockout rate, excess or obsolete inventory, service level, markdowns, order expedites, planner minutes and the time from forecast to released order. Keep a control group under the existing process long enough to observe the same demand events.

The economic mechanism is capacity as well as accuracy. If planners review every line, adding more products or stores adds review work in proportion to the catalogue. A selective policy spends that attention on cases where private information can change the decision. The gain may appear as fewer shortages, less working capital tied up in the wrong stock, a shorter planning cycle or more coverage from the same team. The organisation should choose which of these it is buying before it chooses a model.

The control is the reason, not the edit box

Supplier demonstrations often show a planner moving a forecast up and down. That proves the interface accepts a number. It does not prove that the number will improve an order. Ask the supplier to preserve the system forecast, the planner’s reason, the evidence shown, the final released value and the later actual demand in one trace.

Ask how the system learns from an intervention. Does a repeated promotion become a feature, or does the planner keep typing the same correction each week? Can an owner see when a reason has stopped improving outcomes? Can the policy prevent a large edit when the supporting event has expired? These questions turn a useful human signal into a process the business can inspect.

The pilot also needs an unhappy path. Hold an intervention for a product whose event is cancelled, release the queue and check that the forecast is recomputed or the case is rejected. Compare a planner’s edit with the actual event outcome. An override that cannot be evaluated after the fact is a preference, not a control.

What the Inference Institute can help decide

The useful engagement begins with the replenishment workflow and its economic constraint. It maps where forecasts are consumed, separates model evidence from planner knowledge and defines the smallest intervention policy that can be tested.

The result is a pilot specification a supply chain leader can fund and review. It names the segments, reason codes, human boundary, control group, observation window, operational measures and stopping rule. It also makes clear whether the intended gain is availability, working-capital efficiency, planning capacity or cycle time.

What this does not tell you

The Revilla study concerns one retail forecasting process and one set of products. Its 50-week experiment measures forecast error, not a universal change in profit, stockouts or service levels. The Brau field study ran during the early COVID-19 period and evaluates forecast integration methods rather than a modern generative AI product. Neither paper establishes a threshold, vendor or intervention rule for your operation.

The recommendation is also not to remove expertise from planning. A planner may hold the only timely evidence about a disruption. The recommendation is to record that evidence as an event the system and the business can evaluate, rather than treating an unexplained edit as an approval right.

The operations leader who owns demand planning should ask one question before turning on overrides: what information does the person have, how will the system use it and which later business measure will show that it helped? If the answer is only that somebody prefers a different number, the permission is adding a second forecast and charging the business for both.

Filed under · Data · Supply chain · Demand forecasting · Inventory Inference Institute · 13 Sept 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.