Start a conversation Contact
← Research

Give outcome labels time to mature

A customer who has not bought yet is different from one whose purchase window has closed. Preserve that distinction in training data and performance reports before an AI system learns to divert effort away from slower conversions.

A commercial director reviews yesterday’s leads. The new scoring model sent promising enquiries to the sales team, but few have become orders. The dashboard compares them with last month’s completed pipeline. A retraining job converts every empty purchase field into a negative label. By the next review, the system has learnt that the customers who take longer to decide are less worth pursuing.

That can spend the business case before anybody notices a modelling problem. Sales attention moves towards quick decisions, a slower product loses coverage, and the director cuts a channel whose customers are still considering their purchase. The immediate decision is whether those records are ready to train on or support a budget change.

Give the data owner an explicit rule for when an outcome becomes observable. Preserve pending outcomes, define the purchase horizon and distinguish the customer’s delay from the reporting pipeline’s delay. A fresh extract can contain unfinished evidence.

The label needs an observation window

Start with the actual target. Predicting whether a lead will buy within the agreed sales window differs from predicting whether it will buy by tomorrow. If the commercial team wants the former while the training job labels the latter, a technically successful model will learn the wrong business question.

Google’s conversion-delay guidance explains why recent advertising performance can look weaker than older performance: spend is already reported while some conversions have yet to happen. This is product documentation about Google Ads, not evidence that a particular enterprise’s weak results will recover. It establishes a reason to inspect the age of the observations before interpreting the comparison.

The training consequence is studied directly in Yang and colleagues’ research on delayed feedback, published in the AAAI proceedings. The paper describes the trade-off between waiting for more accurate conversion labels and using fresh data. It proposes elapsed-time sampling and importance weighting, evaluated with Criteo and Taobao data. It also examines bias from estimated weights. The result supports treating label delay as a modelling choice. It does not establish a waiting period or an algorithm for another firm.

The earlier article on backtest leakage asks whether the inputs existed when the prediction was made. Here, the inputs can be perfectly timed. The defect sits in the later outcome: the evaluation has stopped watching before the event it claims to measure could reasonably finish. Both boundaries need to be correct.

Keep pending outcomes visible

For a bounded purchase target, give each lead a prediction time, the target window’s end, an outcome event time and the time the outcome became available to the data pipeline. Keep the original score and model version beside them. A closed window with no recorded purchase supports a negative label only when the relevant reporting feeds are sufficiently complete.

Google separately documents data freshness and reporting adjustments. Processing schedules and later revisions affect what appears in a report. This is a different delay from a customer taking time to purchase. The provider’s schedules describe its own service. The organisation must establish the delays in its CRM, order feed and reconciliation process independently.

In a local training dataset, retain an explicit pending state. Do not silently turn a missing outcome into false. Nor should the team train immediately on successful recent leads while excluding only their unresolved peers. That would make the newest cohort disproportionately positive. A simple baseline is to admit whole cohorts after their common observation window and reporting allowance have elapsed, keeping unresolved feed failures visible.

Waiting has a cost. Mature cohorts describe an older market, and a long sales cycle can make that evidence slow to arrive. Where faster learning is valuable, compare a model that accounts for delay with the mature-cohort baseline. Record its assumptions, validate predictions against subsequently observed outcomes and keep estimates visibly separate from completed observations. Complexity earns its place through that comparison.

Rehearse a late purchase before retraining

Use a constructed record before touching production. Suppose a lead is scored on Monday under a target of purchase by Friday evening. The customer orders on Thursday, and the order reaches the analytical store on Saturday. Those dates are an illustrative test fixture, not an observed sales result.

A Tuesday extract should mark the outcome as pending. A Friday extract also cannot establish a negative merely because the order field is empty. Once the Saturday feed arrives, the label is positive because the purchase occurred inside the target window. Its later arrival must not move the purchase outside that window or rewrite the inputs the model saw on Monday.

Repeat the fixture with no purchase and a feed confirmed complete through the window’s end. That record can become negative. Repeat it with an order occurring after the target window: retain the late order as a business event, but keep the original bounded target negative. Finally, break the feed. The result should be an unresolved data condition, rather than a confident failure to convert.

Then replay real historical cohorts at successive observation ages. Preserve what each extract would have known, and compare its provisional labels with the later reconciled outcome. Segment by product, channel and customer journey. Averages can conceal a slow segment that contributes valuable orders. Keep the cohort denominator fixed so that extra observation time, rather than a changed population, explains the revisions.

The acceptance test for a conversion-label pipeline Fig. 01
Baseline
The current extract and a reference built from whole cohorts after the target window and verified reporting allowance.
Outcome
Stable conversion labels and calibrated predictions when the same cohorts are revisited after reconciliation.
Guardrails
Unresolved feed failures, disproportionate revisions in slower segments, lost sales coverage and estimates presented as completed outcomes.
Decision rule
Hold retraining or budget reallocation when the conclusion depends on immature labels. Compare any delay correction with subsequently reconciled outcomes before widening its use.

Give the commercial owner a report with separate views of completed cohorts, pending exposure and estimated future outcomes. Finance can then see which conclusions rest on observed purchases and which depend on a forecasting model. A delay estimate may support a bounded operating decision, but the later reconciliation must remain visible when it proves wrong.

Decide what can wait

Maturity does not repair missing tracking, disputed attribution or a badly chosen target. It also does not prove that the sales model caused a purchase. That requires a comparison capable of measuring incremental effect. The cited paper studies advertising conversion prediction, and the vendor guidance describes reporting behaviour. Neither establishes a return for this proposed workflow.

Waiting for outcomes must not suspend immediate controls. Broken feeds, inappropriate offers or excessive spending can justify intervention before conversion labels mature. Separate those operating decisions from claims that the model’s predictive performance has deteriorated.

Before the next retraining run, the data owner and commercial director should agree when a lead becomes eligible evidence, how later events revise it and which decisions can use provisional estimates. Otherwise the organisation risks teaching its model that patience is failure, then paying for that lesson in lost sales coverage.

Filed under · Data · Delayed feedback · Training data · Conversion measurement Inference Institute · 28 Sept 2026

Related engagement

The decision behind this article

Independent assessment of an existing AI architecture.

Explore AI Architecture Review →

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.