Request a scoping call Contact
← Research

Your data records a process, not the world

Operational history reflects selection, recording and past policy as well as outcomes. Explain how the dataset was produced, then test whether its labels and population support the decision the proposed model must make.

Data / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

A long operational history can be valuable training material. Its volume and age do not establish that the labels match the intended prediction or that the cases represent future use. Start by identifying what was observed and which process produced the record.

Some labels record an eventual outcome. Others record an assessor’s decision, a policy rule or a proxy for the outcome of interest. Distinguish them. A model can learn patterns beyond a historical decision rule, but its evaluation is limited by the populations and evidence actually recorded.

Document the data-generating process before relying on its labels. Selection, missing outcomes and policy changes can affect both training and evaluation. These are statistical and governance questions with practical consequences for the supported use.

The three ways the record differs from the world

What the data says, and what it is usually taken to say Fig. 01
What the column contains What it is read as
The outcome for cases that were approved The outcome for all cases, including the ones that were declined
What an assessor recorded, in the fields available What was true about the case
Decisions under the policy in force that year A stable relationship between features and outcome
Cases that reached the organisation at all The population the system will be applied to

Selective labels arise when outcomes are observed only for cases admitted by an earlier decision. For example, repayment outcomes may be missing for declined loan applicants. That creates a gap when assessing a candidate on the full applicant population. It does not prove the model can only copy the old rule, but it limits what the observed outcomes alone can establish.

A change of intake channel can change the population. A model evaluated on one channel may behave differently on another. Measure coverage and performance for the intended channels rather than treating every observed difference as unexplained model drift.

The feedback loop that closes afterwards

How a deployed model becomes its own training data Fig. 02
  1. 01 Model scores a case Trained on historical decisions.
  2. 02 The score shapes the decision Directly, or through what a reviewer looks at first.
  3. 03 The decision is recorded As an outcome, indistinguishable from an unassisted one.
  4. 04 The next model trains on it Assess the influence of predecessor-assisted decisions on later labels.

Model-assisted decisions can influence later records and the labels used for retraining. Agreement with those records may then reflect part of that influence rather than independent improvement. Record when the model participated and design the evaluation to distinguish outcomes from its effect on the decision process.

An appropriately designed comparison can preserve evidence about the model’s contribution. Options may include a controlled model-free route, prospective observation or reviewed outcomes independent of the earlier recommendation. Their suitability depends on ethics, law and the workflow. No single comparison design is universally required or sufficient.

What to establish before training on operational history

A dataset can span several policy periods under one schema. Record the changes and test their effect on the intended claim. A policy-period feature may help some tasks but can leak unavailable information or preserve a rule the organisation intends to change. Evaluate it rather than adding it automatically.

Inspect why fields are completed or left blank. Optional notes may be more common in unusual cases or under a particular assessor’s practice. Their presence can become a proxy with unintended consequences. Establish the local pattern and timing before deciding whether the field is suitable.

Why this belongs in a governance conversation

Selection and historical policy can affect groups differently. Removing protected characteristics does not alone remove those effects because other fields and the label process may preserve related information. Evaluate outcomes and relevant slices with the required expertise, rather than relying on feature deletion as a complete fairness control.

This is why the useful artefact is not a fairness metric computed at the end. It is a written account of how the data came to exist, produced before the modelling starts, saying what was observed, what was not, and what changed. That document is what makes an impact assessment answerable, and it is the part of the work that regulators, auditors and courts will find most legible — the UK Information Commissioner’s guidance on AI and data protection puts the same emphasis on being able to explain the provenance of what a system learned from.

What this does not tell you

These limitations do not automatically make the data unusable. They determine which claims need qualification and which uses require further evidence. Compare alternatives, including a purpose-collected set or a simpler process, against their own limits and costs.

It also does not offer a statistical fix. Techniques exist for selective labels and for feedback loops, and they help, and none of them recovers information that was never recorded. The honest position is that some questions cannot be answered from the data available, and saying so is more valuable than a model that answers them anyway.

The data owner should retain an account of sources, selection, label meaning, policy periods and prior model involvement. Use it to define the supported population and unresolved evidence gaps. Approve training and evaluation against that boundary, and revisit it when the process changes.

Filed under · Data · Data · Bias · Method Inference Institute · 02 Oct 2026 (updated)

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.