Request a scoping call Contact
← Research

The output cannot tell you whether the mechanism was ever there

A plausible output does not establish how a supplier produced it. Where a mechanism matters to the purchase, define direct evidence and a test that could expose its absence before accepting the service.

Method / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

A pilot can show that a supplier’s outputs improve on a reference process while leaving another question unanswered: does the service use the mechanism the contract describes? Decide which mechanism claims matter to the purchase and how the buyer will test them.

Some purchases are rightly specified by outcomes. Others also depend on a model, data source, processing route or rights basis. For those, an output is evidence of what was delivered but may be consistent with several ways of producing it. The acceptance criteria should distinguish the two types of claim.

Specify evidence that can challenge each material mechanism claim. A discriminating test, an inspectable operating record or a contractual representation may answer different parts of the question. Agree the required access before signature, when it can still influence the contract.

A worked instance, decided by a regulator

On 27 August 2026 the United States Federal Trade Commission finalised its orders against CMG Media Corporation, MindSift LLC and 1010 Digital Works LLC over a marketing product sold under the name “Active Listening”. It was sold to small advertisers on the proposition that it picked out relevant conversations overheard by consumers’ smart devices and targeted advertising against them.

The Commission’s May 2026 settlement announcement describes allegations that the service did not use captured conversations or voice data and instead resold brokered email lists. The agreed payments totalled USD 930,000, including USD 880,000 from CMG, for customer redress. These are allegations resolved through the named settlements, rather than a measurement of AI suppliers generally.

The case illustrates a narrower procurement problem. Advertising placements and campaign reports can be produced through different targeting inputs. Examining those deliverables alone may not establish that the claimed input was used. A buyer needs evidence that distinguishes the relevant routes rather than assuming consistency with a claim proves it.

Why the output is silent

Several conditions can make an output insufficient to identify its production route. Assess which apply to the specific proposition rather than assuming every AI purchase shares them.

Different methods can produce equally plausible summaries, rankings or drafted text. Evaluation data disclosed to the supplier may permit tuning that weakens a claimed independent comparison. Variability can also make a proposed test too weak to distinguish the relevant difference. These are reasons to design a discriminating test, not proof that a fixed sample size can never support one.

Illustrative propositions and the limits of output inspection Fig. 01
What the proposal claims What the output shows
The system reads your own documents before answering. An answer consistent with those documents — and equally consistent with a general model that never opened one.
A model scores each case and the score drives the routing. Cases routed, in a distribution that four rules would also produce.
The model was tuned on material from your sector. Sector vocabulary, which an instruction in the prompt also produces.
The training data was obtained on a basis the supplier holds. The output alone may not establish the acquisition route or rights basis.

Data provenance often requires records and contractual evidence beyond the output. Assess the supplier’s acquisition and rights claims directly. The buyer’s exposure depends on its role, use and contract, so do not infer a universal liability from the absence of visible provenance.

Designing a test that can fail

The question to settle before the acceptance criteria are written Fig. 02

If the claimed mechanism were quietly replaced with something cruder, would anything you receive change?

  • Something you receive changes measurably Output testing is a real test. Write it into acceptance and run it on cases the supplier has never seen. Establish this rather than assume it — the assumption is usually where the error is.
  • Nothing you receive changes Test the mechanism directly, through its inputs, its operating traces and a held-back set built to require it. If the routes are observationally equivalent, more samples of that deliverable will not distinguish them.

NIST’s Artificial Intelligence Technology Evaluation announcement and testbed describe sequestered evaluation data intended to reduce training and test contamination. That is a useful independence measure. Data disclosure does not automatically invalidate every evaluation, and a protected set still needs a suitable task and test design.

A buyer can retain a protected set with cases designed to require the claimed capability. Choose the sample and conditions against the difference the test must detect. If the supplier runs a hosted service, the cases may need to reach that service during evaluation: protect them from prior tuning and define how they may be retained or reused.

The same test, pointed inwards

The same review applies to internal changes. A service can replace or bypass a mechanism for legitimate cost, latency or reliability reasons. Its records should make that change visible so the owner can reassess the claims that depended on the original route.

Consider a constructed routing example. A service initially sends all cases to a model, then adds rules for selected requests. The resulting outputs may remain acceptable while the route changes. Record which path answered each case and update the design decision when that path affects a material claim.

That is a failure of instrumentation rather than of honesty, and it has the same remedy as the procurement case. Record which path produced each output, and count the paths. Without a route record, the organisation may be unable to substantiate claims about how a particular case was handled. State that evidence limit accurately to reviewers, customers and the team that inherits the service.

What this does not tell you

This is one enforcement action, brought by one agency, against three named respondents, on a record they agreed to settle. It is not evidence about how often AI capability claims are misdescribed, and no figure in this article should be read as measuring that. It is also not legal advice. Whether a capability claim in your own contract is a warranty, a representation or marketing is a question for your counsel, decided on the wording in front of them.

An independent assessment should distinguish demonstrated claims, qualified evidence and claims that remain untestable with the available access. It can establish specific properties within a defined test scope. It cannot certify every supplier statement or turn incomplete evidence into a general assertion that the service operates as described.

The procurement owner should include a test or evidence requirement for each material mechanism in the acceptance criteria. Where that evidence cannot be obtained, record the dependence on a supplier representation and assess its contractual treatment. The buyer can then accept that reliance knowingly or change the purchase.

Filed under · Method · Procurement · Evaluation · Supplier risk Inference Institute · 02 Oct 2026 (updated)

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.