The output cannot tell you whether the mechanism was ever there
The Federal Trade Commission finalised orders on 27 August 2026 against three firms that sold an advertising product on a capability the agency says it never had. No customer could have found that by examining what they received, and that is the ordinary case in AI procurement rather than the strange one.
The pilot review goes well. The supplier’s system has been running against a slice of the business for six weeks, a panel has scored a sample of its outputs against what the old process produced, and the scores are better. Nobody in the room can describe what the system is actually doing, and nobody treats that as a problem, because the results are on the table and the results are what was bought.
The results are not what was bought. What was bought was a mechanism — a model of a particular kind, running over data of a particular kind, reaching the result by a route that was described in the proposal and priced against that description. The results are evidence that something produced them. They are not evidence about what.
The claim: where a supplier’s proposition is a claim about how a result is produced rather than about the result itself, inspecting the result cannot test it, and the buyer has to specify a test that fails when the claimed mechanism is absent. Such a test is cheap to design before signature and close to impossible to construct afterwards, because after signature the only artefact anybody still has is outputs.
A worked instance, decided by a regulator
On 27 August 2026 the United States Federal Trade Commission finalised its orders against CMG Media Corporation, MindSift LLC and 1010 Digital Works LLC over a marketing product sold under the name “Active Listening”. It was sold to small advertisers on the proposition that it picked out relevant conversations overheard by consumers’ smart devices and targeted advertising against them.
The Commission’s announcement of the settlement in May 2026 states that the service “did not, in fact, listen in on consumers’ conversations or use voice data at all”, and that what the companies provided “consisted of reselling — at a significant markup — email lists obtained from other data brokers”. The three firms are to pay USD 930 000 between them, of which USD 880 000 falls on CMG, for redress to affected customers.
Set the deception aside. It is the Commission’s finding on a specific record, and it binds three named respondents. The part worth taking from it is narrower and applies to buyers who will never meet a bad actor: what could a careful customer have observed. They received advertising placements and campaign reporting. A targeting system driven by captured speech and a bought email list both produce advertising placements and campaign reporting. Performance was whatever it was, and there was no parallel campaign run the other way to hold it against. The deliverable was consistent with the claim and equally consistent with the absence of the claim.
Why the output is silent
Three conditions make a deliverable uninformative about the route that produced it, and between them they cover a large part of what enterprises are currently buying.
The first is that the output is a plausible artefact rather than a checkable one. A summary, a ranking, a match, a risk score, a paragraph of drafted text — these are judged by whether they look right, and several very different processes produce things that look right. The second is that the buyer holds no ground truth the supplier has not already seen, so every number quoted in the review was computed on material the supplier could tune against. The third is that the run-to-run variance of the process is wider than the difference the mechanism would make, which means a sample of a few hundred cases cannot separate the two even in principle.
| What the proposal claims | What the output shows |
|---|---|
| The system reads your own documents before answering. | An answer consistent with those documents — and equally consistent with a general model that never opened one. |
| A model scores each case and the score drives the routing. | Cases routed, in a distribution that four rules would also produce. |
| The model was tuned on material from your sector. | Sector vocabulary, which an instruction in the prompt also produces. |
| The training data was obtained on a basis the supplier holds. | Nothing at all. Provenance leaves no trace in the artefact it produced. |
The fourth row is the one that carries commercial exposure rather than disappointment. A claim about where data came from is invisible in every output the system will ever emit, and it is the claim a buyer inherits most directly.
Designing a test that can fail
If the claimed mechanism were quietly replaced with something cruder, would anything you receive change?
- Output testing is a real test. Write it into acceptance and run it on cases the supplier has never seen. Establish this rather than assume it — the assumption is usually where the error is.
- Test the mechanism directly, through its inputs, its operating traces and a held-back set built to require it. No amount of sampling the deliverable will reach the question.
The held-back set does the most work, and it has an unusually good public worked example. NIST’s Artificial Intelligence Technology Evaluation programme, announced in July 2026, is built on the observation that a dataset stops measuring anything the moment the party being measured can see it. Its testbed keeps evaluation data sequestered, which NIST describes as mitigating “the risk of train/test data contamination”. Data providers contribute a task and a dataset that is not published. Model providers submit models and receive measurements on data they never hold.
The enterprise version of that is very much smaller — some tens of cases, kept by the buyer, never sent to the supplier, refreshed whenever the supplier changes something material. It is also the only part of a pilot evaluation that still means anything six months later.
The same test, pointed inwards
Reading this as a supplier-fraud problem is the way to miss most of it. The same asymmetry runs inside an organisation, where a mechanism is rarely misrepresented and quite often replaced.
A system ships with a model in the loop. A latency ceiling appears, or a cost ceiling, or a bad fortnight of outputs, and a rule is added in front of the model to handle the common cases. Nobody decides to stop using the model. The routing simply narrows until the model sees the residue. Reporting does not change, because reporting was built on outputs, and the outputs still arrive. Two years later a governance review asks what makes the decision, and the accurate answer is that nobody has checked since the first month.
That is a failure of instrumentation rather than of honesty, and it has the same remedy as the procurement case. Record which path produced each output, and count the paths. A system that cannot say which of its own routes answered a given case cannot be evidenced to anybody — not to a regulator, not to a customer, and not to the team that inherits it.
What this does not tell you
This is one enforcement action, brought by one agency, against three named respondents, on a record they agreed to settle. It is not evidence about how often AI capability claims are misdescribed, and no figure in this article should be read as measuring that. It is also not legal advice. Whether a capability claim in your own contract is a warranty, a representation or marketing is a question for your counsel, decided on the wording in front of them.
We do not certify a supplier’s claims as true, and no review we run establishes that a system works the way its documentation says. An assessment whose only available conclusion is “as described” would not be an assessment. What a review does is narrower and more useful: decide which material claims are testable, design the test that would fail if a claim were wrong, and write down the claims that cannot be tested at all — so that those ones, at least, are bought knowingly.
The person who decides differently is whoever is drafting the acceptance criteria for the next AI service, in the fortnight before a shortlist becomes a contract. One line belongs in that document: the test that fails if the mechanism is not there. Where nobody in the room can write that line, the thing being purchased is not a claim about the system. It is a claim about the supplier, and it is being bought on trust — which is a defensible thing to buy, as long as everybody signing knows that is what is on the invoice.