Start a conversation Contact
← Research

If a model cannot cite you, you are not in the answer

A growing share of the questions your buyers ask are answered by a system that reads a handful of sources and summarises them. What gets read is decided by properties of your published material that are within your control, and mostly are not being controlled.

The question a prospective client used to type into a search engine now frequently goes to an assistant, and what comes back is not a list of places to look. It is an answer, assembled from a small number of sources the system retrieved, with citations attached.

That changes the unit of competition. It is no longer a ranked position on a results page. It is whether your material was among the handful of documents a retrieval step pulled, and whether it was written in a way that let a model attribute a specific claim to it.

The claim: being citable is a property of how material is written and published, and it is decided by the same things that make material useful to a human reader who is in a hurry.

What a retrieval-based answer engine is doing

Strip away the branding and every one of these systems performs the same sequence: interpret the question, retrieve candidate material, select passages, compose an answer, attach citations. Each stage discards most of what came before it, and the discarding is the part worth designing for.

Where a page is eliminated on its way to being cited Fig. 01
  1. Stage 01 Reachable Fetchable, rendered without script, and not blocked to the crawler that feeds the system.
  2. Stage 02 Retrieved A passage matches the question closely enough to be a candidate. Specificity beats breadth here.
  3. Stage 03 Selected The passage answers the question on its own, without the paragraph before it.
  4. Stage 04 Attributed The claim is stated plainly enough that a citation can be attached to it without hedging.

Stage three is where most corporate material dies. A page written as a narrative — context, then background, then eventually the point — has no passage that answers anything on its own. A page whose sections each state a claim and then support it has several.

Stage four is where hedging is punished. “Organisations may in some circumstances wish to consider” cannot be cited, because there is no proposition in it. A sentence that commits to something can be attributed to you, which is the whole mechanism.

What is actually within your control

The third and fourth lines are also the two that make writing more trustworthy to a person, which is not a coincidence. The systems are trained on and evaluated against human judgements of usefulness, and the shortcut for authors is that there is no separate craft here — material written to be genuinely useful is material that gets retrieved.

The sixth line is the one organisations most often fail without noticing. A description of your practice that differs between the homepage, the services pages and the technical documentation gives a retrieval system three candidate answers and no way to choose. Consistency is a machine-readability property as well as a brand one.

The machine-readable surfaces

There is a small amount of publishing infrastructure worth having, and it is genuinely small.

A crawler policy that is a decision rather than a default. Blocking every AI crawler removes you from the answers as effectively as writing nothing, and allowing everything may not be what a business with licensed content wants. It is a commercial choice and it should be made deliberately in robots.txt.

A structured description of the site for models. The llms.txt convention proposes a single markdown file — a title, a short summary in a blockquote, some prose, then curated link sections — that tells a model what a site contains and where the important material is, without it having to fetch and strip a dozen pages of markup. This site publishes one at /llms.txt, generated from the same registries the pages render from so that it cannot drift, alongside a full-text file at /llms-full.txt carrying every article.

Structured data on the pages themselves, so that an article’s author, date and subject are stated rather than inferred. This is ordinary schema.org markup and it has been good practice for a decade.

What this does not tell you

Nobody can promise placement in a generated answer, and any supplier offering it is selling a guess about systems whose retrieval and ranking behaviour is unpublished, changes without notice, and differs between products. The published academic auditing of these systems finds measurable biases in what they cite, which is a reason to be sceptical of anyone claiming a method.

Measurement is also immature. Attribution in an assistant’s answer is not reported to you the way a click is, the sampling approaches available are noisy, and an organisation that builds a target around a number it cannot verify has built an incentive to game something it cannot see.

What is defensible is the input side. Write material that states claims, sources its figures, stands up passage by passage and says the same thing everywhere. That is worth doing on its own merits, it is what makes a page citable, and it is the only part of this anybody actually controls.

The person who should act is whoever owns the published material. Take the last thing your organisation published and ask whether any single paragraph in it, read alone, answers a question a buyer would type. If none does, the page is not in the answer — and the fix is editorial rather than technical.

Filed under · Method · Retrieval · Method · Publishing Inference Institute · 14 Aug 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.