Publish evidence that readers and retrieval systems can use
Clear, accessible and well-sourced publishing gives retrieval systems usable material. Improve those inputs for readers and automated access, while recognising that they cannot guarantee selection or citation.
Keep development and test groups separate before comparing performance. No measured results are shown.
Some buyers use assistants that retrieve source material and compose an answer with references. For a publisher, this adds a discovery route alongside conventional search. The page must be accessible to the relevant system and provide information that supports a specific answer.
Retrieval and citation depend on more than a page’s search position. A system may select a passage, omit it or attribute it differently according to its own behaviour. Publishers can improve the material available without controlling that selection.
Write specific claims with appropriate sources, dates and qualifications. Make the page readable without unnecessary client-side processing. These practices help people assess the material and provide clearer inputs to systems that retrieve it.
What a retrieval-based answer engine is doing
A retrieval-based answer system typically searches candidate material and uses selected passages when composing an answer. The exact stages differ between products, and some responses use other sources or prior model knowledge. The following sequence is a useful publishing model rather than a universal product specification.
- Stage 01 Reachable Fetchable, rendered without script, and not blocked to the crawler that feeds the system.
- Stage 02 Retrieved A passage matches the question closely enough to be a candidate. Specificity beats breadth here.
- Stage 03 Selected The passage answers the question on its own, without the paragraph before it.
- Stage 04 Attributed The claim is stated plainly enough that a citation can be attached to it without hedging.
Sections should state their subject clearly and supply enough context for the reader to interpret a passage. A narrative can still be useful, but an argument whose key claim appears only after several unrelated paragraphs may be harder to extract accurately.
Use qualifications that express a real evidence limit. Vague wording can obscure the proposition, while a precise condition makes it more trustworthy. Removing justified uncertainty to seek a citation would weaken the publication.
What is actually within your control
Sources and checkable statements improve a reader’s ability to assess the argument. They also give a retrieval system explicit evidence to work with. This does not establish that useful writing will always be retrieved or accurately cited.
Keep the account of the practice consistent across the homepage, services and documentation. Conflicting descriptions create ambiguity for both readers and automated systems. Shared registries help maintain the offer while allowing each page to explain it in context.
The machine-readable surfaces
There is a small amount of publishing infrastructure worth having, and it is genuinely small.
Set a deliberate crawler policy in robots.txt, taking account of public discovery and any licensed content. Access rules differ between crawlers and uses. Blocking one crawler does not necessarily remove material from every generated answer, and permission does not establish that indexing occurred.
A structured description of the site for models. The llms.txt convention proposes a single markdown file — a title, a short summary in a blockquote, some prose, then curated link sections — that tells a model what a site contains and where the important material is, without it having to fetch and strip a dozen pages of markup. This site publishes one at /llms.txt, generated from the same registries the pages render from so that it cannot drift, alongside a full-text file at /llms-full.txt carrying every article.
Structured data on the pages themselves, so that an article’s author, date and subject are stated rather than inferred. This is ordinary schema.org markup and it has been good practice for a decade.
What this does not tell you
Selection and citation remain outside the publisher’s control. Behaviour differs between products and can change. Treat visibility claims as hypotheses to measure with a stated method, rather than a promised outcome of an editorial or technical change.
Measurement is also immature. Attribution in an assistant’s answer is not reported to you the way a click is, the sampling approaches available are noisy, and an organisation that builds a target around a number it cannot verify has built an incentive to game something it cannot see.
Improve the inputs that the organisation controls: accessible pages, clear claims, nearby sources, publication dates and a consistent description of its work. Assess these on their usefulness and accuracy for readers as well as their machine readability.
The publishing owner can review whether each section answers a recognisable reader question and whether its evidence and limitations remain clear when excerpted. Address missing context or accessibility first. A lack of citation alone does not diagnose the cause.