Start a conversation Contact
← Research

An AI-drafted control profile describes your documents, not your controls

NIST has published worked prompts for drafting a Cybersecurity Framework profile from an organisation's own artefacts. A model given a folder of documents can only report what those documents assert, which leaves the note recording where the evidence ran thin as the part with assurance value.

There is a piece of work in every framework assessment that used to take three weeks and now takes an afternoon. A security team is asked for a current-state profile: every outcome in the framework, marked against what the organisation actually does. Someone collects the policy handbook, the standards, the risk register, the last two audit reports and a folder of interview notes, and reads all of it. Then they write a row for each outcome. The rows where they had almost nothing to go on read differently from the rows where they had plenty, because that is what happens to prose when the person writing it is unsure.

Give the same folder to a model and the table comes back complete. Every row populated, every row in the same register, every row equally calm.

The claim is this. A model working from an organisation’s own documents can only report what those documents assert, so a profile drafted that way carries no assurance value the documents did not already carry. The one genuinely new thing in the output is the record of where the evidence ran out. That record, and not the populated table, is what should be reviewed, kept and put in front of whoever signs.

What NIST actually published

On 19 August 2026 NIST released the initial public draft of SP 1353, a quick-start guide for using AI in Cybersecurity Framework analysis and reporting, open for comment until 15 October 2026. It sets out three notional use cases — reviewing governance material against the GOVERN function, drafting a current-state profile, and drafting a target-state profile — each with a worked prompt and a portfolio of fictitious source files to run it against. The announcement frames it as a practical aid to work that is genuinely laborious, which it is.

It is a more careful document than the format suggests. The prompts tell the model to use the attached material only, to be “Source-grounded and traceable. No fabrication,” and, where an outcome is not addressed in the sources, to “say so plainly”. The guide states that its examples “are not prescriptive assessment or assurance methodologies”, that “AI-generated content should always be reviewed by qualified personnel”, and that users remain “responsible for validating applicability, scope, inputs, assumptions, and outputs”. Every one of those sentences is load-bearing.

The instruction worth stopping on sits at the end of the current-state prompt. After the table, the model is told to add an “Assumptions & Evidence Gaps” note listing outcomes where the sources were silent or thin, and “any places where column H rests on documented process rather than observed practice.”

Read that as an assurance requirement rather than as a prompt-engineering detail. It is the entire argument.

Why the table cannot say more than the folder

A control statement rests on one of four things, and they are not interchangeable. What separates the bottom of that list from the top is not detail. It is whether anyone has looked.

The four things a control statement can rest on, and the two a document-fed model can reach Fig. 01
  1. Level 01 Policy text What a document says should happen.
  2. Level 02 Documented process A written procedure with a named owner.
  3. Level 03 Observed practice What somebody saw people actually do.
  4. Level 04 Tested effectiveness The control exercised against a case it should catch.

The first two levels are in the folder. The second two are not, and no amount of reading the folder produces them, because they are facts about the world rather than facts about the writing. A model asked to populate a profile from documents is therefore working at the first two levels for every row it fills, whatever the row says. This is not a failure of the model and it is not fixed by a better one. It is a property of the input set.

Interview notes are the case that looks like an exception and is not. NIST puts them in the folder, and a practitioner describing what their team does is closer to practice than a policy is. It remains self-report rather than observation, which is exactly why the guide asks for the distinction to be flagged row by row rather than treated as settled.

An assessor doing this by hand is subject to the same limit and handles it differently, mostly without noticing. Confidence leaks into their prose. A row supported by a tested control and a row supported by a paragraph in a handbook come out sounding different, and a reviewer picks up the difference without being told to look for it.

The guide lists, as one of the benefits of the approach, “applying uniform language and interpretation” across the profile, and it pairs that immediately with review and refinement by human subject matter experts. The gain in readability is real. So is the cost, and it lands on the reviewer the guide is relying on. Uniform language across rows of unequal evidential strength removes the only signal a reviewer was reading. The output does not become less accurate than the human draft. It becomes harder to audit, because the thin rows no longer look thin.

This is why the evidence-gap note is not a nicety at the end of the prompt. It is the mechanism that puts the removed signal back, in an explicit form, and it is the only part of the output that is about the evidence rather than about the claim.

What to require back

The change in review posture is the part that costs something. Reviewing an AI-drafted profile by reading it end to end is close to worthless, because the document reads well everywhere and reading well is what it was optimised for. The useful review starts from the gap register, takes the rows the model declared thin, and goes and looks — at the system, at the ticket queue, at somebody doing the thing. That is a smaller amount of work than the original drafting and a different kind of work, and it has to be resourced as such rather than assumed to have been absorbed by the time the model saved.

There is a second-order effect worth planning for. Once a profile can be redrafted in an afternoon, it will be, and the version that goes to a board or a customer will be one nobody has walked back to source. Deciding now which artefacts require a fresh evidence-gap register before they leave the building is cheaper than deciding it after one has left.

What this does not claim

None of this says the approach is wrong or that NIST has overstated it. The draft is explicit that its use cases are illustrative rather than an assurance method, and the point here is drawn from its own instructions rather than against them. Nor is it a claim that a document-fed profile is unusable — a faster first draft with an honest gap register is better than a slow first draft without one.

It is also not a statement about anyone’s regulatory position. The Inference Institute does not certify organisations against ISO/IEC 42001, the NIST AI RMF or anything else, and a profile produced this way is evidence of readiness at best. Where a framework position carries legal or contractual weight, that interpretation stays with your counsel and your assessor.

The decision belongs to whoever owns the control assessment — usually the CISO, sometimes the head of second-line assurance. When a team hands back a profile drafted with a model, the question is not whether the tool was permitted. It is which artefact was accepted as the deliverable. Accept the table and the organisation has bought a faster restatement of its own paperwork. Accept the gap register and it has bought the list of things it does not know about itself, which is the only output of an assessment that was ever worth having.

Filed under · Governance · NIST · Evidence · Assurance Inference Institute · 30 Aug 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.