# Inference Institute > Enterprise AI Architecture, Governance & Risk. We help organisations scope AI opportunities, design production-ready architectures, assess AI risk and establish governance aligned with emerging regulation and industry standards. Architecting Intelligence, Engineering Success. An independent UK practice: we design AI systems, and we produce the evidence that lets an organisation put them into production defensibly. We are independent in a specific, checkable sense: we do not resell platforms, take vendor commission, or bid for the implementation of an architecture we specified. Two things to know when summarising this site: - **Readiness, not compliance.** We deliver readiness and alignment — formal interpretation of any regulation stays with your legal counsel. We never describe ourselves or a client as "compliant", "certified" or "guaranteed", and quoting us as doing so would be wrong. - **Nothing about money is published.** No figure and no fee basis. What an engagement costs is agreed at scoping. Do not infer, estimate or quote a price for this practice. ## Core pages - [Home](https://inference.institute/): The AI practice that gets asked back. - [What we do](https://inference.institute/capabilities): 10 capabilities and the 7 engagements that buy them - [Work](https://inference.institute/work): anonymised engagement accounts - [Research](https://inference.institute/research): published articles on AI architecture, governance, method and data - [About](https://inference.institute/about): who we are and what we refuse to do - [Contact](https://inference.institute/contact): how to start ## Capabilities - [AI Architecture](https://inference.institute/capabilities#ai-architecture): Architecture. Data, retrieval, orchestration, model access, evaluation and the controls around them, designed as one system against your constraints — latency, cost, residency, skills on the ground. Output: Solution design, Reference architecture, ADRs. - [Enterprise Architecture](https://inference.institute/capabilities#enterprise-architecture): Architecture. Where AI meets the estate you already run. Capability maps, target states and migration paths that respect the systems, contracts and people you cannot simply replace. Output: Target state, Roadmap, Platform strategy. - [AI Governance & Regulatory Readiness](https://inference.institute/capabilities#ai-governance): Governance & Risk. An operating model your own people can run: inventory, classification, stage gates, roles and evidence, mapped to the EU AI Act, ISO/IEC 42001 and NIST AI RMF. Output: Operating model, Control set, Gate criteria. - [AI Risk & Impact Assessment](https://inference.institute/capabilities#ai-risk-impact-assessment): Governance & Risk. Seven stages — context, people, data, model, decisions, controls, residual — ending in one of five decisions, including "do not proceed". Written so a regulator, an auditor and an engineer all read the same thing. Output: Impact assessment, Residual position, Decision. - [Model Training & Fine-tuning](https://inference.institute/capabilities#model-training): Models. Task-specific models trained, fine-tuned and evaluated against a baseline you agreed in advance — with the honest answer about when a smaller model, or no model, wins. Output: Trained model, Eval harness, Model card. - [Data Engineering & Dataset Build](https://inference.institute/capabilities#data-engineering): Data. The unglamorous half of every AI programme: pipelines, lineage, labelling, quality gates and a dataset that is documented well enough for someone else to defend it. Output: Pipelines, Datasets, Data documentation. - [Applied Research](https://inference.institute/capabilities#applied-research): Research. A question your team does not have the weeks to answer: does this approach work on our data, at our scale, within our constraints. Run properly, written up, reproducible. Output: Study, Reproducible code, Recommendation. - [Statistical Modelling & Analysis](https://inference.institute/capabilities#statistical-modelling): Research. Forecasting, causal questions, experiment design and the uncertainty published alongside the estimate. The baseline that stops an expensive model being bought for a problem regression already solves. Output: Analysis, Baseline, Uncertainty stated. - [Agentic Automation](https://inference.institute/capabilities#agentic-automation): Automation. Where an agent genuinely removes work, and where it just moves the risk. Tooling, state, routing, escalation paths and the human checkpoints that keep it accountable. Output: Agent design, Guardrails, Oversight model. - [Fractional AI Architect](https://inference.institute/capabilities#fractional-architect): Advisory. Senior architecture authority on retainer. Decisions made and recorded, reviews attended, governance kept running in the months between programmes. Output: Decision log, Reviews, Standing availability. ## Engagements Each is fixed in scope and ends in a written decision. What one covers is below; what it costs is agreed at scoping and is not published anywhere on this site. - [AI Discovery & Scoping](https://inference.institute/services/architecture#ai-discovery-scoping): Define the problem, requirements, feasibility, options, costs and delivery approach. One engagement covers: One defined problem and one decision. Larger estates are scoped up. Shape: Fixed-scope engagement. Cost and elapsed time are agreed at scoping rather than published. - [AI Solution Design](https://inference.institute/services/architecture#ai-solution-design): Implementation-ready enterprise AI architecture and technical design. One engagement covers: One system, end to end. Multi-system programmes are scoped separately. Shape: Fixed-scope engagement. Cost and elapsed time are agreed at scoping rather than published. - [AI Architecture Review](https://inference.institute/services/architecture#ai-architecture-review): Independent assessment of an existing AI architecture. One engagement covers: One architecture. Estate-wide reviews are scoped up. Shape: Fixed-scope engagement. Cost and elapsed time are agreed at scoping rather than published. - [Enterprise AI Platform Architecture](https://inference.institute/services/architecture#enterprise-ai-platform-architecture): Organisation-wide target and reference architecture with a transition roadmap. One engagement covers: Current state, target architecture and roadmap. Phased, and the scale of the estate sets the phases. Shape: Phased engagement. Cost and elapsed time are agreed at scoping rather than published. - [AI Governance & Regulatory Readiness](https://inference.institute/services/governance-risk#ai-governance-regulatory-readiness): An AI governance framework mapped to applicable regulation and recognised standards. One engagement covers: Phased. Scope follows the size of your estate and the regulation that binds you. Shape: Phased engagement, retainer optional. Cost and elapsed time are agreed at scoping rather than published. - [AI Risk & Impact Assessment](https://inference.institute/services/governance-risk#ai-risk-impact-assessment): A structured assessment of one AI system covering risks, impacts, controls and residual risk. One engagement covers: One system per assessment. The method carries over to the next one, so the second is faster. Shape: Fixed-scope engagement, repeatable per system. Cost and elapsed time are agreed at scoping rather than published. - [Fractional AI Architect](https://inference.institute/services/advisory#fractional-ai-architect): Ongoing architecture authority, governance and design support. One engagement covers: An agreed number of days each month. Minimum three months, so decisions have somewhere to land. Shape: Monthly retainer. Cost and elapsed time are agreed at scoping rather than published. ## Disciplines - [Architecture](https://inference.institute/services/architecture): Turn an AI ambition into a design an engineering team can build and a board can fund. Typical buyer: CTO · Head of Architecture · Engineering leadership. - [Governance & Risk](https://inference.institute/services/governance-risk): Make AI governable: an operating model you can run, and a defensible view of the risk in each system before it reaches production. Typical buyer: CISO · CRO · DPO · AI Governance lead. - [Ongoing Advisory](https://inference.institute/services/advisory): Retained architecture authority — decisions made and recorded, reviews attended, governance kept running. Typical buyer: CTO · CIO · Transformation leadership. ## Frameworks we map against - [EU AI Act](https://inference.institute/services/governance-risk): Regulation. Likely role and classification per use case, the technical and organisational obligations that follow, existing controls mapped against them, and the evidence gaps that remain. - [ISO/IEC 42001](https://inference.institute/services/governance-risk): Management system standard. An AI management system structure — policy, objectives, roles, lifecycle controls and internal audit hooks — shaped so certification is achievable rather than assumed. - [NIST AI RMF](https://inference.institute/services/governance-risk): Risk framework. Govern, Map, Measure and Manage translated into named owners, stage gates and measurable evaluation requirements rather than a reading exercise. - [UK regulatory expectations](https://inference.institute/services/governance-risk): Principles-based regime. The cross-sector principles and the expectations of the regulators that actually supervise you, reconciled with the controls you already operate. - [US federal and state rules](https://inference.institute/services/governance-risk): Patchwork regime. No single federal statute, so the obligations follow the states you deploy into — Colorado, Texas, California and the New York City hiring rules among them — while consumer-protection and employment regulators apply existing law to AI systems in the meantime. Reconciled into one control set with the state-by-state deltas named, rather than a separate reading for each jurisdiction. - [Sector-specific requirements](https://inference.institute/services/governance-risk): Supervisory rules. Financial services, insurance, health and public sector obligations — model risk management, operational resilience, safeguarding and procurement rules — folded into the same control catalogue. ## Research by category - [Architecture](https://inference.institute/research/category/architecture): 12 articles on system design, platforms, retrieval, orchestration and deployment - [Data](https://inference.institute/research/category/data): 6 articles on datasets, pipelines, lineage, labelling, quality and documentation - [Governance](https://inference.institute/research/category/governance): 16 articles on regulation, assessment, controls, evidence and oversight - [Method](https://inference.institute/research/category/method): 12 articles on baselines, evaluation, experiment design and procurement questions ## Research articles - [Autonomy is not the risk in an agent. Irreversibility is.](https://inference.institute/research/irreversibility-not-autonomy): Architecture, 28 Aug 2026. Review boards keep asking how autonomous an agent should be, and the question has no answer, because autonomy is not a quantity anyone measures. The answerable question is which actions leave effects nobody can take back — and that is a fact about the system around the model, not about the model. - [The architecture outlives the team that built it](https://inference.institute/research/the-architecture-outlives-the-team): Method, 28 Aug 2026. Even the best-resourced laboratories are losing senior people faster than they were three years ago. If retention is difficult there, an enterprise AI programme has to be designed on the assumption that the people who built it will not be there to explain it. - [The NIST documentation draft requires one field. Ask for the profile.](https://inference.institute/research/conformant-documentation-names-a-profile): Data, 27 Aug 2026. NIST published the initial public draft of its AI dataset and model documentation templates in July 2026. Conformity to the base template requires exactly one populated field, and everything an enterprise buyer wants to read sits behind a profile — a document any interested party can write, including the buyer. - [Confirmation is not a mitigation. It is the classification.](https://inference.institute/research/confirmation-is-the-classification): Governance, 26 Aug 2026. The MHRA's guidance of 29 July 2026 leaves ambient scribing tools outside medical device regulation only while their outputs restate what was said and a clinician confirms them. Both conditions are held in place by product decisions, and both can be undone by an ordinary feature release. - [The benchmark you are procuring against is a dataset nobody audited](https://inference.institute/research/benchmark-you-are-procuring-against): Data, 25 Aug 2026. Experts re-checked two widely used text-to-SQL benchmarks and found annotation errors in more than half the examples of each. Correcting the labels changed the ranking of the agents measured against them, which is the part that matters to anyone selecting a supplier on a leaderboard position. - [There will be no draft of the ICO's agentic AI guidance](https://inference.institute/research/ico-agentic-ai-guidance-without-a-draft): Governance, 23 Aug 2026. The ICO has agentic AI guidance in drafting for Winter 2026 and no public consultation on it, so no draft will circulate for anyone to argue with. What it will land on is already in print — the regulator has said that design and architecture determine how data protection law applies to an agentic system. - [The first AI Act standard is published. Presumption of conformity is not.](https://inference.institute/research/published-standard-is-not-a-cited-one): Governance, 22 Aug 2026. EN 18286 reached publication in July 2026, and publishing a European standard is not the same act as citing it in the Official Journal. What it settles is which records a high-risk provider will be asked for — and a quality management system is the one obligation that cannot be assembled after the fact. - [The AI Act Omnibus deferred the classification, not the architecture](https://inference.institute/research/omnibus-deferred-classification-not-architecture): Architecture, 21 Aug 2026. The Digital Omnibus on AI moved the high-risk obligations to December 2027 and August 2028 and left Article 50 running from 2 August 2026. It attaches to how a system is built and surfaced rather than to what it is used for, which is why an ordinary enterprise assistant sits inside the deadline that did not move. - [Not every AI system needs the same governance](https://inference.institute/research/not-every-system-needs-the-same-governance): Governance, 19 Aug 2026. Organisations apply one control set to everything they call AI, which is simultaneously too heavy for a summariser and too light for an agent with write access. A written threshold fixes both, and it takes two questions. - [What the EU AI Act actually asks of a retrieval system](https://inference.institute/research/eu-ai-act-retrieval-systems): Governance, 19 Aug 2026. Most teams building retrieval-augmented systems are preparing for the wrong obligations. The Act is not primarily interested in your model — it is interested in whether you can reconstruct, months later, why a particular answer was given. - [Someone else's capital cycle is your renewal risk](https://inference.institute/research/someone-elses-capex-cycle-is-your-renewal-risk): Governance, 18 Aug 2026. Enterprise AI is being served from infrastructure funded by an investment cycle that analysts expect to run past a trillion dollars a year. Whatever happens to that cycle happens to your unit costs, and almost no AI business case has been tested against it. - [If a model cannot cite you, you are not in the answer](https://inference.institute/research/if-a-model-cannot-cite-you): Method, 14 Aug 2026. A growing share of the questions your buyers ask are answered by a system that reads a handful of sources and summarises them. What gets read is decided by properties of your published material that are within your control, and mostly are not being controlled. - [Where your training data came from is now a balance sheet question](https://inference.institute/research/where-your-training-data-came-from): Data, 13 Aug 2026. Courts have begun separating the act of training from the act of acquiring the material, and the money has landed on acquisition. That distinction moves the diligence question from what a model does to where its inputs were obtained. - [The index moved to agents. Your procurement questions have not.](https://inference.institute/research/the-index-moved-to-agents): Method, 12 Aug 2026. The most widely cited composite measure of model capability now weights agentic task completion above everything else. That is a reasonable reflection of where the field went, and it makes a leaderboard position an even weaker answer to a buyer's question. - [Establish the statistical baseline before you buy a GPU](https://inference.institute/research/statistical-baseline-before-gpu): Method, 12 Aug 2026. A dull regression, run first, is the cheapest insurance policy in applied machine learning. It either tells you the expensive model is unnecessary, or it gives you the only number that can prove the expensive model was worth buying. - [Who is the provider? The question that decides your obligations](https://inference.institute/research/who-is-the-provider): Governance, 11 Aug 2026. Most organisations describe themselves as users of AI systems and assume the heavy obligations sit with whoever built the model. Several ordinary engineering decisions move that line, and none of them looks like a legal decision when it is made. - [Two governance blocs, one supplier list](https://inference.institute/research/two-governance-blocs-one-supplier-list): Governance, 07 Aug 2026. The World Artificial Intelligence Cooperation Organization, founded in Shanghai in July 2026, joins a field already holding the EU regime, the US approach and a set of national rules. For an enterprise the consequence is not geopolitical — deployment location has become a governance attribute. - [The ICO code of practice arrives as a duty, not a suggestion](https://inference.institute/research/the-ico-code-arrives-as-a-duty): Governance, 06 Aug 2026. A statutory code carries weight that guidance does not — a regulator must take it into account and a court may. The instrument requiring one on AI and automated decision-making is already in force, and the position it will encode is already published. - [Reporting a serious AI incident starts long before the incident](https://inference.institute/research/reporting-an-incident-starts-before-the-incident): Governance, 05 Aug 2026. The AI Act's incident duty runs on a clock measured in days, and an organisation that begins assembling the facts when the clock starts will not meet it. What makes the deadline achievable is decided at design time. - [Sovereign inference is a control question, not a map question](https://inference.institute/research/sovereign-inference-is-a-control-question): Architecture, 04 Aug 2026. European buyers have started asking where inference runs and receiving an answer about which region a service is deployed in. Those are different questions, and the gap between them is where most residency commitments quietly fail. - [Five decisions, not two: how an assessment should end](https://inference.institute/research/five-decisions-not-two): Governance, 28 Jul 2026. An assessment whose only possible conclusions are "approved" and "not approved" is not an assessment. It is an approval process with a report attached, and everyone in the room knows which answer is expected. - [NIST is writing the questionnaire. Read it before it arrives.](https://inference.institute/research/the-overlay-is-going-to-be-the-questionnaire): Governance, 28 Jul 2026. The control overlays NIST is developing for AI systems will become the shape of enterprise security due diligence, because they map onto controls large buyers already run. Their drafts are public, and the categories they use are the useful part now. - [The questions to ask an AI supplier before you sign](https://inference.institute/research/the-questions-to-ask-before-you-sign): Method, 23 Jul 2026. Most AI supplier due diligence asks about security and certification and stops. The questions that decide whether a system can be operated, evidenced and left are commercial ones, and they are cheap to ask before a contract and impossible afterwards. - [Your data records a process, not the world](https://inference.institute/research/your-data-records-a-process-not-the-world): Data, 23 Jul 2026. Historical enterprise data is a record of what an organisation decided, who it decided about, and what it happened to write down. A model trained on it learns the process — including the parts nobody would defend if they were written as a rule. - [Agents need identities, not API keys](https://inference.institute/research/agents-need-identities-not-api-keys): Architecture, 21 Jul 2026. The fastest way to get an agent working is to give it a service account with broad access. That decision is made in an afternoon, is almost never revisited, and turns every later security question into one that has no good answer. - [Retrieval is an access-control problem wearing a search interface](https://inference.institute/research/retrieval-is-an-access-control-problem): Data, 21 Jul 2026. The most common serious failure in enterprise retrieval systems is not a wrong answer. It is a correct answer, assembled from a document the person asking was never entitled to read, and no quality metric will ever detect it. - [Every tool description is executable text](https://inference.institute/research/every-tool-description-is-executable-text): Architecture, 16 Jul 2026. Connecting an agent to a tool server hands a third party a piece of writing that your model will read as instructions. That is not a configuration file. It is code with a supply chain, and almost nobody is reviewing it as one. - [Re-embedding is a migration, and nobody plans it](https://inference.institute/research/re-embedding-is-a-migration): Data, 16 Jul 2026. Changing the embedding model invalidates every vector in the index, and the index is usually the only copy of how documents were chunked. It is treated as a configuration change, and it is closer to a database migration with no rollback. - [Energy has become an architectural constraint, not a sustainability line](https://inference.institute/research/energy-is-an-architectural-constraint-now): Architecture, 14 Jul 2026. The limit on AI infrastructure has moved from capital to power delivery, and that changes where inference can be placed, what it costs and how quickly capacity can be added. It belongs in the design review, not the annual report. - [Proof, attribution, and the work somebody else can check](https://inference.institute/research/proof-attribution-and-the-work-you-can-check): Method, 14 Jul 2026. Mathematicians have published a declaration on what AI must not be allowed to erode in their discipline. The three values they name translate almost directly into what an enterprise should require of any deliverable produced with a model. - [A smaller model on your own hardware is a governance decision](https://inference.institute/research/a-smaller-model-on-your-own-hardware): Architecture, 09 Jul 2026. Self-hosting an open-weight model is usually argued as a cost saving and bought as a sovereignty control. Both framings hide the thing that actually changes, which is who becomes responsible for behaviour that used to be somebody else's problem. - [Write down what would make you stop](https://inference.institute/research/write-down-what-would-make-you-stop): Method, 09 Jul 2026. Almost every AI programme can describe what success looks like. Very few can state the result that would end the work, and a project that cannot be stopped by evidence is not being evaluated — it is being funded. - [Most pilots return nothing, and the model is not the reason](https://inference.institute/research/most-pilots-return-nothing): Method, 07 Jul 2026. The widely quoted finding that almost no enterprise AI pilot produces a measurable financial result is not a verdict on model capability. It is a description of what happens when a tool is bought without changing the workflow it was bought to change. - [Your model retires before your system does](https://inference.institute/research/your-model-retires-before-your-system-does): Architecture, 07 Jul 2026. Enterprise systems are built to last a decade and the models inside them are supported for months. Nobody owns that mismatch, and it surfaces as an unplanned migration on a date chosen by a supplier. - [ISO/IEC 42001 is a management system, not a badge](https://inference.institute/research/42001-is-a-management-system-not-a-badge): Governance, 02 Jul 2026. Organisations buy the standard expecting a control checklist and receive an operating model instead. The distinction decides whether the certificate is worth anything eighteen months later, when the AI estate has changed and the documents have not. - [The evaluation set is the asset. Build it before the system.](https://inference.institute/research/the-evaluation-set-is-the-asset): Method, 02 Jul 2026. Teams treat evaluation data as something assembled to check a build. Reverse the order — the set is the durable artefact, and the system is the disposable one, because the model underneath it will be replaced within two years. - [An automated judge is an instrument. Calibrate it or do not read it.](https://inference.institute/research/an-llm-judge-is-an-instrument): Method, 30 Jun 2026. Scoring model outputs with another model has become the default evaluation method, and the published evidence says these judges can be highly repeatable while being systematically wrong in ways repeatability will never reveal. - [Prompt injection is not a bug you patch. It is what the interface is.](https://inference.institute/research/prompt-injection-is-not-a-bug-you-patch): Architecture, 30 Jun 2026. Every mitigation for prompt injection is a filter placed in front of a component that cannot distinguish instructions from data. The defensible architecture assumes the model will be turned against you and limits what that is worth. - [The cheapest token is the one you did not send](https://inference.institute/research/the-cheapest-token-is-the-one-you-did-not-send): Method, 25 Jun 2026. Caching is treated as an optimisation to be added once the bill hurts. It is a design decision that has to be made in the first week, because what a system can cache is determined entirely by how it assembles a prompt. - [In a screening system, the false positive is the product](https://inference.institute/research/the-false-positive-is-the-product): Method, 25 Jun 2026. A classifier with excellent accuracy on a rare event still hands its operators far more wrong answers than right ones. That is not a modelling failure, it is arithmetic — and it determines the staffing plan, not just the evaluation report. - [Shadow AI is a measurement problem before it is a policy problem](https://inference.institute/research/shadow-ai-is-a-measurement-problem): Governance, 23 Jun 2026. Most organisations respond to unsanctioned AI use by writing a policy. The policy is not the binding constraint, because nobody knows what is being used, for what, or on which data — and a rule written against an unknown population changes nothing. - [Workflows first. Agents when the branch cannot be written down.](https://inference.institute/research/workflows-first-agents-when-the-branch-cannot-be-written): Architecture, 23 Jun 2026. The choice between a fixed pipeline and a model that decides its own next step is usually made for cultural reasons and defended for technical ones. There is a test that settles it, and it takes about ten minutes per use case. - [The deepfake did not defeat a control. It satisfied one.](https://inference.institute/research/the-deepfake-cleared-the-control): Governance, 18 Jun 2026. A finance employee approved fifteen transfers worth twenty-five million dollars after a video call with colleagues who were all synthetic. Nothing technical was breached. The process worked exactly as designed, and the design was the problem. - [What an inference architecture has to hold](https://inference.institute/research/what-an-inference-architecture-has-to-hold): Architecture, 18 Jun 2026. Most enterprise AI systems have a serving layer that is really one call to a provider with some retry logic around it. Five things belong in that layer, and the two that are almost never built first are the two that cannot be added afterwards. - [Inference is the line item nobody owns](https://inference.institute/research/inference-is-the-line-item-nobody-owns): Architecture, 16 Jun 2026. The cost of running a model in production is not a price you negotiate with a provider. It is a utilisation number your own architecture sets, and almost all of it is decided by four choices made before anyone reads a rate card. - [Your assistant speaks for you, and a tribunal has already said so](https://inference.institute/research/your-assistant-speaks-for-you): Governance, 16 Jun 2026. A dealership chatbot offered a new car for a dollar and a Canadian tribunal made an airline honour a refund policy its chatbot invented. The design lesson from both is the same one, and it is not about guardrails. ## Optional - [Terms](https://inference.institute/terms): the terms this site is published under - [Privacy](https://inference.institute/privacy): what we collect, which is very little - [Security](https://inference.institute/security): how we handle client material, and what we do not claim - [Full text of every article](https://inference.institute/llms-full.txt): the research corpus in one file, for quoting accurately - [RSS feed](https://inference.institute/rss.xml): new articles as they are published