Articles
Proof, attribution, and the work somebody else can check Mathematicians have published a declaration on what AI must not be allowed to erode in their discipline. The three values they name translate almost directly into what an enterprise should require of any deliverable produced with a model. A smaller model on your own hardware is a governance decision Self-hosting an open-weight model is usually argued as a cost saving and bought as a sovereignty control. Both framings hide the thing that actually changes, which is who becomes responsible for behaviour that used to be somebody else's problem. Write down what would make you stop Almost every AI programme can describe what success looks like. Very few can state the result that would end the work, and a project that cannot be stopped by evidence is not being evaluated — it is being funded. Most pilots return nothing, and the model is not the reason The widely quoted finding that almost no enterprise AI pilot produces a measurable financial result is not a verdict on model capability. It is a description of what happens when a tool is bought without changing the workflow it was bought to change. Your model retires before your system does Enterprise systems are built to last a decade and the models inside them are supported for months. Nobody owns that mismatch, and it surfaces as an unplanned migration on a date chosen by a supplier. ISO/IEC 42001 is a management system, not a badge Organisations buy the standard expecting a control checklist and receive an operating model instead. The distinction decides whether the certificate is worth anything eighteen months later, when the AI estate has changed and the documents have not. The evaluation set is the asset. Build it before the system. Teams treat evaluation data as something assembled to check a build. Reverse the order — the set is the durable artefact, and the system is the disposable one, because the model underneath it will be replaced within two years. An automated judge is an instrument. Calibrate it or do not read it. Scoring model outputs with another model has become the default evaluation method, and the published evidence says these judges can be highly repeatable while being systematically wrong in ways repeatability will never reveal. Prompt injection is not a bug you patch. It is what the interface is. Every mitigation for prompt injection is a filter placed in front of a component that cannot distinguish instructions from data. The defensible architecture assumes the model will be turned against you and limits what that is worth. The cheapest token is the one you did not send Caching is treated as an optimisation to be added once the bill hurts. It is a design decision that has to be made in the first week, because what a system can cache is determined entirely by how it assembles a prompt. In a screening system, the false positive is the product A classifier with excellent accuracy on a rare event still hands its operators far more wrong answers than right ones. That is not a modelling failure, it is arithmetic — and it determines the staffing plan, not just the evaluation report. Shadow AI is a measurement problem before it is a policy problem Most organisations respond to unsanctioned AI use by writing a policy. The policy is not the binding constraint, because nobody knows what is being used, for what, or on which data — and a rule written against an unknown population changes nothing.
Bring us the question
Something here already on your desk?
That is the conversation we are best at. Thirty minutes, a written summary, no obligation.