Start a conversation Contact
← Research

Research · page 04 of 04

Further back in the archive.

Articles

30 Jun 2026 An automated judge is an instrument. Calibrate it or do not read it. Scoring model outputs with another model has become the default evaluation method, and the published evidence says these judges can be highly repeatable while being systematically wrong in ways repeatability will never reveal. Method 30 Jun 2026 Prompt injection is not a bug you patch. It is what the interface is. Every mitigation for prompt injection is a filter placed in front of a component that cannot distinguish instructions from data. The defensible architecture assumes the model will be turned against you and limits what that is worth. Architecture 25 Jun 2026 The cheapest token is the one you did not send Caching is treated as an optimisation to be added once the bill hurts. It is a design decision that has to be made in the first week, because what a system can cache is determined entirely by how it assembles a prompt. Method 25 Jun 2026 In a screening system, the false positive is the product A classifier with excellent accuracy on a rare event still hands its operators far more wrong answers than right ones. That is not a modelling failure, it is arithmetic — and it determines the staffing plan, not just the evaluation report. Method 23 Jun 2026 Shadow AI is a measurement problem before it is a policy problem Most organisations respond to unsanctioned AI use by writing a policy. The policy is not the binding constraint, because nobody knows what is being used, for what, or on which data — and a rule written against an unknown population changes nothing. Governance 23 Jun 2026 Workflows first. Agents when the branch cannot be written down. The choice between a fixed pipeline and a model that decides its own next step is usually made for cultural reasons and defended for technical ones. There is a test that settles it, and it takes about ten minutes per use case. Architecture 18 Jun 2026 The deepfake did not defeat a control. It satisfied one. A finance employee approved fifteen transfers worth twenty-five million dollars after a video call with colleagues who were all synthetic. Nothing technical was breached. The process worked exactly as designed, and the design was the problem. Governance 18 Jun 2026 What an inference architecture has to hold Most enterprise AI systems have a serving layer that is really one call to a provider with some retry logic around it. Five things belong in that layer, and the two that are almost never built first are the two that cannot be added afterwards. Architecture 16 Jun 2026 Inference is the line item nobody owns The cost of running a model in production is not a price you negotiate with a provider. It is a utilisation number your own architecture sets, and almost all of it is decided by four choices made before anyone reads a rate card. Architecture 16 Jun 2026 Your assistant speaks for you, and a tribunal has already said so A dealership chatbot offered a new car for a dollar and a Canadian tribunal made an airline honour a refund policy its chatbot invented. The design lesson from both is the same one, and it is not about guardrails. Governance

Bring us the question

Something here already on your desk?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.