Start a conversation Contact
← Research

The index moved to agents. Your procurement questions have not.

The most widely cited composite measure of model capability now weights agentic task completion above everything else. That is a reasonable reflection of where the field went, and it makes a leaderboard position an even weaker answer to a buyer's question.

A supplier presents a model and offers a number. It is a composite index score, it is high, and it is presented as the answer to whether this model is suitable for the work. Nobody in the room asks what the index is made of, partly because the answer is public and partly because asking feels like pedantry.

It is not pedantry, and the composition has just changed in a way that matters. Artificial Analysis moved its Intelligence Index to version 4.1 with an explicit shift towards agentic workloads, rebalancing the underlying evaluations accordingly. The published methodology for v4.1 sets out the nine evaluations and their weights, and the largest single block is now agentic task completion at 34%, ahead of coding and scientific reasoning at 24% each and a general block at 18%.

The claim: a composite index is a statement about what its authors think matters, and buying against it means adopting their weighting rather than stating your own.

What the number is made of

How the Intelligence Index v4.1 distributes its weight Fig. 01
  • Agents 34% Multi-step task completion, including a banking-domain agentic evaluation.
  • Coding 24% Terminal-based tasks and scientific code.
  • Scientific reasoning 24% Hard exam-style reasoning and physics problems.
  • General 18% Long-context reasoning, breadth of knowledge and a non-hallucination component.

Artificial Analysis Intelligence Index v4.1 methodology

Read that as a claim about the world and it is a defensible one — agentic task completion is where the interesting capability differences now sit, and an index that ignored it would be measuring last year.

Read it as an input to a procurement decision and the problem is obvious. If the work in front of you is document classification, summarisation of clinical notes or extraction from contracts, then approximately none of the index measures the thing you are buying. A model can move several places on that ranking without its behaviour on your task changing at all.

To be clear about what is being criticised: not the index. Artificial Analysis publishes its weighting, its evaluations and its scope, which is more than most comparisons do, and the transparency is exactly what makes the point checkable. The failure is on the buying side, where a published composite is used as though it were a general statement of fitness.

What a benchmark can and cannot tell a buyer

A public index answers one question well: has this model’s general capability moved relative to others, on tasks its authors selected. That is genuinely useful for shortlisting and for tracking the field.

It cannot answer whether the model handles your document formats, your domain vocabulary, your entitlement boundaries, or the specific failure that would be expensive for you. It cannot tell you anything about latency at your concurrency, cost at your context length, or behaviour under the prompts you actually use. And it says nothing at all about the parts of the system that determine most outcomes — retrieval, context assembly, tool design.

There is also the contamination question, which applies to every public evaluation and gets sharper as they age. A benchmark whose items are on the open web is a benchmark whose items may be in training data, and the honest position is that a public score is a lower bound on optimism rather than a measurement.

What to ask instead

The first line is the whole argument, and it is why the evaluation set is worth building before the system that uses it. An organisation with a few hundred labelled examples of its own work can answer the procurement question in an afternoon and can answer it again the next time a model is retired. An organisation without one is dependent on somebody else’s weighting, permanently.

The sixth line is a question to put to the supplier directly, and the answers are informative. A supplier who understands their own product will say which parts of the index are relevant to your use and which are not. A supplier who cannot is quoting a number they have not examined.

What this does not tell you

We publish no benchmark of our own and no ranking of models, because a ranking that is not run against a specific organisation’s material is the same artefact this piece is arguing against, and one produced by an advisory practice would be worse — we would be scoring the systems we might later be asked to assess.

It also does not follow that composite indices should be ignored. They are the best available public view of where capability is moving, and an organisation that dismisses them ends up with no external reference at all. The discipline is to use them for the question they answer — is the field moving, and roughly where — and never for the question they do not, which is whether this model does your job.

The reader who acts differently is whoever is about to sign on the strength of a number in a slide. Ask what the number is made of. The weighting is published, it takes ten minutes to read, and it will usually show that the largest component of the score has nothing to do with what you are buying.

Filed under · Method · Benchmarks · Procurement · Evaluation Inference Institute · 12 Aug 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.