Request a scoping call Contact
← Research

Choose analysis work with an independent check

BootLoops gives AI agents a toolkit for exact and checkable quantitative work. Its useful design lesson is to choose analyses where an independent computation can verify the result, while a domain expert still decides whether the question matters.

Method / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

An analyst does not need an agent that sounds certain. They need one that can find a suitable method, run it through inspectable tools and return a result that can be checked independently. A new BootLoops preprint, published on 1 October, describes an open-source harness for AI agents doing quantitative science. The practical idea is broader than a scientific assistant: choose analytical work where a second computation can verify the result, and keep a domain expert responsible for deciding whether the result answers a worthwhile question.

BootLoops is a collection of programs and working protocols, not a new foundation model. Its tools include exact reductions, high-precision evaluation, interval arithmetic with error bounds and exhaustive enumeration. An agent can search the tool documentation, compose the programs and write the code needed to apply them. The software and acceptance tests then provide checks that are not just another language model’s opinion. The project’s repository describes it as model-independent and says the code was written by Claude under the author’s supervision. It is maintained by Matthew Schwartz, not an officially supported Anthropic product.

Put the check beside the calculation

The useful mechanism is a separation of roles. The model can propose a route through existing methods and tools. Ordinary programs perform the calculation. A separate implementation, held-out input or mathematical certificate checks the result. Where the method supports it, the output can be an exact value or an interval with a proven error radius under the method’s assumptions, rather than a rounded estimate with an unverified error bar.

In the preprint, the author reports 30 Feynman integrals computed in closed form. Fifteen had already appeared in the literature. For the other fifteen, the author reports finding no prior computation. The proposed forms were tested against independent numerical evaluations at points that were not used to fit them. The paper notes two qualifications to its comparison table, where the reported digits match published results and independent numerical checks reached different precisions. These are author-reported research results, not an independent replication or evidence that an agent can choose important questions without expert direction.

The distinction matters for enterprise analysis. A repeated calculation can be made faster without giving the agent authority over the assumptions. Imagine a research team asking an agent to reconstruct a published result from public inputs. The agent could find the method, draft the implementation and run it. The team could compare it with a separately written program, check known cases and record the full input and software versions. Agreement would support the calculation under those assumptions. It would not show that the assumptions describe the organisation, that the source data are complete or that the result justifies a business decision. This example is hypothetical.

hypothetical fixture

Set the acceptance test before the agent runs

Baseline
An analyst’s current script and documented calculation
Outcome
Reconstruct a bounded analysis from fixed public or synthetic inputs
Guardrails
No sensitive records or live decisions; preserve inputs, code, tool versions and reviewer notes
Decision rule
Agree with an independent implementation within a predeclared tolerance, and reject a planted error
Decision owner
Research lead
Population
A small, fixed set of calculations with known answers

An acceptance brief for a reproducible calculation. It does not establish that the selected model or research question is appropriate.

A hypothetical bounded analysis is accepted only when it agrees with a separately implemented calculation and a planted error is rejected.

Reviewed 2026-10-08

Start with a bounded analytical task

The first question is not which model to use. It is whether the task has a checkable acceptance condition. Recomputing a result from a published method, testing a fixed formula, enumerating a small finite space or deriving a bounded numerical value may qualify. Open-ended synthesis and judgment about scientific significance do not become mechanically verifiable merely because an agent writes code for them.

Before a trial, an analyst should write down the inputs, expected result, tolerance and an independent route to the answer. Include a planted mistake that the check must catch. Preserve the model and tool versions, generated code, data transformations, random seeds where relevant and the reviewer’s assessment. Compare the workflow with the existing analyst-led method on the same small set of public or synthetic cases. Measure review time, successful reproductions, detected errors and the work needed to repair a failure. The model’s explanation is context for review, not a substitute for executing the check.

The expert still owns the question and the interpretation. Schwartz’s account of the project says that some cross-disciplinary connections found by Claude were technically correct but scientifically unremarkable until researchers in those fields helped redirect the work. That is a useful division of labour: software can make a calculation exact, while people with subject knowledge decide whether the calculation is relevant.

Keep the claim within the evidence

BootLoops is an early, author-led research project. The author’s preprint and repository describe a substantial collection of methods and checks, but they do not establish lower costs, shorter analysis cycles or improved business decisions in an enterprise. Nor does an independently verified numerical result validate the model, input data or interpretation that produced it.

For an organisation, the opportunity is a bounded experiment in reproducibility, not a general delegation of analysis. Select one calculation with public or synthetic inputs, define how a second route can fail, and keep an analyst accountable for the assumptions and conclusion. If no independent route can challenge the result, label it as a model-generated candidate rather than a verified calculation.

Filed under · Method · Scientific computing · Agents · Reproducibility Inference Institute · 08 Oct 2026

Related engagement

The decision behind this article

A clear design your team or chosen delivery partner can build from.

Explore AI Architecture →

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.