Choose analysis work with an independent check
BootLoops gives AI agents a toolkit for exact and checkable quantitative work. Its useful design lesson is to choose analyses where an independent computation can verify the result, while a domain expert still decides whether the question matters.
Keep development and test groups separate before comparing performance. No measured results are shown.
An analyst does not need an agent that sounds certain. They need one that can find a suitable method, run it through inspectable tools and return a result that can be checked independently. A new BootLoops preprint, published on 1 October, describes an open-source harness for AI agents doing quantitative science. The practical idea is broader than a scientific assistant: choose analytical work where a second computation can verify the result, and keep a domain expert responsible for deciding whether the result answers a worthwhile question.
BootLoops is a collection of programs and working protocols, not a new foundation model. Its tools include exact reductions, high-precision evaluation, interval arithmetic with error bounds and exhaustive enumeration. An agent can search the tool documentation, compose the programs and write the code needed to apply them. The software and acceptance tests then provide checks that are not just another language model’s opinion. The project’s repository describes it as model-independent and says the code was written by Claude under the author’s supervision. It is maintained by Matthew Schwartz, not an officially supported Anthropic product.
Put the check beside the calculation
The useful mechanism is a separation of roles. The model can propose a route through existing methods and tools. Ordinary programs perform the calculation. A separate implementation, held-out input or mathematical certificate checks the result. Where the method supports it, the output can be an exact value or an interval with a proven error radius under the method’s assumptions, rather than a rounded estimate with an unverified error bar.
In the preprint, the author reports 30 Feynman integrals computed in closed form. Fifteen had already appeared in the literature. For the other fifteen, the author reports finding no prior computation. The proposed forms were tested against independent numerical evaluations at points that were not used to fit them. The paper notes two qualifications to its comparison table, where the reported digits match published results and independent numerical checks reached different precisions. These are author-reported research results, not an independent replication or evidence that an agent can choose important questions without expert direction.
The distinction matters for enterprise analysis. A repeated calculation can be made faster without giving the agent authority over the assumptions. Imagine a research team asking an agent to reconstruct a published result from public inputs. The agent could find the method, draft the implementation and run it. The team could compare it with a separately written program, check known cases and record the full input and software versions. Agreement would support the calculation under those assumptions. It would not show that the assumptions describe the organisation, that the source data are complete or that the result justifies a business decision. This example is hypothetical.
Set the acceptance test before the agent runs
- Baseline
- An analyst’s current script and documented calculation
- Outcome
- Reconstruct a bounded analysis from fixed public or synthetic inputs
- Guardrails
- No sensitive records or live decisions; preserve inputs, code, tool versions and reviewer notes
- Decision rule
- Agree with an independent implementation within a predeclared tolerance, and reject a planted error
- Decision owner
- Research lead
- Population
- A small, fixed set of calculations with known answers
An acceptance brief for a reproducible calculation. It does not establish that the selected model or research question is appropriate.
A hypothetical bounded analysis is accepted only when it agrees with a separately implemented calculation and a planted error is rejected.
- BootLoops preprint · Sections 3.3 and 6
- BootLoops code and protocols · README, Status and provenance
Reviewed 2026-10-08
Start with a bounded analytical task
The first question is not which model to use. It is whether the task has a checkable acceptance condition. Recomputing a result from a published method, testing a fixed formula, enumerating a small finite space or deriving a bounded numerical value may qualify. Open-ended synthesis and judgment about scientific significance do not become mechanically verifiable merely because an agent writes code for them.
Before a trial, an analyst should write down the inputs, expected result, tolerance and an independent route to the answer. Include a planted mistake that the check must catch. Preserve the model and tool versions, generated code, data transformations, random seeds where relevant and the reviewer’s assessment. Compare the workflow with the existing analyst-led method on the same small set of public or synthetic cases. Measure review time, successful reproductions, detected errors and the work needed to repair a failure. The model’s explanation is context for review, not a substitute for executing the check.
The expert still owns the question and the interpretation. Schwartz’s account of the project says that some cross-disciplinary connections found by Claude were technically correct but scientifically unremarkable until researchers in those fields helped redirect the work. That is a useful division of labour: software can make a calculation exact, while people with subject knowledge decide whether the calculation is relevant.
Keep the claim within the evidence
BootLoops is an early, author-led research project. The author’s preprint and repository describe a substantial collection of methods and checks, but they do not establish lower costs, shorter analysis cycles or improved business decisions in an enterprise. Nor does an independently verified numerical result validate the model, input data or interpretation that produced it.
For an organisation, the opportunity is a bounded experiment in reproducibility, not a general delegation of analysis. Select one calculation with public or synthetic inputs, define how a second route can fail, and keep an analyst accountable for the assumptions and conclusion. If no independent route can challenge the result, label it as a model-generated candidate rather than a verified calculation.