Skip to content
AI.info

The Pulse

BootLoops Brings Certified Computation to AI Science Agents

BootLoops combines computational tools with protocols for checking AI-assisted quantitative research. Its examples span mathematical physics, population genetics, ecology and phylogenetics, with researchers involved in assessing the scienti

BootLoops Brings Certified Computation to AI Science Agents

AI.info Team ·

“Claude and GPT are good at science, but they are not scientists: yes, they are smart, but it can take a lot of hand-holding to get them to produce anything of scientific value.”

Matthew Schwartz, professor and author of the Anthropic guest post

That distinction frames BootLoops, an open-source toolkit for using large language models on quantitative scientific work. Schwartz introduced the project in a guest post published by Anthropic on October 1, 2026. The BootLoops 1.0 repository combines computational software with protocols meant to make an agent’s output testable, rather than asking users to accept an answer on the model’s authority.

From mathematical physics to other fields

Schwartz says he began by asking Claude to build tools for mathematical physics, then expanded the work as the software found applications elsewhere. The toolkit covers multidimensional integrals, exact recurrences, Bayesian evidence calculations and numerical methods with certified error bounds. Its documentation tells an agent what each instrument does and what checks its result must pass.

The project’s starting point was scattering amplitudes, calculations used in particle physics. Schwartz describes using Claude to port existing methods into a shared framework and extend them to elliptic integrals. In his account, the work produced 30 end-to-end integral calculations: 15 reproductions of known results and 15 that had not previously been computed using that approach.

Applications described in the Anthropic post reach beyond physics. In population genetics, Claude used methods from mathematical physics to solve an integral expression concerning how natural selection shapes rare mutations; Schwartz says the team applied it to the gnomAD catalog of human genetic variation. A separate population-genetics analysis examined 5.7 billion pairs of nearby mutations and found evidence for gene conversion. In phylogenetics, the work on Bayesian evidence for evolutionary trees is a distinct project. The repository describes Mixalot as a statistics codebase and Popcorn as a population-genetics codebase within the toolkit; JaCKandJill is the separate sibling repository for Bayesian phylogenetics.

Checks are part of the toolkit

BootLoops includes more than code for carrying out calculations. Its protocols call for independent checks at points not used to fit a result, positive controls that demonstrate a check can detect failure, planted-truth tests before analyzing real data, and records of which sources contributed to a fit. The guidance also tells agents to test a small run before committing to a larger one.

The repository lists 49 tool packages and provides a self-test command that runs their included test batteries. Its documentation says the run takes a few minutes on a laptop; three packages stop with named errors because the repository lacks data they need. Tests that require an external engine skip when it is not installed, and the documentation identifies what must be added.

BootLoops is written mainly in Python 3.12, with Julia components in some packages, and is released under the MIT license. The README reports validation on Debian 12 with Python 3.12 and Julia 1.11 across x86_64 and arm64 containers, while flagging an arm64 limitation in the Blade installer. macOS was not part of the release testing.

Schwartz still puts scientists in charge

The framework does not make an agent’s conclusion self-validating. Schwartz writes that Claude can misjudge whether a result matters, declare work finished with a key gap unresolved, and draw incorrect conclusions from calculations. His described workflow keeps human researchers involved to choose problems, approve plans, inspect results and consult domain experts.

That separation is also reflected in the project’s ownership: Schwartz says BootLoops is not an Anthropic project, even though he worked as a visiting researcher at the company during the work. The repository says he created it and that Claude wrote the code under his direction. BootLoops asks researchers to verify its results before relying on them and excludes clinical, actuarial, payment, regulatory and public-safety decisions from its intended uses.

Sources

Explore

More articles