TL;DR
Paper2Agent turns a research paper and its supporting code into tested tools that an AI agent can call and reuse.
In a 100-paper computational-biology sample, 74 became working agents. The other 26 exposed missing code or data, broken dependencies and scripts that couldn’t run beyond the original experiment.
Once a paper can execute, reproducing a result requires knowing which build ran. The citation identifies the source. It doesn’t identify what actually ran.

A digital object identifier, or DOI, can tell me which paper you used.
It can’t tell me which code you ran.
That gap has always existed in computational research. A new system called Paper2Agent makes it harder to ignore because it turns a paper’s methods into tools that an AI agent can call.
The headline version is that you can now talk to a scientific paper. That’s true, but it undersells the change.
Paper2Agent goes beyond a chat interface over a PDF. It takes the manuscript, public codebase, supplementary material, datasets and example workflows, then builds a Model Context Protocol server. MCP is a standard interface that lets an AI client discover and call tools or retrieve structured resources. In this case, the tools are functions derived from the paper’s code.
The resulting paper agent can reproduce analyses, regenerate figures, apply the method to new data and combine capabilities from several papers.
That turns the publication into something closer to a software dependency. And dependencies need a different kind of receipt.
The Paper Isn’t Doing This Alone
The phrase “executable paper” can create the wrong picture. Paper2Agent isn’t reading prose and inventing a working scientific method from scratch.
It starts with the paper and the research artifacts around it. The system locates the associated repository, creates an isolated environment, runs tutorials to obtain reference outputs, extracts useful functions and exposes them as MCP tools. It then tests those tools against the source implementation. A test can check whether expected files were created, whether numerical outputs fall within a tolerance or whether a generated figure matches the reference.
Tools that repeatedly fail are left out. Tools that pass are included in the generated server without further changes and include references back to the original source code.
That validation layer matters. In the AlphaGenome case study, removing the test-and-repair stage dropped accuracy on tutorial-derived questions from 98.7% to 69.3%. The interface isn’t the main achievement. The work is in getting from a repository that a person can inspect to a callable function another system can use reliably.
The researchers tested Paper2Agent beyond a few polished demonstrations. They sampled 100 computational-biology papers from bioRxiv without filtering for documentation quality, repository maintenance or code completeness. Seventy-four were successfully converted. The system proposed 599 tools and 593 passed automated validation.
That’s an impressive result. The 26 failures may be more revealing.
Twenty-Six Papers Didn’t Survive the Build
The failed conversions included papers with missing executable code, unavailable data or model files, dependency and environment failures and scripts that couldn’t generalize beyond the original experiment.
Paper2Agent didn’t create those problems. It encountered them.
A paper can look complete to a reader while still failing a more demanding test: can a system unfamiliar with the research reconstruct enough of the environment and workflow to run the method without the original researchers quietly filling in the gaps?
That makes the conversion process useful as a reproducibility probe. The authors make this point directly. They suggest that the ease with which a paper becomes an agent could become a practical measure of reproducibility.
But it isn’t a universal measure of scientific quality. The study sampled computational papers with public artifacts. A theoretical proof, an ethnography or a wet-lab result with no computational method shouldn’t be judged by whether it can become an MCP server.
Even within computational research, success means something narrower than it first appears.
It means the generated tool can reproduce the source repository’s expected behavior closely enough to pass Paper2Agent’s validation tests.
It doesn’t mean the method is correct.
The Tool Can Be Faithful and the Conclusion Can Still Be Wrong
This is the distinction I’d want attached to every executable paper.
Paper2Agent validates generated tools against reference outputs from the original implementation. That’s exactly the right test for whether the wrapper faithfully executes the underlying code.
It isn’t an independent validation of the code, the statistical assumptions or the scientific conclusion.
If the original implementation contains a mistake and the generated tool reproduces it perfectly, the tool has passed a fidelity test. It hasn’t passed a truth test.
The authors acknowledge the boundary. For open-ended scientific questions, they say benchmark agreement should be read primarily as faithful execution rather than analytical validity. Researchers still have to choose the scientific direction and evaluate the evidence.
This is familiar territory. In Receipts Everywhere, I argued that a receipt can prove who signed something, when they signed it and whether it changed. It can’t prove the underlying claim is true.
An executable paper has the same split. A validation log can show that a tool produced the expected output in a defined environment. It can’t tell you that the expected output deserves your confidence.
That isn’t a criticism of Paper2Agent. It’s a reason to be precise about what the new artifact gives us.
A Citation Is Not a Build Manifest
The DOI points to a version of record. The executable artifact has more moving parts.
The codebase can change. Dependencies can change. A dataset can be revised. An external API can return different results. A maintenance agent can repair a broken path or replace a deprecated function. Any one of those changes may be reasonable, but the service running today may no longer be the same build that was validated when the paper was published.
The Paper2Agent authors treat maintenance as an inherent part of executable research. They’re right. A useful paper agent can’t remain frozen while the software around it decays.
But maintenance creates a second clock.
The paper has a publication date and a version of record. The executable artifact has releases, dependency changes, repairs and new validation runs.
If I use the artifact in a later analysis, the citation alone leaves several questions unanswered:
Which paper version did the tool represent?
Which repository commit or release supplied the code?
Which generated tool version ran?
Which environment and dependencies were installed?
Which data version went in?
Which validation tests passed, with what tolerances?
What changed after the tool was first validated?
Paper2Agent already supplies part of this chain through source-code references, isolated environments, reference outputs and automated test records. The next step is to make the complete execution record portable with the result.
That record could look less like a bibliography entry and more like a build manifest: stable identifiers for the paper, source code, generated tool, environment, input data and validation run, plus a history of any repairs.
Without it, two researchers can cite the same paper, invoke what appears to be the same agent and still run materially different artifacts.
The Method Now Has a Runtime
I recently wrote about laboratory equipment becoming accessible through a common agent interface. Paper2Agent moves the same interface idea from the equipment to the published method. The agent can now call the method itself.
That will make scientific software easier to use. It may also make weak research packages easier to identify because the conversion pipeline has to find the missing file, broken dependency or undocumented assumption that a reader might never see.
The risk is that the clean interface hides the machinery again.
A researcher asks a question. The agent selects a tool. The server runs the method. A result appears. The path from publication to output may feel shorter, but it now includes a generated wrapper, a runtime environment and a maintenance history.
Those layers don’t make the result less scientific. They become part of the record another researcher needs to reproduce it.
The institution of citation was built to identify an intellectual source. Executable research asks it to do another job: identify the exact computational object that produced a result.
One reference may not be able to do both.
Cite the paper for the idea. Record the build for the result.
Resources
Jiacheng Miao et al., “Reimagining research papers as interactive and reliable AI agents”, Nature, September 16, 2026.
Paper2Agent source repository, including the framework, generated examples and supplementary material.
Paper2Agent supplementary note, including large-scale evaluation details, ablations and maintenance considerations.
Source note: Paper2Agent’s 74-of-100 result comes from computational-biology papers sampled retrospectively from bioRxiv’s bioinformatics category. It isn’t an estimate of how much of science can be converted into agents. The comparison benchmarks measure the tested systems and tasks described in the paper, not scientific validity in general.
Analytical note: the proposed build-manifest requirements are my extension of the paper’s architecture and maintenance discussion. The Paper2Agent paper describes source references, isolated environments, validation records and ongoing maintenance, but it doesn’t present the complete manifest proposed here as a current feature.

