AI-assisted reproducibility

Research,
reproduced.

We reconstruct and execute published computational workflows, then deliver an auditable account of what reproduced, what did not, and why.

B Reproduction report
BR-0042

PAPER

Scaling behavior in sparse
computational systems

Submitted artifact · 12 Mar 2025

8/ 10

REPORTED RESULTS REPRODUCED

Core findings verified

2 results blocked by missing assumptions

✓
Figure 2Matched within reported tolerance
REPRODUCED
✓
Table 1All metrics independently verified
REPRODUCED
!
Figure 4Random seed not specified
NEEDS INPUT

From publication to execution.

A paper is a claim.
We turn it into evidence.

01 / THE WORKFLOW

Reproduction,
end to end.

Bullet Research’s AI agents read, reconstruct, execute, and document the full computational workflow under expert supervision.

01

Ingest the work

We analyze the paper, code repository, public data, and stated dependencies as one connected artifact.

02

Rebuild & execute

We reconstruct the environment, resolve assumptions, and run the reported analyses in an isolated workspace.

03

Verify the claims

Reported metrics and figures are compared against fresh outputs, with tolerances and deviations recorded.

04

Deliver the record

Authors receive a private runnable workflow and a precise report of results, blockers, and missing assumptions.

02 / BUILT FOR TRUST

More signal for authors.
Less friction for reviewers.

Private by default

Reports and runnable workflows stay with authors and organizers. Nothing is publicly scored.

Fully auditable

Every environment, command, output, deviation, and intervention is recorded.

Human in the loop

AI accelerates the work. Researchers remain in control of interpretation and judgment.

Pilot program

Make reproducibility
part of the work.

We’re partnering with a small number of reproducibility-focused workshops and artifact-evaluation groups for a free pilot of up to 10 voluntary submissions.

  • Private runnable workflow for every author
  • Detailed, auditable reproduction report
  • Anonymized aggregate findings for organizers
  • No automated acceptance decisions or public scoring
Discuss a pilot