Autonomous Science Bench0.1
Autonomous Science Bench evaluates AI agents for autonomous scientific discovery on real workflows across programmable cloud laboratories.
Written by Yibo Wen on
Autonomous Science Bench is a benchmark for AI systems that run autonomous laboratories. Its tasks come from the experimental workflows of 20 programmable cloud lab nodes across the United States, and it scores the whole system, from model and agent to tools and harness, on how well it plans experiments, interprets measurements, and improves its decisions over time. This first release sets a common evaluation framework and four pilot tasks.
Autonomous Science Bench 0.1 leaderboard
Autonomous ScienceBench 0.1- GPT-5.6 TerraCodex
- GLM-5.3 Flashmini-SWE-agent
- Qwen3.8 27Bmini-SWE-agent
- Sonnet 5.5Claude Code
- GPT-5.6 LunaCodex
- Haiku 4.5Claude Code
- MiniMax M3mini-SWE-agent
Overview
Autonomous laboratories run planning, automation, and measurement as one continuous loop. Being useful in that loop takes more than accurate prediction. An agent must:
- Choose the evidence: decide which experiments are worth running.
- Explore or exploit: know when to search and when to commit.
- Adapt: respond when outcomes defy expectations.
- Know the limits: recognize when the data cannot support a conclusion.
- Respect the lab: stay within its operations, costs, and measurement limits.
Autonomous Science Bench tests this with bounded campaigns, each a scientific objective, an experimental interface, and a fixed budget. Success is what the agent achieves with them, judged by the nodes’ own scientists.
A shared task format
Every task is a design–build–test–learn loop. From a provided starting batch, the agent designs experiments, receives measurements, and adapts, round after round.
- Untested well
- Measured, low to high value
- Picked by the agent
- Model before and after the new results
Every task specifies
The rules of the campaign. How the agent plays within them is up to it.
- Available actions and feedback
- Rounds, batch sizes, and budgets
- Success criteria
Final evaluation
After the last round, the agent submits one of these for independent verification.
- A ranking of candidates
- A predictive model
- Final designs
Scientific Tasks
Tasks come from the nodes’ own research workflows and are built with their domain experts. They test optimization, model identification, experimental design, and trade-offs between objectives, and each ties the agent’s decisions to an independently measured outcome.
- Biology
Protein variants, cell-free expression, and growth conditions.
PilotProtein active learning
- Chemistry
Reaction conditions, catalysts, yield, and selectivity.
PilotPropylene active learning
- Materials
Compositions and processing that reach target properties.
PilotSparse defect scan
- Semiconductors & electronics
Fabrication recipes, device behavior, and consistency.
PilotInverse lithography
The 20-node network sets the scope of Autonomous Science Bench; coverage grows as each workflow is validated and approved.
- PCL Test Bed nodes
- 20
- subawards to partner institutes
- 47
- institutional collaborators
- 409
- industry partners
- 201
- DOE national labs
- 3
- PCL node, marked with its primary domain
- Subaward to a partner institution
- Concentration of nodes and partners
Biology3
Chemistry3
Materials8
Semiconductors & electronics4
Cross-domain2
Results
The pilot leaves substantial room for progress. Seven configurations ran four tasks across biology, chemistry, materials, and semiconductors, with one trial per task. No configuration solves every task. GPT-5.6 Terra with Codex leads at 75%, followed by GLM-5.3 Flash at 50% and Sonnet 5.5 and Qwen3.8 27B at 25%. Haiku 4.5, GPT-5.6 Luna, and MiniMax M3 resolve none. With so few trials, small gaps between configurations are not yet meaningful.
Autonomous Science Bench 0.1 cost vs. resolution rate
How results are scored
Each task scores progress during the campaign and the quality of its final outcome, against simple baselines, under matched starting conditions and budgets. Results name the full model–agent–harness configuration.
During the campaign
- Improvement over the starting batch
- Experiments or cost to reach a target
Final outcome
- Ranking quality or predictive accuracy
- Constraints met across objectives
Compared against
- Random selection
- Bayesian optimization
- Node reference workflows
Roadmap
Each task can run on three backends. The agent’s interface never changes; only what answers its experiments gets closer to the real lab.
- v0Replay oracleResults come from recorded experiments, replayed exactly each time.
- v1Digital twinResults come from a simulator or model, with noise, limits, and failures.
- v2Live PCLResults come from real instruments, with real queues, costs, and faults.
Built to stay trustworthy
- Against contamination: private final evaluation, controlled data access, and fresh experiments.
- For fair comparisons: versioned tasks, documented protocols, and expert review.
The goal: a shared foundation that shows where autonomous labs work, where they fail, and which advances lead to measurable scientific progress.
Leadership
- NameRoleAffiliation
- NameRoleAffiliation
- NameRoleAffiliation
Citation
If you find this work useful, please cite it.
@misc{autonomoussciencebench2026,
title={Autonomous Science Bench: Evaluating AI for Autonomous Scientific Discovery},
author={{Autonomous Science Bench Team}},
year={2026},
url={https://yibow.me/autonomous-science-bench},
}Acknowledgements
Autonomous Science Bench is developed with researchers and engineers across the programmable cloud lab community. We thank the participating nodes for their scientific expertise and experimental infrastructure, and acknowledge the broader ecosystem established through the U.S. National Science Foundation’s PCL Test Bed initiative.