Autonomous Science Bench0.1

Autonomous Science Bench evaluates AI agents for autonomous scientific discovery on real workflows across programmable cloud laboratories.

Written by Yibo Wen on

Autonomous Science Bench is a benchmark for AI systems that run autonomous laboratories. Its tasks come from the experimental workflows of 20 programmable cloud lab nodes across the United States, and it scores the whole system, from model and agent to tools and harness, on how well it plans experiments, interprets measurements, and improves its decisions over time. This first release sets a common evaluation framework and four pilot tasks.

Autonomous Science Bench 0.1 leaderboard

Autonomous ScienceBench 0.1
  1. GPT-5.6 TerraCodex
  2. GLM-5.3 Flashmini-SWE-agent
  3. Qwen3.8 27Bmini-SWE-agent
  4. Sonnet 5.5Claude Code
  5. GPT-5.6 LunaCodex
  6. Haiku 4.5Claude Code
  7. MiniMax M3mini-SWE-agent
Resolution rates across 4 pilot tasks on Autonomous Science Bench 0.1

Overview

Autonomous laboratories run planning, automation, and measurement as one continuous loop. Being useful in that loop takes more than accurate prediction. An agent must:

  • Choose the evidence: decide which experiments are worth running.
  • Explore or exploit: know when to search and when to commit.
  • Adapt: respond when outcomes defy expectations.
  • Know the limits: recognize when the data cannot support a conclusion.
  • Respect the lab: stay within its operations, costs, and measurement limits.

Autonomous Science Bench tests this with bounded campaigns, each a scientific objective, an experimental interface, and a fixed budget. Success is what the agent achieves with them, judged by the nodes’ own scientists.

A shared task format

Every task is a design–build–test–learn loop. From a provided starting batch, the agent designs experiments, receives measurements, and adapts, round after round.

01Starting batch123456ABCDE02Agent picks a batch123456ABCDE03Lab returns results123456ABCDE04Agent learnsafterbeforecandidatesvalueinitial datarun on cloud labresults to agent↻ next batch · repeat for R rounds
  • Untested well
  • Measured, low to high value
  • Picked by the agent
  • Model before and after the new results
One campaign. The agent starts from a provided batch of measured wells, then runs a closed loop for a fixed number of rounds: pick a batch, receive its measurements (recorded, or run by a cloud lab in prospective campaigns), refit its model, and pick the next batch. Here the new results (ringed) lift the model’s estimate to a peak the starting batch missed.

Every task specifies

The rules of the campaign. How the agent plays within them is up to it.

  • Available actions and feedback
  • Rounds, batch sizes, and budgets
  • Success criteria

Final evaluation

After the last round, the agent submits one of these for independent verification.

  • A ranking of candidates
  • A predictive model
  • Final designs

Scientific Tasks

Tasks come from the nodes’ own research workflows and are built with their domain experts. They test optimization, model identification, experimental design, and trade-offs between objectives, and each ties the agent’s decisions to an independently measured outcome.

  • Biology

    Protein variants, cell-free expression, and growth conditions.

    PilotProtein active learning

  • Chemistry

    Reaction conditions, catalysts, yield, and selectivity.

    PilotPropylene active learning

  • Materials

    Compositions and processing that reach target properties.

    PilotSparse defect scan

  • Semiconductors & electronics

    Fabrication recipes, device behavior, and consistency.

    PilotInverse lithography

The 20-node network sets the scope of Autonomous Science Bench; coverage grows as each workflow is validated and approved.

PCL Test Bed nodes
20
subawards to partner institutes
47
institutional collaborators
409
industry partners
201
DOE national labs
3
  • PCL node, marked with its primary domain
  • Subaward to a partner institution
  • Concentration of nodes and partners

Biology3

Chemistry3

Materials8

Semiconductors & electronics4

Cross-domain2

The 20 nodes of the NSF PCL Test Bed, which set the scope of Autonomous Science Bench. Each marker is a node at its lead institution; arcs trace its subawards to partner institutions, and the glow shows where the network concentrates. Hover or tap a node to follow its partners.Source: NSF PCL Test Bed dashboard (data as of August 21, 2026) and NSF award records. Primary domains are our reading of each award.

Results

The pilot leaves substantial room for progress. Seven configurations ran four tasks across biology, chemistry, materials, and semiconductors, with one trial per task. No configuration solves every task. GPT-5.6 Terra with Codex leads at 75%, followed by GLM-5.3 Flash at 50% and Sonnet 5.5 and Qwen3.8 27B at 25%. Haiku 4.5, GPT-5.6 Luna, and MiniMax M3 resolve none. With so few trials, small gaps between configurations are not yet meaningful.

Autonomous Science Bench 0.1 cost vs. resolution rate

0%20%40%60%80%100%$0$5$10$15$20$25Total cost (USD)Resolution rateGPT-5.6 Terra (max), Codex: 75% (3/4) at $21.40Haiku 4.5, Claude Code: 0% (0/4) at $1.75Sonnet 5.5 (high), Claude Code: 25% (1/4) at $2.54GPT-5.6 Luna (max), Codex: 0% (0/4) at $3.07MiniMax M3, mini-SWE-agent: 0% (0/4) at $9.96Qwen3.8 27B, mini-SWE-agent: 25% (1/4) at $9.76GLM-5.3 Flash, mini-SWE-agent: 50% (2/4) at $6.59GPT-5.6 TerraCodexHaiku 4.5Claude CodeSonnet 5.5Claude CodeGPT-5.6 LunaCodexMiniMax M3mini-SWE-agentQwen3.8 27Bmini-SWE-agentGLM-5.3 Flashmini-SWE-agent
Each square is one configuration, pooled over every pilot trial; cost is its total across them. The line traces the cost frontier, the best resolution rate reached at each budget.

How results are scored

Each task scores progress during the campaign and the quality of its final outcome, against simple baselines, under matched starting conditions and budgets. Results name the full model–agent–harness configuration.

During the campaign

  • Improvement over the starting batch
  • Experiments or cost to reach a target

Final outcome

  • Ranking quality or predictive accuracy
  • Constraints met across objectives

Compared against

  • Random selection
  • Bayesian optimization
  • Node reference workflows

Roadmap

Each task can run on three backends. The agent’s interface never changes; only what answers its experiments gets closer to the real lab.

  1. v0
    Replay oracleResults come from recorded experiments, replayed exactly each time.
  2. v1
    Digital twinResults come from a simulator or model, with noise, limits, and failures.
  3. v2
    Live PCLResults come from real instruments, with real queues, costs, and faults.
Same tasks · same agent APIand more physical reality at each step

Built to stay trustworthy

  • Against contamination: private final evaluation, controlled data access, and fresh experiments.
  • For fair comparisons: versioned tasks, documented protocols, and expert review.

The goal: a shared foundation that shows where autonomous labs work, where they fail, and which advances lead to measurable scientific progress.

Leadership

  • NameRoleAffiliation
  • NameRoleAffiliation
  • NameRoleAffiliation

Citation

If you find this work useful, please cite it.

BibTeX
@misc{autonomoussciencebench2026,
  title={Autonomous Science Bench: Evaluating AI for Autonomous Scientific Discovery},
  author={{Autonomous Science Bench Team}},
  year={2026},
  url={https://yibow.me/autonomous-science-bench},
}

Acknowledgements

Autonomous Science Bench is developed with researchers and engineers across the programmable cloud lab community. We thank the participating nodes for their scientific expertise and experimental infrastructure, and acknowledge the broader ecosystem established through the U.S. National Science Foundation’s PCL Test Bed initiative.