Run Autonomous Science Bench
How to run Autonomous Science Bench 0.1
01Install
PlaceholderRequirements and the harness Autonomous Science Bench runs in.
shell
# Placeholder: install the harness
<install asbench>02Choose a task
The pilot has 4 tasks. Each is a campaign with fixed rounds, batch sizes, and budget, scored on its own metrics.
- Protein active learning
protein-active-learningNDCG@50 · Precision@50 - Propylene active learning
propylene-active-learningNDCG@20 · Precision@20 - Inverse lithography
inverse-lithographyNormalized XOR · Morphology change - Sparse defect scan
sparse-defect-scanDefect F1 · Full-field NRMSE
03Choose a backend
The same task runs on any backend; only what answers the agent’s experiments changes. See the roadmap.
- v0
--backend replayReplay oracleRecorded experiments, replayed exactly. The pilot results use it. - v1
--backend twinDigital twinA simulator or model, with noise and failures. Planned. - v2
--backend liveLive PCLReal instruments on a cloud lab node. Planned.
04Run a campaign
PlaceholderSupported agents and models, and how trials and seeds are set.
shell
# Placeholder: run one campaign
asbench run \
--task protein-active-learning \
--backend replay \
--agent <agent> \
--model <provider/model> \
--trials <n>05Analyze results
PlaceholderViewing a campaign trajectory, round by round, and the metrics it reports.
shell
# Placeholder: inspect a finished campaign
asbench view <run-dir>06Submit results
PlaceholderWhat a leaderboard submission must include: the full model–agent–harness configuration, seeds, and matched starting batches.
shell
# Placeholder: submit results to the leaderboard
asbench submit <run-dir>07Join our community
If you enjoy the benchmark or have feedback on how to improve it, join the community (link coming soon) and let us know.
08Citation
If you find this work useful, please cite it.
BibTeX
@misc{autonomoussciencebench2026,
title={Autonomous Science Bench: Evaluating AI for Autonomous Scientific Discovery},
author={{Autonomous Science Bench Team}},
year={2026},
url={https://yibow.me/autonomous-science-bench},
}