Scientific ML Agent Benchmark

NatureBench

Can coding agents match the published SOTA of Nature-family papers?

We evaluate frontier coding-agent configurations on 90 Nature-family scientific ML tasks in isolated containers.

Leaderboard

Metric Definitions

g
SOTA-normalized relative gap, g = dir * (m - m_sota) / |m_sota|; dir handles metric direction.
Surpass-SOTA
Share of judge-valid tasks with g > 0.1.
Match-SOTA
Share of judge-valid tasks with g >= 0; g = 0 matches paper SOTA.
Median g (all)
Median g across all tasks, assigning g = -1 to judge-invalid and no-score tasks.
CR
Completion Rate, the valid-score rate.
SR
Score Rate, the any-score rate.

Main Leaderboard

Ranked by Surpass-SOTA, then Match-SOTA; equal values share a rank.

Click the beside a model to view its evaluation setup.

All 90 NatureBench tasks for full-benchmark performance reporting.

90
Tasks
6
Scientific domains
23.3%
Best Surpass-SOTA
57.8%
Best Match-SOTA
91.1%
Top completion
Rank Model Agent Run source Surpass-SOTA Match-SOTA Median g (all) CR SR Invalid

Score Distribution

Full track · 90 tasks

Task-level g bins per configuration

Domain Breakdown

Full track · 90 tasks

Domain labels follow the six-domain taxonomy used in the NatureBench paper. Within each domain, configurations are ranked by Surpass-SOTA, then Match-SOTA, as on the main leaderboard.

Domain Ranking

Domain Winners

Best configuration within each scientific domain

Domain N Winner Surpass-SOTA Match-SOTA Median g (all)

Case-Level View

Full track · 90 tasks

Each cell shows task-level normalized relative gap g. Values above 0.1 count as Surpass-SOTA; nonnegative values match the paper's reported SOTA.

Surpass-SOTA, g > 0.1 Match-SOTA, 0 <= g <= 0.1 gray/white: below SOTA, g < 0(darker = farther below) Invalid No score/submission Outlined cell: best configuration for that task

Per-Case Score Matrix

Task name · Case ID · Domain · ML task type · Best model · Best harness · Best harness + model

Submit Results

We welcome community submissions to the NatureBench leaderboard.

Cite NatureBench

If you use NatureBench in your research, please cite our work.

BibTeX
@misc{wang2026naturebench,
  title         = {NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?},
  author        = {Yuru Wang and Lejun Cheng and Yuxin Zuo and Sihang Zeng and Bingxiang He and Che Jiang and Junlin Yang and Yuchong Wang and Kaikai Zhao and Weifeng Huang and Kai Tian and Zhenzhao Yuan and Jincheng Zhong and Weizhi Wang and Ning Ding and Bowen Zhou and Kaiyan Zhang},
  year          = {2026},
  eprint        = {2606.24530},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2606.24530}
}