CoolFace
Datasetpublic

ByteDance-Seed/EdgeBench

Overview EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/EdgeBench.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
84likes7.5kdownloads
Dataset Card

<p align="center"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/ByteDance-Seed/EdgeBench/resolve/main/assets/logo-dark.png"> <img src="assets/logo.jpg" alt="ByteDance Seed" width="420"> </picture> </p>

<p align="center"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/ByteDance-Seed/EdgeBench/resolve/main/assets/title-dark.svg"> <img src="assets/title.svg" alt="EdgeBench" width="280"> </picture> </p>

<br>

<!-- <p align="center"> <strong>Measuring How AI Agents Learn from Real-World Environments</strong> </p>

<p align="center"> 134 real-world tasks &nbsp;|&nbsp; 51 open-source &nbsp;|&nbsp; 6 capability categories &nbsp;|&nbsp; 38,000+ hours of agent interaction </p> -->

<p align="center"> <a href="https://edge-bench.org/"><img src="https://img.shields.io/badge/Project-edge--bench.org-blue" alt="Project"></a> <a href="https://arxiv.org/abs/2607.05155"><img src="https://img.shields.io/badge/Tech%20Report-PDF-red?logo=adobeacrobatreader" alt="Tech Report"></a> <a href="https://github.com/ByteDance-Seed/EdgeBench"><img src="https://img.shields.io/badge/GitHub-EdgeBench-green?logo=github" alt="GitHub"></a> <a href="https://bytedance-seed.github.io/EdgeBench/"><img src="https://img.shields.io/badge/Docs-SForge%20Harness-purple" alt="Docs"></a> <a href="https://github.com/ByteDance-Seed/EdgeBench/blob/main/assets/wechat_qr.jpg"><img src="https://img.shields.io/badge/WeChat-Group-07C160?logo=wechat&logoColor=white" alt="WeChat Group"></a> <a href="https://discord.gg/p2JZ26ku8"><img src="https://img.shields.io/badge/Discord-Join-5865F2?logo=discord&logoColor=white" alt="Discord"></a> </p>


Overview

EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks along with the full evaluation framework.

Analyzing ~38,000 hours of agent interaction on all 134 tasks, we find that performance follows a log-sigmoid scaling law as a function of interaction time ($R^2 = 0.998$). See the tech report for details.

<p align="center"> <img src="assets/figfull136curvefitsideby_side.png" alt="Log-sigmoid scaling fit across 134 tasks" width="800"> </p>

Leaderboard

Full Benchmark (134 tasks)

Model@2h@4h@6h@8h@10h**@12h**
Claude Opus 4.839.045.748.149.850.951.3
GPT-5.536.842.144.546.347.648.4
GPT-5.429.734.036.538.038.939.3
GLM-5.126.030.432.934.936.537.4
DS-V4-Pro23.327.129.029.930.931.0

<details> <summary><b>Category Scores @12h (134 tasks)</b></summary>

ModelScientific & MLSystems & SEOptimizationKnowledgeFormalGames
Claude Opus 4.848.567.436.547.055.039.3
GPT-5.544.365.033.645.750.039.1
GPT-5.433.554.127.938.840.829.0
GLM-5.133.850.926.443.524.629.3
DS-V4-Pro30.043.021.537.014.116.9

</details>

Open-Source Subset (51 tasks)

Model@2h@4h@6h@8h@10h**@12h**
Claude Opus 4.833.238.540.842.143.344.2
GPT-5.531.236.038.240.342.143.1
GPT-5.425.028.230.332.133.334.2
GLM-5.121.424.226.828.229.130.4
DS-V4-Pro17.121.122.923.825.125.7

<details> <summary><b>Category Scores @12h (51 tasks)</b></summary>

ModelScientific & MLSystems & SEOptimizationKnowledgeFormalGames
Claude Opus 4.838.962.038.238.740.939.3
GPT-5.533.260.532.338.449.039.1
GPT-5.424.650.129.931.630.229.0
GLM-5.126.843.626.731.019.929.3
DS-V4-Pro31.137.624.133.212.716.9

</details>

<details> <summary><b>Per-Task Scores by Time Budget (51 tasks)</b></summary>

Each model cell reports scores at @2h / @4h / @6h / @8h / @10h / @12h. Missing valid results are shown as —.

TaskCategoryOpus 4.8GPT-5.5GPT-5.4GLM-5.1DS-V4-Pro
bipedalwalkerlocomotionrlScientific & ML16.7/20.8/22.4/23.3/23.3/23.314.7/14.9/15.2/15.2/16.0/21.013.9/13.9/13.9/14.5/14.5/17.513.9/20.3/21.5/22.5/22.5/22.58.9/14.8/17.6/20.4/20.4/20.6
bordensourceinversionScientific & ML7.5/19.8/26.2/28.5/38.5/48.420.1/27.0/29.4/37.8/38.1/38.57.2/7.3/7.6/7.9/8.0/8.07.0/10.3/12.0/12.3/12.5/15.17.0/11.6/15.1/26.6/36.7/38.2
dabicgravityinversionScientific & ML9.5/15.2/15.7/17.4/17.5/17.515.9/16.2/16.7/17.0/17.2/17.314.6/14.6/15.5/15.5/15.0/15.09.2/13.7/16.0/16.5/16.5/17.1—/12.7/12.7/12.7/13.0/13.8
graphnodeclassificationScientific & ML59.4/62.7/65.0/65.6/66.5/66.654.7/55.1/55.1/55.3/55.9/56.054.9/56.2/56.5/56.9/57.5/57.649.4/52.3/52.3/52.3/52.3/52.346.0/48.2/49.2/51.3/51.7/51.8
annvectorsearch_qpsSystems & SE26.2/57.0/58.6/58.7/59.4/59.722.3/34.3/35.1/36.0/40.0/40.727.5/30.2/44.5/45.2/49.7/50.26.7/24.4/25.6/25.6/26.1/38.39.4/19.6/22.4/22.8/23.8/23.8
arccompilerruntimeSystems & SE49.3/52.0/52.0/52.0/52.0/52.055.5/56.5/60.9/70.3/71.0/72.445.1/46.5/49.8/49.8/50.0/50.047.7/48.0/48.4/48.7/48.7/48.740.3/41.7/44.2/44.2/44.2/44.2
exchangecorethroughputSystems & SE40.7/57.0/58.5/58.9/59.7/59.715.4/37.2/39.9/44.3/51.3/53.214.3/40.8/41.0/45.2/46.4/47.329.2/43.7/46.5/48.6/50.3/52.632.9/33.8/45.0/47.7/48.4/48.6
ffmpegswscalereimplementationSystems & SE9.9/17.6/19.8/20.9/21.1/21.18.8/14.3/15.1/15.3/15.3/15.35.4/8.5/9.4/11.6/13.3/13.90.3/0.3/0.4/2.2/2.2/2.20.1/1.9/2.0/3.8/3.8/3.8
gitrewritein_zigSystems & SE22.0/22.8/22.8/22.8/23.1/23.116.1/16.9/17.7/18.2/18.2/18.49.6/13.8/14.0/14.2/14.2/15.412.0/20.2/23.3/23.4/23.4/23.58.5/13.5/16.0/17.6/17.8/17.9
integercompressioncodecSystems & SE69.4/69.7/74.8/74.9/75.2/75.361.1/67.6/73.9/73.9/74.3/74.438.6/40.9/41.2/42.2/42.2/42.323.5/27.3/28.5/28.7/28.9/28.915.9/16.0/16.2/16.2/16.2/16.2
julietvulnerabilityanalyzerSystems & SE71.9/74.9/75.4/75.6/75.6/75.681.0/83.2/85.4/86.8/87.4/89.852.9/66.1/74.3/76.0/76.8/77.259.3/60.7/62.8/63.5/63.5/63.546.8/63.1/66.1/66.2/66.2/66.2
rustmulticratereconstructionSystems & SE—/—/—/—/—/—27.5/42.6/53.1/54.9/57.8/57.816.7/19.9/21.3/21.4/21.4/21.424.8/24.8/25.2/25.2/37.5/38.520.5/21.7/22.7/23.1/23.5/23.6
schemathesisconfigmodernizationSystems & SE82.5/85.0/86.1/87.4/87.4/87.779.1/82.2/82.9/83.2/83.6/84.067.2/68.8/68.8/71.7/71.7/71.958.3/59.7/60.4/61.2/61.7/61.754.3/54.3/55.3/55.3/55.3/55.6
schemathesisdatagenpipelineSystems & SE68.0/70.2/70.2/70.2/70.2/70.254.6/54.6/56.7/56.7/56.7/56.756.6/56.6/56.6/56.6/56.6/56.662.1/64.2/64.2/67.0/67.0/67.047.9/50.1/52.3/52.3/52.3/52.3
schemathesisreportingobservabilitySystems & SE73.9/75.6/76.2/76.2/76.2/76.276.6/76.6/76.6/76.6/77.1/77.170.0/73.7/74.7/75.7/76.2/76.261.9/61.9/61.9/61.9/61.9/61.959.4/62.4/63.0/63.0/65.0/65.0
vliwkerneloptimizationSystems & SE74.0/76.0/77.7/79.5/79.6/80.971.6/75.7/77.1/79.5/83.1/85.675.7/77.0/77.2/78.7/79.1/79.15.6/9.5/27.5/35.0/35.9/35.90.2/24.9/28.1/33.0/33.9/34.1
adplacementoptimizationOptimization65.2/66.1/66.9/67.1/67.4/67.744.0/53.3/59.5/61.6/62.9/62.941.8/42.4/43.1/47.7/47.9/48.148.7/52.7/53.3/56.5/58.5/58.825.5/28.5/35.2/35.8/36.2/36.2
appleincrementalgameOptimization42.7/44.9/45.9/48.6/49.9/50.626.6/29.8/30.6/32.7/33.1/33.628.3/30.3/32.0/33.3/33.9/34.919.0/19.0/19.1/19.1/19.1/19.119.6/19.7/19.7/19.7/19.7/19.7
equivalenceclassdivideandconquerOptimization11.2/15.3/17.0/20.1/20.8/21.311.8/15.5/15.8/21.3/22.2/22.414.5/17.0/18.3/18.7/20.2/20.33.8/4.2/10.0/8.0/10.6/10.60.7/1.8/3.2/3.2/3.4/3.4
gridturingrobotOptimization34.7/37.1/37.3/39.6/40.3/40.340.4/41.6/41.9/42.0/42.1/42.226.8/26.8/27.2/28.9/28.9/28.920.0/21.0/24.6/24.6/24.6/25.723.7/24.1/24.1/24.1/24.2/24.2
jaguanestingoptimizationOptimization11.2/17.8/24.5/31.4/41.0/44.215.9/19.4/20.0/20.6/21.3/21.622.4/23.0/23.9/24.0/24.1/24.18.9/9.0/10.0/12.2/12.3/12.410.7/20.2/23.7/26.7/28.1/28.4
molecularselfassemblyOptimization22.4/33.4/34.0/34.1/34.4/34.720.2/20.3/20.5/20.7/20.7/20.720.8/21.1/21.1/21.5/21.5/21.610.0/12.5/12.9/13.0/13.1/13.219.4/21.7/21.8/21.8/21.9/21.9
orderadditionpermutation_optimizationOptimization22.6/31.6/34.0/34.4/35.7/36.416.7/20.5/21.5/22.4/23.0/23.31.6/10.6/13.1/14.0/14.2/14.32.0/2.1/23.6/25.8/25.8/33.24.6/16.5/17.8/22.9/25.4/30.8
smt_solverOptimization10.3/17.4/19.0/23.1/23.3/23.97.2/7.8/8.4/8.6/8.6/8.66.7/7.9/8.9/9.1/9.1/9.22.7/2.7/2.7/2.7/2.7/3.61.4/2.8/3.3/3.3/3.3/3.3
treant_forestOptimization14.5/15.9/16.1/16.2/16.4/18.012.1/14.2/14.9/15.2/15.5/15.612.2/12.2/12.7/13.0/13.2/13.38.0/11.6/11.7/14.1/14.5/16.96.8/8.1/9.7/10.1/12.7/13.5
treeblockpartitioningOptimization21.5/30.1/32.4/36.8/37.7/37.728.8/31.1/33.0/33.0/35.0/36.423.1/26.8/28.8/32.9/34.3/34.312.1/15.4/17.1/19.3/20.3/23.411.2/11.8/11.9/11.9/14.6/16.1
triangulationcoloringoptimizationOptimization70.8/71.4/71.9/73.2/73.3/73.473.7/74.3/74.5/75.0/75.1/75.274.1/74.2/74.3/74.3/74.3/74.368.8/71.2/71.6/72.0/72.7/73.056.1/58.0/59.0/59.1/59.1/59.3
vehicleroutingtime_windowsOptimization72.5/72.6/72.9/73.6/73.7/74.088.7/89.0/89.4/89.7/89.7/90.885.3/88.6/88.7/89.5/89.5/89.676.6/76.6/76.6/76.6/76.6/77.954.7/76.8/81.9/82.2/82.9/83.1
vibratingpathgraph_coloringOptimization19.7/21.1/21.4/22.5/24.5/25.310.1/10.5/10.6/10.7/10.7/11.418.1/19.4/19.8/23.4/23.6/24.19.6/18.3/20.3/22.9/22.9/22.912.4/14.4/19.3/19.4/21.8/22.1
warehouseforkliftroutingOptimization7.7/9.5/10.4/10.5/11.1/11.29.8/11.0/11.8/11.9/12.1/12.60.0/0.0/0.0/0.0/0.0/0.0—/0.0/0.0/0.6/0.7/0.50.0/0.0/0.0/0.0/0.0/0.0
wirelesselectricitylayoutOptimization6.5/13.7/14.4/14.5/14.5/14.56.2/6.9/7.1/7.1/7.1/7.210.9/11.1/11.1/11.1/11.1/11.17.2/9.4/6.6/8.1/9.4/9.50.0/0.0/0.0/0.0/0.0/0.0
collegeenglishexam_bankKnowledge24.8/28.3/34.8/35.5/35.8/39.824.5/35.5/35.5/35.5/37.8/37.830.7/30.7/31.3/34.0/34.0/34.522.2/26.0/29.3/30.0/32.3/32.519.2/21.7/22.5/22.7/29.2/34.7
ctariskbudget_optimizationKnowledge42.7/44.8/45.3/45.3/45.3/46.143.8/45.8/46.7/46.7/46.7/46.746.0/49.0/49.0/49.0/49.8/49.838.1/44.8/49.0/49.6/49.6/49.644.0/45.6/46.9/46.9/48.1/48.1
k12mathrecommendationKnowledge23.6/38.5/41.4/42.0/43.7/44.338.5/42.4/42.9/43.5/43.9/44.025.9/29.0/30.0/30.8/31.1/31.424.8/25.7/31.9/32.5/32.7/32.725.6/26.3/26.8/25.7/26.0/26.3
portfolioriskcalibrationKnowledge20.1/21.6/23.0/23.6/23.6/24.517.3/21.3/22.7/23.5/24.4/25.06.0/9.6/10.7/10.7/10.7/10.70.0/8.4/8.5/8.9/9.2/9.410.4/16.3/16.6/16.7/23.7/23.7
carleson_formalizationFormal4.3/7.7/11.0/12.7/15.0/16.86.0/9.5/13.2/16.5/25.3/26.51.8/3.5/4.6/5.4/6.3/7.11.0/1.7/2.0/2.2/2.2/2.20.8/1.3/2.0/2.0/2.3/2.5
combinatorialgamesformalizationFormal14.5/23.2/27.6/32.1/34.5/35.512.0/18.8/24.6/27.2/33.4/38.25.9/8.3/11.5/13.5/16.3/17.86.7/9.8/14.3/14.9/16.2/16.24.3/6.7/7.3/7.4/7.7/7.8
fltregularformalizationFormal31.0/41.8/50.6/50.6/50.6/50.643.7/48.3/50.6/66.7/75.1/75.11.5/19.5/28.4/41.8/46.0/48.314.4/13.4/16.5/18.8/18.8/38.75.7/11.9/14.6/14.9/17.2/17.6
leananalysisproofsFormal17.9/25.1/28.6/30.2/32.6/33.016.8/23.2/28.4/33.9/39.0/42.53.6/8.1/10.8/12.9/15.5/16.45.2/5.9/5.9/5.9/5.9/5.95.8/7.3/8.2/8.8/9.3/9.5
newfoundationsconsistencyFormal28.9/36.2/50.0/62.7/64.2/65.113.7/38.2/55.1/56.4/65.1/66.53.3/12.2/14.9/20.5/30.7/39.82.2/3.3/5.1/21.9/24.6/27.02.2/3.4/6.5/7.2/10.5/11.4
ordinalnotationwell_foundednessFormal10.6/18.4/24.7/24.7/24.7/24.713.7/24.7/24.7/24.7/24.7/24.71.2/5.5/13.7/15.3/18.4/21.62.0/3.5/5.1/5.9/5.9/5.93.5/3.5/4.7/4.7/4.7/4.7
pfr_formalizationFormal32.4/36.9/38.8/40.2/45.6/46.330.7/38.3/41.9/47.6/52.7/60.010.2/13.7/27.1/34.0/35.9/38.98.3/14.9/22.5/26.5/31.3/33.59.9/14.7/16.5/17.8/18.5/19.1
sphereeversionformalizationFormal41.7/47.4/49.1/50.4/54.1/55.445.0/51.1/55.0/55.9/56.9/58.513.3/20.2/32.7/43.7/50.4/51.415.5/24.2/26.7/28.7/30.2/30.22.9/14.1/22.3/24.9/28.6/29.3
anchorheadtextadventureGames13.3/19.3/19.7/20.3/22.3/22.315.0/26.3/31.7/34.3/35.3/36.35.0/11.7/13.0/13.3/14.7/17.710.7/17.3/19.7/20.3/20.3/20.32.0/6.0/7.3/8.0/12.3/14.7
dcssdungeonaiGames4.2/4.9/5.9/6.3/6.3/8.38.9/9.7/10.0/10.0/13.3/13.42.6/5.6/5.6/5.6/6.1/6.12.8/3.0/3.3/3.3/5.1/7.62.8/3.6/4.4/4.5/5.1/5.7
nethackdungeonagentGames29.7/35.3/36.7/37.3/41.9/41.916.6/17.6/18.1/20.6/21.3/22.510.9/14.1/15.2/15.8/17.0/20.42.3/2.3/15.3/21.6/21.6/21.61.0/1.4/2.9/3.2/3.2/3.3
openrct2themepark_aiGames24.4/24.4/26.0/26.0/27.5/27.528.5/28.6/32.7/37.3/37.4/37.623.0/23.1/23.1/23.1/23.1/23.135.1/36.2/36.2/36.2/36.2/36.224.4/24.4/24.4/26.0/26.0/26.0
openttdtransportaiGames50.0/50.4/50.6/51.7/51.8/52.010.1/11.6/13.2/21.9/25.6/28.110.8/11.4/11.6/11.9/11.9/11.90.0/0.0/0.0/0.0/0.0/0.04.8/9.2/9.3/12.3/15.2/15.2
trinitytextadventureGames25.0/28.0/29.3/30.0/30.0/30.022.3/26.7/28.7/36.3/36.3/40.016.3/19.7/22.7/23.3/23.7/27.016.0/18.7/24.3/26.0/26.7/26.716.3/16.3/17.7/20.0/20.0/20.3
trysttextadventureGames18.1/33.8/36.7/40.0/40.0/44.332.1/42.4/44.3/48.6/55.2/55.719.5/20.0/20.0/31.0/38.6/44.318.6/28.6/36.2/40.5/42.9/43.38.6/11.4/11.4/11.4/11.4/13.8
wesnothtacticalaiGames84.0/85.3/87.7/87.7/87.7/88.064.7/73.0/76.3/78.0/78.3/79.379.7/79.7/80.3/80.3/81.3/81.375.7/78.3/78.3/78.3/78.3/78.317.0/36.3/36.3/36.3/36.3/36.3

</details>

Task Taxonomy

EdgeBench contains 134 realistic, diverse tasks spanning six capability categories, of which 51 are publicly released. Each task is designed as a day-scale challenge with a performance ceiling high enough that no current agent can saturate it. Recorded human expert effort averages 57.2 hours per task (up to 320 hours).

<p align="center"> <img src="assets/edgebench_taxonomy.png" alt="EdgeBench Task Taxonomy" width="850"> </p>

Evaluation Harness: SForge

EdgeBench is powered by **SForge**, a two-container evaluation harness built for long-horizon agent evaluation. See the SForge documentation for setup and usage instructions.

Citation

If you find EdgeBench useful in your research, please cite our tech report:

bibtex
@misc{edgebench2026,
  title  = {EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments},
  author = {Deyao Zhu and Xin Zhou and Shengling Qin and Xuekai Zhu and Hangliang Ding and Shu Zhong and others},
  year   = {2026},
  url    = {https://arxiv.org/abs/2607.05155},
}

License

EdgeBench task datasets are released under CC BY 4.0.

Contact

To evaluate on the full 134-task suite, please contact zhongshu@bytedance.com.

<p align="center"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/ByteDance-Seed/EdgeBench/resolve/main/assets/logo-dark.png"> <img src="assets/logo.jpg" alt="ByteDance Seed" width="200"> </picture> <br> <sub>Built by <a href="https://github.com/ByteDance-Seed">ByteDance Seed</a></sub> </p>