CoolFace
Datasetpublic

wanlilll/WeaveBench

WeaveBench A long-horizon, real-world benchmark for computer-use agents with hybrid GUI + CLI + code interfaces. ๐ŸŽŠ Accepted to EMNLP 2026 Main Conference โ€” see you in Budapest! ๐Ÿ“„ Paper: arXiv:2606.09426 ๐Ÿ’ป Code: github.com/weavebench/WeaveBench ๐ŸŒ Website: weavebench.github.io 114 long-horizon, real-world tasks across 8 work domains, where every task requires the agent to interleave GUI clicks with shell/code in one trajectory. Scored by a trajectory-aware Agent-as-Judgeโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/wanlilll/WeaveBench.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
8likes2.5kdownloads

wanlilll/WeaveBench ยท main ยท files are served by the source, never re-hosted here