wanlilll/WeaveBench
WeaveBench A long-horizon, real-world benchmark for computer-use agents with hybrid GUI + CLI + code interfaces. ๐ Accepted to EMNLP 2026 Main Conference โ see you in Budapest! ๐ Paper: arXiv:2606.09426 ๐ป Code: github.com/weavebench/WeaveBench ๐ Website: weavebench.github.io 114 long-horizon, real-world tasks across 8 work domains, where every task requires the agent to interleave GUI clicks with shell/code in one trajectory. Scored by a trajectory-aware Agent-as-Judgeโฆ See the full description on the dataset page: https://huggingface.co/datasets/wanlilll/WeaveBench.
This repository belongs to wanlilll on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
