wanlilll/WeaveBench
WeaveBench A long-horizon, real-world benchmark for computer-use agents with hybrid GUI + CLI + code interfaces. ๐ Accepted to EMNLP 2026 Main Conference โ see you in Budapest! ๐ Paper: arXiv:2606.09426 ๐ป Code: github.com/weavebench/WeaveBench ๐ Website: weavebench.github.io 114 long-horizon, real-world tasks across 8 work domains, where every task requires the agent to interleave GUI clicks with shell/code in one trajectory. Scored by a trajectory-aware Agent-as-Judgeโฆ See the full description on the dataset page: https://huggingface.co/datasets/wanlilll/WeaveBench.
Card: note EMNLP 2026 Main Conference acceptance
Fix DOC_task_2_pdf_form_fill warmup: resolve Master PDF Editor .deb URL dynamically (pinned 5.9.91 URL now 404s)
Add arXiv link: 2606.09426 (paper line + citation eprint)
README: rewrite to match v0.1.0 layout (vm + judge subdirs, --domains flag, openai/gpt-5.5 default, weavebench-demo, mirror, REPRODUCE pointer)
Update GitHub URL to weavebench/WeaveBench (org-owned)
Add judge_template.tar.gz (host-side OpenClaw profile + workspace for agent-as-judge)
Add v3_eyeson_apps Ubuntu qcow2 (paper-canonical VM image)
Add 114 tasks under 8-domain flat layout
Replace 4-batch internal layout with flat 8-domain layout
Add runtime asset: openclaw.tar.gz
Add runtime asset: codex.tar.gz
Add runtime asset: hermes.tar.gz
Add runtime asset: claudecode.tar.gz
Add runtime asset: hermes_mcp_wheels.tar.gz
Add 114 tasks across 8 domains + per-task workspace scaffolds
Upload README.md with huggingface_hub
initial commit
