clawbench
ClawBench
ClawBench Dataset
ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites.
|💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website |
🚀 What's New
[2026.05.12] Added the V2 corpus (130 newer tasks across 63 platforms) and 7 new models judged with… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ClawBench.ClawBenchPro
ClawBenchPro
ClawBenchPro is a compact, builder-based workplace-agent benchmark package exported from Nanoclaw.
It contains task YAML files, prompts, task-local environment builders, skills, evaluation manifests,
provenance metadata, and checksums.
Included Splits
Dataset
Tasks
Groups
round_01_aligned_mix_800
800
base, hard_aligned, multi_turn_aligned, skills_aligned
persona_aligned_mix_200
200
base, hard, multi_turn, skills
Directory Layout… See the full description on the dataset page: https://huggingface.co/datasets/ErenJaegerYeager/ClawBenchPro.ClawBenchV2Trace
ClawBench V2 Traces
Full execution traces for every V2 model run scored on ClawBench.
|🏆 Leaderboard | 📊 Benchmark | 📖 Paper | 💻 Code | 🎬 V1 Traces |
Companion to TIGER-Lab/ClawBench (task definitions) and NAIL-Group/ClawBenchV1Trace (V1 traces). This dataset publishes the raw execution data for every V2 model run we've evaluated — one directory per (task × model × attempt), each with the screen recording, network capture, browser actions, agent reasoning, and the final… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ClawBenchV2Trace.ClawBench
ClawBench — A Benchmark for AI Web Agents
Can AI Agents Complete Everyday Online Tasks?
|💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website |
ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. The corpus ships in two slices: V1 — 153 tasks across 144 websites (the original… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBench.ClawBenchV1Trace
ClawBench V1 Traces
Full execution traces for every model run scored in ClawBench V1.
|🏆 Leaderboard | 📊 Benchmark | 🎞 V2 Traces | 📖 Paper | 💻 Code | 🌐 Website |
This is the companion dataset to NAIL-Group/ClawBench. Where the main dataset publishes the task definitions (instructions, rubrics, eval schemas), this one publishes the raw execution data — one directory per (task × model × attempt), each with the screen recording, network capture, browser actions, agent reasoning… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBenchV1Trace.qwen36-35b-a3b-clawbench
Qwen3.6-35B-A3B ClawBench Full Results
This repository contains organized ClawBench run artifacts from bench/ClawBench/test-output/qwen36-35b-a3b-full-20260510-2140.
The formal result directories are preserved under results/. Quarantined infrastructure-failure directories and top-level runtime control files such as .pid, .proc, .logpath, and .monitor-hermes-state/ are intentionally excluded from this organized export.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/whywhywhyyy/qwen36-35b-a3b-clawbench.
