datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ClawBench
ClawBench Dataset
ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites.
|💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website |
🚀 What's New
[2026.05.12] Added the V2 corpus (130 newer tasks across 63 platforms) and 7 new models judged with… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ClawBench.ClawBenchPro
ClawBenchPro
ClawBenchPro is a compact, builder-based workplace-agent benchmark package exported from Nanoclaw.
It contains task YAML files, prompts, task-local environment builders, skills, evaluation manifests,
provenance metadata, and checksums.
Included Splits
Dataset
Tasks
Groups
round_01_aligned_mix_800
800
base, hard_aligned, multi_turn_aligned, skills_aligned
persona_aligned_mix_200
200
base, hard, multi_turn, skills
Directory Layout… See the full description on the dataset page: https://huggingface.co/datasets/ErenJaegerYeager/ClawBenchPro.ClawBenchV2Trace
ClawBench V2 Traces
Full execution traces for every V2 model run scored on ClawBench.
|🏆 Leaderboard | 📊 Benchmark | 📖 Paper | 💻 Code | 🎬 V1 Traces |
Companion to TIGER-Lab/ClawBench (task definitions) and NAIL-Group/ClawBenchV1Trace (V1 traces). This dataset publishes the raw execution data for every V2 model run we've evaluated — one directory per (task × model × attempt), each with the screen recording, network capture, browser actions, agent reasoning, and the final… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ClawBenchV2Trace.ClawBench
ClawBench — A Benchmark for AI Web Agents
Can AI Agents Complete Everyday Online Tasks?
|💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website |
ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. The corpus ships in two slices: V1 — 153 tasks across 144 websites (the original… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBench.ClawBenchV1Trace
ClawBench V1 Traces
Full execution traces for every model run scored in ClawBench V1.
|🏆 Leaderboard | 📊 Benchmark | 🎞 V2 Traces | 📖 Paper | 💻 Code | 🌐 Website |
This is the companion dataset to NAIL-Group/ClawBench. Where the main dataset publishes the task definitions (instructions, rubrics, eval schemas), this one publishes the raw execution data — one directory per (task × model × attempt), each with the screen recording, network capture, browser actions, agent reasoning… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBenchV1Trace.qwen36-35b-a3b-clawbench
Qwen3.6-35B-A3B ClawBench Full Results
This repository contains organized ClawBench run artifacts from bench/ClawBench/test-output/qwen36-35b-a3b-full-20260510-2140.
The formal result directories are preserved under results/. Quarantined infrastructure-failure directories and top-level runtime control files such as .pid, .proc, .logpath, and .monitor-hermes-state/ are intentionally excluded from this organized export.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/whywhywhyyy/qwen36-35b-a3b-clawbench.ClawBenchV1Trace
ClawBench V1 Traces
Full execution traces for every model run scored in ClawBench V1.
|🏆 Leaderboard | 📊 Benchmark | 📖 Paper | 💻 Code | 🌐 Website |
This is the companion dataset to NAIL-Group/ClawBench. Where the main dataset publishes the task definitions (instructions, rubrics, eval schemas), this one publishes the raw execution data — one directory per (task × model × attempt), each with the screen recording, network capture, browser actions, agent reasoning, and the final… See the full description on the dataset page: https://huggingface.co/datasets/ZhangArthurHao/ClawBenchV1Trace.ClawBenchV2Trace
ClawBench V2 Traces
Full execution traces for every V2 model run scored on ClawBench.
|🏆 Leaderboard | 📊 Benchmark | 📖 Paper | 💻 Code | 🎬 V1 Traces |
Companion to NAIL-Group/ClawBench (task definitions) and NAIL-Group/ClawBenchV1Trace (V1 traces). This dataset publishes the raw execution data for every V2 model run we've evaluated — one directory per (task × model × attempt), each with the screen recording, network capture, browser actions, agent reasoning, and the final… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBenchV2Trace.clawbench
ClawBench (nearai-bench packaging)
A flat, self-contained repackaging of ClawBench
— 319 agent tasks across 35 domains, difficulty levels L1–L4. Task
content, environments and verifiers are unmodified, so scores stay
comparable to the upstream ClawBench leaderboard.
Why this exists
Upstream ships a git repo of nested task directories
(tasks/<domain>/<task>/{task.toml,instruction.md,environment/,verifier/,solution/}).
Cloning that per worker is wasteful for an… See the full description on the dataset page: https://huggingface.co/datasets/NEAR-AI/clawbench.ClawBench-resultsClawBench-test
ClawBench
Can AI Agents Complete Everyday Online Tasks?
ClawBench evaluates AI agents on 153 everyday tasks (such as booking flights, ordering groceries, submitting job applications) across 144 live websites. We capture 5 layers of behavioral data (session replay, screenshots, HTTP traffic, agent reasoning traces, and browser actions), collect human ground-truth for every task, and score with an agentic evaluator that provides step-level traceable diagnostics.
Paper… See the full description on the dataset page: https://huggingface.co/datasets/Duke313/ClawBench-test.clawbench-results
ClawBench Results
Persistent queue state and benchmark submissions for the ClawBench HF Space.
strl-main-ec-clawbench_flawedflowed_intercept-gc-claude_query-mc-claude_agent_sonnet-rbc-r0strl-main-ec-clawbench_flawedflowed_intercept-gc-claude_query-mc-claude_agent_sonnet-rbc-r1clawbenchstrl-main-ec-clawbench_flawedflowed-gc-claude_query-mc-claude_agent_sonnet-rbc-memory_bm-r0strl-main-ec-clawbench_flawedflowed_intercept-gc-claude_client_strl-mc-claude_agent_sonn-r0strl-main-ec-clawbench_default-gc-claude_query-mc-claude_agent_sonnet-rbc-memory_bm25-s-r0
