datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ClawBench
ClawBench Dataset
ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites.
|💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website |
🚀 What's New
[2026.05.12] Added the V2 corpus (130 newer tasks across 63 platforms) and 7 new models judged with… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ClawBench.ClawBench
ClawBench — A Benchmark for AI Web Agents
Can AI Agents Complete Everyday Online Tasks?
|💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website |
ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. The corpus ships in two slices: V1 — 153 tasks across 144 websites (the original… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBench.ClawBench-test
ClawBench
Can AI Agents Complete Everyday Online Tasks?
ClawBench evaluates AI agents on 153 everyday tasks (such as booking flights, ordering groceries, submitting job applications) across 144 live websites. We capture 5 layers of behavioral data (session replay, screenshots, HTTP traffic, agent reasoning traces, and browser actions), collect human ground-truth for every task, and score with an agentic evaluator that provides step-level traceable diagnostics.
Paper… See the full description on the dataset page: https://huggingface.co/datasets/Duke313/ClawBench-test.
