CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TIGER-Lab /ClawBench ClawBench Dataset ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. |💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website | 🚀 What's New [2026.05.12] Added the V2 corpus (130 newer tasks across 63 platforms) and 7 new models judged with… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ClawBench.tabulartext-generationn<1K1 likes1k downloads3mo agoHugging Face02ErenJaegerYeager /ClawBenchPro ClawBenchPro ClawBenchPro is a compact, builder-based workplace-agent benchmark package exported from Nanoclaw. It contains task YAML files, prompts, task-local environment builders, skills, evaluation manifests, provenance metadata, and checksums. Included Splits Dataset Tasks Groups round_01_aligned_mix_800 800 base, hard_aligned, multi_turn_aligned, skills_aligned persona_aligned_mix_200 200 base, hard, multi_turn, skills Directory Layout… See the full description on the dataset page: https://huggingface.co/datasets/ErenJaegerYeager/ClawBenchPro.texttext-generation1K<n<10K0 likes822 downloads4mo agoHugging Face03TIGER-Lab /ClawBenchV2Trace ClawBench V2 Traces Full execution traces for every V2 model run scored on ClawBench. |🏆 Leaderboard | 📊 Benchmark | 📖 Paper | 💻 Code | 🎬 V1 Traces | Companion to TIGER-Lab/ClawBench (task definitions) and NAIL-Group/ClawBenchV1Trace (V1 traces). This dataset publishes the raw execution data for every V2 model run we've evaluated — one directory per (task × model × attempt), each with the screen recording, network capture, browser actions, agent reasoning, and the final… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ClawBenchV2Trace.1K<n<10K2 likes638 downloads4mo agoHugging Face04NAIL-Group /ClawBench ClawBench — A Benchmark for AI Web Agents Can AI Agents Complete Everyday Online Tasks? |💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website | ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. The corpus ships in two slices: V1 — 153 tasks across 144 websites (the original… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBench.tabulartext-generationn<1K2 likes364 downloads4mo agoHugging Face05NAIL-Group /ClawBenchV1Trace ClawBench V1 Traces Full execution traces for every model run scored in ClawBench V1. |🏆 Leaderboard | 📊 Benchmark | 🎞 V2 Traces | 📖 Paper | 💻 Code | 🌐 Website | This is the companion dataset to NAIL-Group/ClawBench. Where the main dataset publishes the task definitions (instructions, rubrics, eval schemas), this one publishes the raw execution data — one directory per (task × model × attempt), each with the screen recording, network capture, browser actions, agent reasoning… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBenchV1Trace.1K<n<10K1 likes311 downloads4mo agoHugging Face06whywhywhyyy /qwen36-35b-a3b-clawbench Qwen3.6-35B-A3B ClawBench Full Results This repository contains organized ClawBench run artifacts from bench/ClawBench/test-output/qwen36-35b-a3b-full-20260510-2140. The formal result directories are preserved under results/. Quarantined infrastructure-failure directories and top-level runtime control files such as .pid, .proc, .logpath, and .monitor-hermes-state/ are intentionally excluded from this organized export. Contents… See the full description on the dataset page: https://huggingface.co/datasets/whywhywhyyy/qwen36-35b-a3b-clawbench.imagen<1K1 likes196 downloads4mo agoHugging Face07ZhangArthurHao /ClawBenchV1Trace ClawBench V1 Traces Full execution traces for every model run scored in ClawBench V1. |🏆 Leaderboard | 📊 Benchmark | 📖 Paper | 💻 Code | 🌐 Website | This is the companion dataset to NAIL-Group/ClawBench. Where the main dataset publishes the task definitions (instructions, rubrics, eval schemas), this one publishes the raw execution data — one directory per (task × model × attempt), each with the screen recording, network capture, browser actions, agent reasoning, and the final… See the full description on the dataset page: https://huggingface.co/datasets/ZhangArthurHao/ClawBenchV1Trace.1K<n<10K0 likes191 downloads4mo agoHugging Face08NAIL-Group /ClawBenchV2Trace ClawBench V2 Traces Full execution traces for every V2 model run scored on ClawBench. |🏆 Leaderboard | 📊 Benchmark | 📖 Paper | 💻 Code | 🎬 V1 Traces | Companion to NAIL-Group/ClawBench (task definitions) and NAIL-Group/ClawBenchV1Trace (V1 traces). This dataset publishes the raw execution data for every V2 model run we've evaluated — one directory per (task × model × attempt), each with the screen recording, network capture, browser actions, agent reasoning, and the final… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBenchV2Trace.1K<n<10K0 likes128 downloads4mo agoHugging Face09NEAR-AI /clawbench ClawBench (nearai-bench packaging) A flat, self-contained repackaging of ClawBench — 319 agent tasks across 35 domains, difficulty levels L1–L4. Task content, environments and verifiers are unmodified, so scores stay comparable to the upstream ClawBench leaderboard. Why this exists Upstream ships a git repo of nested task directories (tasks/<domain>/<task>/{task.toml,instruction.md,environment/,verifier/,solution/}). Cloning that per worker is wasteful for an… See the full description on the dataset page: https://huggingface.co/datasets/NEAR-AI/clawbench.texttext-generationn<1K0 likes102 downloads2mo agoHugging Face10NoahMiller /ClawBench-resultsimage1K<n<10K0 likes78 downloads4mo agoHugging Face11Duke313 /ClawBench-test ClawBench Can AI Agents Complete Everyday Online Tasks? ClawBench evaluates AI agents on 153 everyday tasks (such as booking flights, ordering groceries, submitting job applications) across 144 live websites. We capture 5 layers of behavioral data (session replay, screenshots, HTTP traffic, agent reasoning traces, and browser actions), collect human ground-truth for every task, and score with an agentic evaluator that provides step-level traceable diagnostics. Paper… See the full description on the dataset page: https://huggingface.co/datasets/Duke313/ClawBench-test.tabulartext-generationn<1K0 likes53 downloads5mo agoHugging Face12ElPipila /clawbench-results ClawBench Results Persistent queue state and benchmark submissions for the ClawBench HF Space. 0 likes12 downloads1mo agoHugging Face13mzio /strl-main-ec-clawbench_flawedflowed_intercept-gc-claude_query-mc-claude_agent_sonnet-rbc-r0tabularn<1K0 likes10 downloads3mo agoHugging Face14mzio /strl-main-ec-clawbench_flawedflowed_intercept-gc-claude_query-mc-claude_agent_sonnet-rbc-r1tabularn<1K0 likes10 downloads3mo agoHugging Face15geminiDeveloper /clawbench0 likes8 downloads6mo agoHugging Face16mzio /strl-main-ec-clawbench_flawedflowed-gc-claude_query-mc-claude_agent_sonnet-rbc-memory_bm-r0tabularn<1K0 likes7 downloads3mo agoHugging Face17mzio /strl-main-ec-clawbench_flawedflowed_intercept-gc-claude_client_strl-mc-claude_agent_sonn-r0tabularn<1K0 likes7 downloads3mo agoHugging Face18mzio /strl-main-ec-clawbench_default-gc-claude_query-mc-claude_agent_sonnet-rbc-memory_bm25-s-r0tabularn<1K0 likes6 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.