CoolFace
8 results

clawbench

TIGER-Lab /ClawBench ClawBench Dataset ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. |💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website | 🚀 What's New [2026.05.12] Added the V2 corpus (130 newer tasks across 63 platforms) and 7 new models judged with… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ClawBench.tabulartext-generationn<1K1 likes1k downloads3mo agoHugging FaceErenJaegerYeager /ClawBenchPro ClawBenchPro ClawBenchPro is a compact, builder-based workplace-agent benchmark package exported from Nanoclaw. It contains task YAML files, prompts, task-local environment builders, skills, evaluation manifests, provenance metadata, and checksums. Included Splits Dataset Tasks Groups round_01_aligned_mix_800 800 base, hard_aligned, multi_turn_aligned, skills_aligned persona_aligned_mix_200 200 base, hard, multi_turn, skills Directory Layout… See the full description on the dataset page: https://huggingface.co/datasets/ErenJaegerYeager/ClawBenchPro.texttext-generation1K<n<10K0 likes822 downloads4mo agoHugging FaceTIGER-Lab /ClawBenchV2Trace ClawBench V2 Traces Full execution traces for every V2 model run scored on ClawBench. |🏆 Leaderboard | 📊 Benchmark | 📖 Paper | 💻 Code | 🎬 V1 Traces | Companion to TIGER-Lab/ClawBench (task definitions) and NAIL-Group/ClawBenchV1Trace (V1 traces). This dataset publishes the raw execution data for every V2 model run we've evaluated — one directory per (task × model × attempt), each with the screen recording, network capture, browser actions, agent reasoning, and the final… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ClawBenchV2Trace.1K<n<10K2 likes638 downloads4mo agoHugging FaceNAIL-Group /ClawBench ClawBench — A Benchmark for AI Web Agents Can AI Agents Complete Everyday Online Tasks? |💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website | ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. The corpus ships in two slices: V1 — 153 tasks across 144 websites (the original… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBench.tabulartext-generationn<1K2 likes364 downloads4mo agoHugging FaceNAIL-Group /ClawBenchV1Trace ClawBench V1 Traces Full execution traces for every model run scored in ClawBench V1. |🏆 Leaderboard | 📊 Benchmark | 🎞 V2 Traces | 📖 Paper | 💻 Code | 🌐 Website | This is the companion dataset to NAIL-Group/ClawBench. Where the main dataset publishes the task definitions (instructions, rubrics, eval schemas), this one publishes the raw execution data — one directory per (task × model × attempt), each with the screen recording, network capture, browser actions, agent reasoning… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBenchV1Trace.1K<n<10K1 likes311 downloads4mo agoHugging Facewhywhywhyyy /qwen36-35b-a3b-clawbench Qwen3.6-35B-A3B ClawBench Full Results This repository contains organized ClawBench run artifacts from bench/ClawBench/test-output/qwen36-35b-a3b-full-20260510-2140. The formal result directories are preserved under results/. Quarantined infrastructure-failure directories and top-level runtime control files such as .pid, .proc, .logpath, and .monitor-hermes-state/ are intentionally excluded from this organized export. Contents… See the full description on the dataset page: https://huggingface.co/datasets/whywhywhyyy/qwen36-35b-a3b-clawbench.imagen<1K1 likes196 downloads4mo agoHugging Face