CoolFace
Datasetpublic

claw-eval/Claw-Eval

Claw-Eval End-to-end transparent benchmark for AI agents acting in the real world. Paper | Leaderboard | Code Dataset Structure Splits Split Examples Description general 161 Core agent tasks across 24 categories (communication, finance, ops, productivity, etc.) multimodal 101 Multimodal agentic tasks requiring perception and creation (webpage generation, video QA, document extraction, etc.) multi_turn 38 Multi-turn conversational tasks… See the full description on the dataset page: https://huggingface.co/datasets/claw-eval/Claw-Eval.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
32likes3.4kdownloads
settings

This repository belongs to claw-eval on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameClaw-Eval
visibilitypublic
licencemit
gatedno
ownerclaw-eval
Account settings
claw-eval/Claw-Eval · CoolFace