claw-eval/Claw-Eval
Claw-Eval End-to-end transparent benchmark for AI agents acting in the real world. Paper | Leaderboard | Code Dataset Structure Splits Split Examples Description general 161 Core agent tasks across 24 categories (communication, finance, ops, productivity, etc.) multimodal 101 Multimodal agentic tasks requiring perception and creation (webpage generation, video QA, document extraction, etc.) multi_turn 38 Multi-turn conversational tasks… See the full description on the dataset page: https://huggingface.co/datasets/claw-eval/Claw-Eval.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face