datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claw-eval-live
Claw-Eval-Live
A live benchmark for workflow agents: 105 controlled tasks with fixtures,
mock services, sandboxed workspaces, task-specific graders, and recorded
execution evidence. The release is a time-stamped snapshot built from public
workflow-demand signals, and the signal-to-task pipeline is designed to be
rerun as demand and models evolve.
This dataset accompanies an anonymous submission to the NeurIPS 2026
Evaluations and Datasets Track.
Quick facts
105… See the full description on the dataset page: https://huggingface.co/datasets/claw-eval-live/claw-eval-live.eval_CLAW_cupThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 2,
"total_frames": 1500,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/smanni/eval_CLAW_cup.
