datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gr00t-n15-robocasa-gr1-eval
GR00T N1.5 on RoboCasa GR-1 Tabletop — Evaluation Trajectories
Per-simulator-step recordings of 1,200 evaluation episodes (24 tasks × 50 episodes) of
NVIDIA's GR00T N1.5 vision-language-action model on the
RoboCasa GR-1 Tabletop Tasks benchmark.
Each episode stores every low-level transition — ego-view frames, robot state, executed actions,
rewards, full MuJoCo state, and 3D poses of every object and scene body — so that rollouts can be
re-analyzed or re-rendered without… See the full description on the dataset page: https://huggingface.co/datasets/junhee1998/gr00t-n15-robocasa-gr1-eval.MVU-Eval-Data
MVU-Eval Dataset
Paper | Code | Project Page
Dataset Description
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MVU-Eval-Team/MVU-Eval-Data.MVU-Eval-Data
MVU-Eval Dataset
Paper | Code | Project Page
Dataset Description
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/stephenbasd/MVU-Eval-Data.AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.zr0-robocasa-gr1-eval
ZR-0 on RoboCasa GR-1 Tabletop — Evaluation Trajectories
Per-simulator-step recordings of 2,400 evaluation episodes (24 tasks × 100 episodes) of the
ZR-0 vision-language-action model on the
RoboCasa GR-1 Tabletop Tasks benchmark.
Each episode stores every low-level transition — ego-view frames, robot state, executed actions,
rewards, full MuJoCo state, and 3D poses of every object and scene body — so that rollouts can be
re-analyzed or re-rendered without re-running the policy.… See the full description on the dataset page: https://huggingface.co/datasets/junhee1998/zr0-robocasa-gr1-eval.camerabench_vqa_lmms_evaleval2_180_2phase_resolvedcolor_clean256_v1vlm-eval-videos
VLM Eval Videos
A video benchmark dataset for evaluating Vision–Language Models (VLMs) on short-form action recognition.
Each clip is paired with a fixed question and a ground-truth short-sentence answer,
making it suitable for automated VLM inference pipelines and LLM-as-a-judge scoring.
Dataset Details
Description
VLM Eval Videos contains 693 short MP4 video clips drawn from YouTube, organised into five categories.
Four categories contain clips of… See the full description on the dataset page: https://huggingface.co/datasets/gnitoahc/vlm-eval-videos.crimson-ma2-evaluation
Crimson MolmoAct2: preliminary real-robot evaluation
Right-arm bottle pick-and-place rollouts recorded on 2026-09-24 (Asia/Bangkok).
This release contains three checkpoints, 31 logged attempts, videos from three cameras,
selected recorded telemetry, and joint trajectories. It is an evaluation evidence release,
not a training dataset or a completed six-checkpoint benchmark.
Model source: Kavin60606/crimson-ma2-ckpts.
The policy task text was pick and place the bottle. Each… See the full description on the dataset page: https://huggingface.co/datasets/CrimsonRobot/crimson-ma2-evaluation.egxo-household-egocentric-video-evaluation
EGXO Household Egocentric Video Dataset
EGXO maintains a continuously growing first-party household egocentric video catalogue. This repository documents the current commercial gold-standard 111-video, 10-hour evaluation release across 49 task families; it does not represent the size of the current catalogue.
Collection provenance
This release was collected through licensed GIG Rewards collection programs operated with telco partners. It is first-party inventory… See the full description on the dataset page: https://huggingface.co/datasets/egxodata/egxo-household-egocentric-video-evaluation.evaluation1evaluation5evaluation_videosevaluation2evaluation3evaluation4
