datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InternData-fractal20220817_dataRoboInter-Data
RoboInter-Data: Intermediate Representation Annotations for Robot Manipulation
Rich, dense, per-frame intermediate representation annotations for robot manipulation, built on top of DROID and RH20T. Developed as part of the RoboInter project. You can try our Online demo.
The annotations cover 230k episodes and include: subtasks,
primitive skills, segmentation, gripper/object bounding boxes, placement proposals, affordance boxes,
grasp poses, traces, contact points, etc. And each… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/RoboInter-Data.minWM-dataProLongVid_data
Dataset Card for ProLongVid-data
Uses
This dataset is used for the training of the ProLongVid model. We only allow the use of this dataset for academic research and education purpose.
Paper: For more details, please check our paper
Code: For training recipe and other update, please refer to github repo.
Citation
@inproceedings{
wang2025prolongvid,
title={ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning},
author={Rui Wang… See the full description on the dataset page: https://huggingface.co/datasets/prolongvid/ProLongVid_data.VBVR-Bench-Data
VBVR: A Very Big Video Reasoning Suite
Overview
Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture,
enabling intuitive reasoning over motion, interaction, and causality. Rapid progress in video models has focused primarily on visual quality.
Systematically studying video reasoning and its scaling behavior suffers from a lack of… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Bench-Data.Cosmos-Reason1-SFT-Dataset
Dataset Description:
The data format is a pair of video and text annotations. We summarize the data and annotations in Table 4 (SFT), Table 5 (RL), and Table 6 (Benchmark) of the Cosmos-Reason1 paper. We release the annotations for embodied reasoning tasks for BridgeDatav2, RoboVQA, Agibot, HoloAssist, AV, and the videos for the RoboVQA and AV datasets. We additionally release the annotations and videos for the RoboFail dataset for benchmarks. By releasing the dataset, NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Cosmos-Reason1-SFT-Dataset.OraRL-Data
OraRL-Data
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Video-ORA-9B] [💻 Code]
We release OraRL-Data, the official evaluation suite for Video-ORA and OraRL.
It packages the canonical annotations and referenced raw media used by the OraRL evaluation suite: 109,374 examples across 16 benchmark configs and 29 splits, with 518.9 GiB of manifested files. The complete evaluation release lives under OraRL-eval-data/, leaving room for the separate OraRL training release in this repository.… See the full description on the dataset page: https://huggingface.co/datasets/OraRL/OraRL-Data.DataSpace
DataSpace
Paper ·
Code ·
Leaderboard ·
KDD Cup 2026
DataSpace is a benchmark for data agents that perform verifiable analytics over
heterogeneous, task-local workspaces. Each task provides a natural-language
question and a workspace containing a combination of CSV, JSON, SQLite,
Markdown, PDF, and video artifacts. The required output is a complete tabular
result.
DataSpace is also the official benchmark of the
KDD Cup 2026 Data Agent Track.
Release policy
This… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTDial/DataSpace.MuseVLA-dataset
MuseVLA Dataset
Multi-modal robot manipulation dataset with synchronized RGB, depth, acoustic,
thermal, and radar streams. Released as two parts (dataset_01/,
dataset_02/) sharing the same per-episode layout. Together they cover
~1400 episodes across 11 instructions (towel / clothes / box / item / drink
manipulation).
Per-episode contents
{episode_name}/
├── video.mp4 # RGB, 1280×720, 30 fps
├── mask/video.mp4 #… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MuseVLA-dataset.minuszero-indian-autonomous-driving-dataset
Minus Zero Indian Urban Autonomous Driving Dataset
Overview
This dataset provides original multicamera autonomous-driving recordings in MCAP format. It is designed for non-commercial research on surround-view perception, temporal and cross-camera synchronization, H.265 video pipelines, localization, GNSS/pose integration, and robotics data tooling.
Recordings include camera and GNSS/pose streams, with machine-state telemetry present in a small subset. Camera… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset.Seamless_Dummy_Dataset_Fixed_3
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
qualcomm-interactive-cooking-dataset-ego-mistake-corrections
Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark
Description
This dataset contains cooking videos with timestamped instruction and feedback for task guidance.
Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp.
Dataset Details
Release files:
annotations/annotations.json
videos/*.MP4
Release statistics:
Total videos: 40
Total released annotations: 1,597
Text type counts in… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.LTX2.3-22B_IC-LoRA-CrossView-Prompt-Dataset
CrossView Prompt Dataset
The training dataset behind the
CrossView Prompt IC-LoRA
for LTX-Video 2.3 — a "virtual second camera" adapter that re-renders a scene
from a new viewpoint described by a short prompt.
Each sample is a pair of static-camera clips of the same scene (a reference
view and a target view) plus a camera-delta caption describing how the
target camera differs from the reference.
Contents
clips/<scene>/<cam>.mp4 # 504 unique clips, native… See the full description on the dataset page: https://huggingface.co/datasets/Cseti/LTX2.3-22B_IC-LoRA-CrossView-Prompt-Dataset.ViTeX-Dataset
ViTeX-Dataset
🌐 Project page ·
📊 Dataset ·
🧪 Benchmark code ·
🤖 Model & Inference code ·
🏆 Leaderboard
Paired real-video dataset for video scene text editing: given a source video, a binary text-region mask, and a (source string → target string) pair, replace only the masked scene text across all frames while preserving the rest of the scene.
Anonymous release under double-blind review at NeurIPS 2026 Datasets and Benchmarks Track. Author list and DOI updated after… See the full description on the dataset page: https://huggingface.co/datasets/ViTeX-Bench/ViTeX-Dataset.Inst-It-Dataset
Inst-IT Dataset: An Instruction Tuning Dataset with Multi-level Fine-Grained Annotations
introduced in the paper Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning
🌐 Homepage | Code | 🤗 Paper | 📖 arXiv
Inst-IT Dataset Overview
We create a large-scale instruction tuning dataset, the Inst-it Dataset. To the best of our knowledge, this is the first dataset that provides fine-grained annotations centric on specific… See the full description on the dataset page: https://huggingface.co/datasets/Inst-IT/Inst-It-Dataset.UE_dataenhanced-fall-dataset
Enhanced Fall Dataset (total = 25,691)
Prerequisite
Download bear7011/gemma-4-e4b-kinetics_54K first for some overlapped videos. (This setting prevents cascading forgetting.)
Notification!
Do not mix the sora-accident dataset into the training process.
Video sources:
videos/kinetics_fall, videos/kinetics_neg — Kinetics dataset
videos/oops — OOPS! dataset (Columbia)
File Structure
├── annotations
│ ├── prompts.json… See the full description on the dataset page: https://huggingface.co/datasets/bear7011/enhanced-fall-dataset.Test-Dataset
AXIS held-out 20 — a cross-embodiment few-shot adaptation benchmark
20 tasks never seen in pretraining, 20 demonstrations each, LoRA adaptation, rollout evaluation.
This release is the SPECIFICATION and the INDEX, not the demonstration data. It is published
first and on purpose: everything here is what you need to render the benchmark on your own
embodiment, and none of it depends on our video encoding being finished.
Status: the task list is CANDIDATES. The learnability gate… See the full description on the dataset page: https://huggingface.co/datasets/axisrobotics/Test-Dataset.LongLive2.0-Toy-Dataset
LongLive2.0 Toy Dataset
This dataset is a toy format-checking dataset for the LongLive2.0 release
code. It is intended to help users verify AR diffusion training, DMD
distillation, and prompt formatting before preparing a larger dataset.
Dataset placeholder:
https://huggingface.co/datasets/Efficient-Large-Model/LongLive2-Toy-Dataset
Expected Layout
The released toy dataset will contain two separate training folders:
ar_training/: paired video/caption data for AR… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/LongLive2.0-Toy-Dataset.ebm-role-play-datasetCosmos-Reason1-RL-Dataset
Dataset Description:
The data format is a pair of video and text annotations. We summarize the data and annotations in Table 4 (SFT), Table 5 (RL), and Table 6 (Benchmark) of the Cosmos-Reason1 paper. We release the annotations for embodied reasoning tasks for BridgeDatav2, RoboVQA, Agibot, HoloAssist, AV, and the videos for the RoboVQA and AV datasets. We additionally release the annotations and videos for the RoboFail dataset for benchmarks. By releasing the dataset, NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Cosmos-Reason1-RL-Dataset.HDI-Dataset
High-Density Human Interaction (HDI) Dataset
Official ECCV 2026 release for the paper "Temporal and Cross-modal Alignment for Enhanced Audiovisual Video Captioning."
The High-Density Human Interaction (HDI) Dataset contains 3,527 short audiovisual clips paired with dense captions. The dataset is curated for challenging audio-visual captioning cases, including multi-speaker dialogue, physical actions coupled with sound, and cross-modal causal chains.
Files… See the full description on the dataset page: https://huggingface.co/datasets/26TCA/HDI-Dataset.PhaForce-Dataset
PhaForce Dataset
This repository hosts the real-robot contact-rich manipulation dataset used by PhaForce: Phase-Scheduled Visual-Force Policy Learning with Slow Planning and Fast Correction for Contact-Rich Manipulation.
Project page: https://thu-wangmx.github.io/phaforce/
Contents
The dataset is converted to FTP-1 / Open-X-Tactile-style zarr keys. Each task is stored as one zarr group:
open_drawer.zarr: 50 trajectories, 21,653 frames
plug_in_charger.zarr: 100… See the full description on the dataset page: https://huggingface.co/datasets/wangmingxinthu/PhaForce-Dataset.Fence-Climbing-Action-Recognition-Dataset
Fence Climbing Action Recognition Dataset
The current security industry faces challenges from people climbing over walls, fences, and other security hazards. Traditional surveillance methods often cannot timely and effectively recognize these abnormal behaviors. Existing solutions are insufficient in the accuracy and real-time detection of actions, resulting in the inability to quickly respond to potential dangers. This dataset aims to support the training of action recognition… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Fence-Climbing-Action-Recognition-Dataset.MCF-Dataset
MCF: Text LLMS For Multimodal Emotional Causality
Data
Dataset task definition and annotation example of the MCF framework. The framework contains two core subtasks:
five-tuple
element extraction (identifying Target, Holder, Aspect, Opinion, Sentiment, and Rationale) and sentiment chain
analysis (constructing causal
relationship chains between emotional events).
The dataset is provided with the following structure. Each sample includes video, audio, and dialogue… See the full description on the dataset page: https://huggingface.co/datasets/ZHANGYUXUAN-zR/MCF-Dataset.Spatial-Scene-Synthetic-DatasetDataSet_mix_duck_oct_cablsfb_datasetChenLong_Embodied_Intelligence_Dataset
ChenLong Embodied Intelligence Dataset
本仓库用于统一管理辰龙机器人实习中的数据集、模型权重、训练结果和说明文档。后续新增不同任务、采集批次、模型版本或实验资源时,都放在这里统一维护。
当前目录
embodied_dataset/:具身智能采集数据集,采用 LeRobot v3.0 结构,包含 data/、meta/、videos/。
yolo_dataset/:YOLO 目标检测数据、模型权重、训练参数和评估结果,当前包含 blue_bucket_yolov8/。
待新增新的数据集或模型。
新增数据集要求
具身数据优先采用 LeRobot v3.0 格式:meta/info.json、meta/stats.json、tasks、episodes、逐帧 Parquet 数据和按相机划分的视频。新增数据集至少写清:
任务:任务文本、目标物、成功标准、失败标准。
硬件:机器人型号、自由度、夹爪、相机位置、分辨率、FPS。… See the full description on the dataset page: https://huggingface.co/datasets/vvzc/ChenLong_Embodied_Intelligence_Dataset.VBVR-Bench-Data
VBVR: A Very Big Video Reasoning Suite
Overview
Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture,
enabling intuitive reasoning over motion, interaction, and causality. Rapid progress in video models has focused primarily on visual quality.
Systematically studying video reasoning and its scaling behavior suffers from a lack of… See the full description on the dataset page: https://huggingface.co/datasets/abs794/VBVR-Bench-Data.
