datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLaVA-OneVision-2-Data
LLaVA-OneVision-2-Data
Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training.
At a Glance
The dataset is split across two Hugging Face repositories because of its size:
Repository
What it contains
Part 1 (this repository)
~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.HiFi-UMI-2K
HiFi-UMI-2K: High-Fidelity Robot-Free Manipulation Data
2,000 hours released · 6 synchronized camera views · 480+ scenes · 3 mm pose accuracy · <40 µs synchronization
🌐 Project Website |
📦 Dataset |
📄 Paper: arXiv:2607.25895
Examples from the HiFi-UMI corpus. Click the image to play the video.
📚 Introduction
HiFi-UMI is a portable, high-fidelity bimanual capture system for collecting robot-free manipulation demonstrations.… See the full description on the dataset page: https://huggingface.co/datasets/simple-world-lab/HiFi-UMI-2K.EgoLifeData cleaning, stay tuned! Please refer to https://egolife-ai.github.io/ first for general info.
Checkout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information.
Code: https://github.com/egolife-ai/EgoLife
LLaVA-Video-178K
Dataset Card for LLaVA-Video-178K
Uses
This dataset is used for the training of the LLaVA-Video model. We only allow the use of this dataset for academic research and education purpose. For OpenAI GPT-4 generated data, we recommend the users to check the OpenAI Usage Policy.
Data Sources
For the training of LLaVA-Video, we utilized video-language data from five primary sources:
LLaVA-Video-178K: This dataset includes 178,510 caption entries, 960,792 open-ended… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K.humanoid-everyday
Humanoid Everyday
A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
Overview
Humanoid Everyday is a large-scale, diverse humanoid manipulation dataset designed for open-world robotic learning and embodied intelligence.
It contains over 260 tasks across 7 major categories, covering dexterous manipulation, human–humanoid interaction, and locomotion-integrated activities.All data were collected through a human-supervised teleoperation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/USC-PSI-Lab/humanoid-everyday.robocasa365-pretrain-mg
Pretraining (MimicGen) — atomic
MimicGen-generated rollouts across 60 atomic tasks (~10,000 demos/task). 1,615 hours total, generated by scripted augmentation from human demonstrations.
Part of the RoboCasa365 collection. Flat LeRobot v3.0 mirror of RoboCasa365 — standard layout, drop-in loadable.
Stats
Episodes: 536,030
Frames: 116,246,439 (20 fps → 1615 h)
Tasks: 720 (natural-language phrasings; underlying RoboCasa task classes: 60)
Cameras: 3 × 256×256 h264… See the full description on the dataset page: https://huggingface.co/datasets/ember-lab-berkeley/robocasa365-pretrain-mg.minWM-dataGUI-World
GUI-World: A Dataset for GUI-Orientated Multimodal Large Language Models
Dataset: GUI-World
Overview
GUI-World introduces a comprehensive benchmark for evaluating MLLMs in dynamic and complex GUI environments. It features extensive annotations covering six GUI scenarios and eight types of GUI-oriented questions. The dataset assesses state-of-the-art ImageLLMs and VideoLLMs, highlighting their limitations in handling dynamic and multi-step tasks. It provides… See the full description on the dataset page: https://huggingface.co/datasets/ONE-Lab/GUI-World.Inf-Stream-Train
Inf-Stream-Eval
Inf-Stream-Eval is a benchmark for evaluating vision-language models (VLMs) on near-infinite video streams. It consists of videos averaging over two hours in length that require dense, per-second alignment between video frames and text.
This dataset was introduced in the paper StreamingVLM: Real-Time Understanding for Infinite Video Streams.
Project Page: https://streamingvlm.hanlab.ai
GitHub Repository: https://github.com/mit-han-lab/streaming-vlm… See the full description on the dataset page: https://huggingface.co/datasets/mit-han-lab/Inf-Stream-Train.robocasa365-pretrain-composite
Pretraining (Human) — composite
Human teleoperation data for 235 composite (multi-step) tasks (~100 demos/task). The long-horizon half of RoboCasa365's 482-hour pretraining-human dataset.
Part of the RoboCasa365 collection. Flat LeRobot v3.0 mirror of RoboCasa365 — standard layout, drop-in loadable.
Stats
Episodes: 24,687
Frames: 27,610,913 (20 fps → 383 h)
Tasks: 4,428 (natural-language phrasings; underlying RoboCasa task classes: 235)
Cameras: 3 × 256×256 h264… See the full description on the dataset page: https://huggingface.co/datasets/ember-lab-berkeley/robocasa365-pretrain-composite.m3evalM³Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks
Jie Huang1,*
Ruixun Liu1,*
Sirui Sun1
Xinyi Yang1
Yin Li2
Yixin Zhu1
Yiwu Zhong1,†
1Peking University
2University of Wisconsin-Madison
* Equal contribution. † Corresponding author.
News
2026-6-4: We released the M³Eval benchmark, code, and project page.
M³Eval Overview
Abstract
As multi-modal models advance… See the full description on the dataset page: https://huggingface.co/datasets/PKU-VaLuE-Lab/m3eval.SuperMemory-VQA
SuperMemoryVQA
SuperMemory-VQA is an egocentric visual question answering benchmark for
evaluating long-horizon memory in augmented reality assistant settings. The
dataset is designed around practical questions a person might ask a wearable
memory assistant, such as where an object was left, what someone said earlier,
whether a planned step was completed, or what happened next in a longer event.
The benchmark contains 4,853 human-verified question-answer pairs grounded in
52.9… See the full description on the dataset page: https://huggingface.co/datasets/OSU-AIoT-MLSys-Lab/SuperMemory-VQA.VideoDetailCaptionPhysics-IQ-Verified
Physics-IQ Verified Dataset
This repository hosts the Physics-IQ Verified benchmark data for evaluating physical understanding in generative video models.
Physics-IQ Verified is derived from the original Physics-IQ benchmark dataset.
Original Physics-IQ
Paper: Do generative video models understand physical principles?
Repository: Code | Dataset in Google Cloud
Physics-IQ Verified (Recommended)
Paper: Physics-IQ Verified
Repository: Code | Dataset: Here in this repo :)
We… See the full description on the dataset page: https://huggingface.co/datasets/Anates-Labs-Research/Physics-IQ-Verified.Wanda
WANDA: Worlds in One Demo
A Synthetic Data Engine for Learning Open-World Mobile Manipulation
🌐 Project page: https://wanda.lecar-lab.org/ · 📄 Paper (PDF) · 💻 Code (coming soon) · 🕹️ Interactive 4D viewer
Lingxiao Guo*, Huanyu Li*, Guanya Shi — Carnegie Mellon University
*Equal contribution; order decided by a coin flip.
WANDA is a synthetic data engine that turns one human demonstration into diverse training data for
open-world mobile manipulation. From a… See the full description on the dataset page: https://huggingface.co/datasets/LeCAR-Lab/Wanda.robocasa365-target-composite-seen
Target (Human) — composite-seen
Human teleoperation data for 16 composite tasks (500 demos/task) recorded in 10 held-out target kitchens. All tasks are also represented in the pretraining datasets ('composite-seen').
Part of the RoboCasa365 collection. Flat LeRobot v3.0 mirror of RoboCasa365 — standard layout, drop-in loadable.
Stats
Episodes: 8,077
Frames: 6,002,265 (20 fps → 83 h)
Tasks: 382 (natural-language phrasings; underlying RoboCasa task classes: 16)… See the full description on the dataset page: https://huggingface.co/datasets/ember-lab-berkeley/robocasa365-target-composite-seen.Humanoid-Everyday-G1compas3d
CoMPAS3D: A Dataset and Benchmark for Interactive Motion
CoMPAS3D (Complex Multi-Level Person-Interaction Annotated Salsa Dataset) is a large-scale motion capture dataset designed to support research on nonverbal, physical communication through dance. It contains over 3 hours of improvised salsa duet performances by 18 dancers across beginner, intermediate, and professional skill levels. Each sequence features high-fidelity 3D motion data in the form of SMPL-X (.npz) files… See the full description on the dataset page: https://huggingface.co/datasets/Rosie-Lab/compas3d.egoinfinity
EgoInfinity (preview)
Derivative scene assets for a curated subset of Action100M (Meta FAIR) clips.
Preview — not for general release. Schema and contents may change.
License
FAIR Noncommercial Research License v1 (see LICENSE-Action100M). Noncommercial research only.
Built by the Rice RobotPI Lab.
Retarget
For 104 of the clips under samples/, we provide retargeting results on four
robot embodiments under samples/<clip>/retarget/<robot>/:
franka —… See the full description on the dataset page: https://huggingface.co/datasets/Rice-RobotPI-Lab/egoinfinity.OSSL-v2
Open Screen Soundtrack Libary Version 2 (OSSL-v2)
Paired film video ↔ soundtrack music clips for video-to-music generation.
Paper: Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections
Layout
ossl-v2-hf/
├── metadata.csv # one row per movie (film + source metadata)
├── splits/{train,test}.txt # clip_ids per split
├── public_train_test_remapped.pkl # {"train":[clip_id...], "test":[clip_id...]}
├── train/{video… See the full description on the dataset page: https://huggingface.co/datasets/McAuley-Lab/OSSL-v2.RoboTwin-LeRobotphysical-ai-bench-conditional-generation
Physical AI Bench - Conditional Generation
Paper | Code
This dataset (Phsical AI benchmark, PAI-Bench) consisting of 600 examples across three key scenarios: robotic arm operations, driving, and ego-centric everyday life scenes, each representing a critical aspect of Physical AI. This dataset is constructed by sampling a number of videos from three different datasets. The specific details are provided below.
Dataset
Category
Sample Nums
Agibot World
Robotics
200
OpenDV… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-conditional-generation.robocasa-composite-raw-videosWorldSenseopenarm-packingbench-v2-rawrobocasa365-pretrain-atomic
Pretraining (Human) — atomic
Human teleoperation data for 65 atomic single-skill tasks (~100 demos/task). Part of RoboCasa365's 482-hour pretraining-human dataset.
Part of the RoboCasa365 collection. Flat LeRobot v3.0 mirror of RoboCasa365 — standard layout, drop-in loadable.
Stats
Episodes: 7,356
Frames: 1,495,313 (20 fps → 21 h)
Tasks: 617 (natural-language phrasings; underlying RoboCasa task classes: 65)
Cameras: 3 × 256×256 h264 video (robot0_agentview_left /… See the full description on the dataset page: https://huggingface.co/datasets/ember-lab-berkeley/robocasa365-pretrain-atomic.robocasa365-target-composite-unseen
Target (Human) — composite-unseen
Human teleoperation data for 16 composite tasks (500 demos/task) recorded in 10 held-out target kitchens. These tasks are NOT represented in the pretraining datasets ('composite-unseen') — held-out generalization benchmark.
Part of the RoboCasa365 collection. Flat LeRobot v3.0 mirror of RoboCasa365 — standard layout, drop-in loadable.
Stats
Episodes: 8,104
Frames: 6,724,287 (20 fps → 93 h)
Tasks: 608 (natural-language phrasings;… See the full description on the dataset page: https://huggingface.co/datasets/ember-lab-berkeley/robocasa365-target-composite-unseen.ha-lab-dataset
Hugging Face Dataset Upload Notes
This note summarizes common commands for uploading and managing selected
folders in the Hugging Face dataset repository.
Repository:
sehoonha/ha-lab-dataset
Login
Use a Hugging Face token with write permission.
hf auth login
hf auth whoami
Upload One Folder
Upload only the local test_folder/ directory to test_folder/ in the dataset
repository.
hf upload sehoonha/ha-lab-dataset test_folder test_folder --repo-type… See the full description on the dataset page: https://huggingface.co/datasets/sehoonha/ha-lab-dataset.datagen-stack-v1-joint-5cam
datagen-stack-v1-joint-5cam
Auto-generated SFT dataset for the stack_retrieve family — cuRobo-planned, physics- &
LTL-safety-checked demos, converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench — all 28 stack_retrieve base tasks.
Per task: 40 success + LTL-safe trajectories → 1120 episodes.
Contents
Episodes
1120 (28 base tasks × 40)
Frames
2,652,083
Unique language tasks
8… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-stack-v1-joint-5cam.Humanoid-Everyday-H1
