datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RoVid-X
Rethinking Video Generation Model for the Embodied World
If you like our project, please give us a star ⭐ on GitHub for the latest update.
Key features
4M robotic video clips(10K+ hours) for large-scale video generation training.
1300+ fine-grained robotic skills, covering diverse actions and task primitives.
Multi-modal physical annotations, including RGB, depth, and optical flow.
Multi-robot and multi-task diversity… See the full description on the dataset page: https://huggingface.co/datasets/DAGroup-PKU/RoVid-X.ROVER
The ROVER Visual SLAM Benchmark
Paper
News
[2024/12/05] Initial code release.
[2025/05/20] ROVER is accepted to IEEE Transactions on Robotics.
[2025/05/25] Dataset released on HuggingFace, see Utility section for HuggingFace download script.
Getting Started
The only required software is Docker. Each SLAM method comes with its own Docker container, making setup straightforward. We recommend using VSCode with the Docker extension for an enhanced… See the full description on the dataset page: https://huggingface.co/datasets/iis-esslingen/ROVER.jointavbench
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
Overview
JointAVBench is a benchmark for evaluating omni-modal large language models on joint audio-visual reasoning tasks. Each multiple-choice question is designed to require both visual and auditory information.
This repository contains the audited release of JointAVBench under the roverx12345 namespace. The benchmark keeps the original 2,853-question split while refining answer… See the full description on the dataset page: https://huggingface.co/datasets/roverx12345/jointavbench.rovi
Dataset Card for ROVI
Dataset Summary
ROVI is the first language-driven video inpainting dataset. Please check our project page for more details.
To prevent data contamination and reduce evaluation time. We only provide part of the testing data in test.json.
@article{wu2024lgvi,
title={Towards language-driven video inpainting via multimodal large language models},
author={Wu, Jianzong and Li, Xiangtai and Si, Chenyang and Zhou, Shangchen and Yang, Jingkang and Zhang… See the full description on the dataset page: https://huggingface.co/datasets/jianzongwu/rovi.gemma_rover_scoop_up_to_5
gemma_rover_scoop_up_to_5
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
office-rover-2task-raw
office-rover-2task-raw
LeRobot v2.1 dataset for a single SO-101 arm: "pick up the red / green cube
and put it in the box" with both cubes on the table. Two language-conditioned
tasks (instructions in Russian). Built for fine-tuning NVIDIA Isaac GR00T N1.7
on one 24 GB RTX 4090 — recipe, patches and 1800 evaluated attempts:
https://github.com/VShirokun/gr00t-on-4090
What is in it: the raw, un-engineered demonstrations: cubes in fixed orientation, 240×320 cameras. Trained as-is… See the full description on the dataset page: https://huggingface.co/datasets/VShirokun/office-rover-2task-raw.office-rover-miss-v21
office-rover-miss-v21
LeRobot v2.1 dataset for a single SO-101 arm: "pick up the red / green cube
and put it in the box" with both cubes on the table. Two language-conditioned
tasks (instructions in Russian). Built for fine-tuning NVIDIA Isaac GR00T N1.7
on one 24 GB RTX 4090 — recipe, patches and 1800 evaluated attempts:
https://github.com/VShirokun/gr00t-on-4090
What is in it: random cube yaw (0–90°) and 30 % deliberately failed grasps followed by scripted recovery — the data… See the full description on the dataset page: https://huggingface.co/datasets/VShirokun/office-rover-miss-v21.rovibook
RoVI-Book Dataset
🎉 **CVPR 2025** 🎉
Official dataset for Robotic Visual Instruction
This is an example to demonstrate the RoVI Book dataset, adapted from the Open-X Embodiments dataset. The bottom displays the proportion of each task type.
Paper:Robotic Visual Instruction
Project Page: https://robotic-visual-instruction.github.io/
Code: https://github.com/RoboticsVisualInstruction/RoVI-Book
Introduction
The RoVI-Book dataset is introduced alongside Robotic… See the full description on the dataset page: https://huggingface.co/datasets/yanbang/rovibook.so-arm101-pick-place
Reproduction and extension of GPU-DAD's SO-101 Pick-Cube Dataset with MuJoCo
This repository is our reproduction of the environment behind
gpudad/so101_pick_cube,
GPU-DAD's SO-101 pick-and-place data set. We rebuilt that scene in MuJoCo and
used the replica to generate a 2,000-episode pick-and-place dataset of our own:
rovolabs/so-arm101-pick-place.
The original Environment
This is a MuJoCo replica of the scene in
gpudad/so101_pick_cube
by gpudad. All fidelity… See the full description on the dataset page: https://huggingface.co/datasets/rovolabs/so-arm101-pick-place.nasa-mars-rover-images
NASA Mars Rover Image Catalog
Credit: NASA/JPL-Caltech/MSSS
Part of a dataset collection on Hugging Face.
Dataset description
The NASA Mars Rover Image Catalog contains metadata for every raw image captured by the Perseverance (Mars 2020) and Curiosity (MSL) rovers on the surface of Mars. Perseverance has been exploring Jezero Crater since February 2021, investigating an ancient river delta for signs of past microbial life and caching samples for future… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/nasa-mars-rover-images.ROVR-Open-Dataset
ROVR Open Dataset
Introduction
Welcome to the ROVR Open Dataset repository! This dataset is designed to empower autonomous driving and robotics research by providing rich, real-world data captured from ADAS cameras and LiDAR sensors. The dataset spans 50+ countries with over 20 million kilometers of driving data, making it ideal for training and developing advanced AI algorithms for depth estimation, object detection, and semantic segmentation.… See the full description on the dataset page: https://huggingface.co/datasets/ROVR-Network/ROVR-Open-Dataset.video_holmesoffice-rover-grasp-hd-v21
office-rover-grasp-hd-v21
LeRobot v2.1 dataset for a single SO-101 arm: "pick up the red / green cube
and put it in the box" with both cubes on the table. Two language-conditioned
tasks (instructions in Russian). Built for fine-tuning NVIDIA Isaac GR00T N1.7
on one 24 GB RTX 4090 — recipe, patches and 1800 evaluated attempts:
https://github.com/VShirokun/gr00t-on-4090
What is in it: grasp-centred demonstrations at 480×640 without cube rotation — the base skill. The other half of… See the full description on the dataset page: https://huggingface.co/datasets/VShirokun/office-rover-grasp-hd-v21.gemma_rover_drop_dirt_up_to_5
gemma_rover_drop_dirt_up_to_5
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
wsROVI
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
Overview
ROVI is a high-quality synthetic dataset featuring 1M curated web images with comprehensive image descriptions and bounding box annotations. Using a novel VLM-LLM re-captioning strategy, ROVI exceeds existing detection-centric datasets in image description, quality, and resolution, while containing two orders of magnitude more categories with an open-vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/CHang/ROVI.so101-rover-datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 47,
"total_frames": 38166,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:47"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aaron-ser/so101-rover-dataset.rovergemma_rover_scoop_up_to_4
gemma_rover_scoop_up_to_4
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
gemma_rover_scoop_drop_5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 10,
"total_frames": 6000,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vladfatu/gemma_rover_scoop_drop_5.gemma_rover_scoop_drop_6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 10,
"total_frames": 6000,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vladfatu/gemma_rover_scoop_drop_6.ROVERgemma_rover_dig_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 10,
"total_frames": 6000,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vladfatu/gemma_rover_dig_4.rovochat-g431b-evalsMS-OCT-CSC-ERM-MS-ROVsgemma_rover_scoop_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 10,
"total_frames": 6000,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vladfatu/gemma_rover_scoop_4.gemma_rover_scoop_drop_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 10,
"total_frames": 6000,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vladfatu/gemma_rover_scoop_drop_4.RoVid-20K-10s
RoVid-20K-10s
RoVid-20K-10s is a curated, fixed-length subset of
DAGroup-PKU/RoVid-X for
robotic image-to-video training. Each sample is a single 10.5417-second shot normalized to
1280x704, 24 FPS, and exactly 253 decoded frames.
The intended teacher-forcing task is:
video frame 0 + task instruction -> future video frames
Frame 0 is decoded directly from the MP4; it is not stored as a duplicate image file. The
model-facing caption field is exactly the original RoVid-X… See the full description on the dataset page: https://huggingface.co/datasets/Perflow-Shuai/RoVid-20K-10s.somos-clean-alpaca-es
Dataset Card for "somos-clean-alpaca-es"
More Information needed
gemma_rover_scoop_drop_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 10,
"total_frames": 6000,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vladfatu/gemma_rover_scoop_drop_1.
