datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EmbodiedGenDatahttps://huggingface.co/spaces/HorizonRobotics/EmbodiedGen-Gallery-Explorer
Hy-Embodied-0.5-VLA-Data
Hy-Embodied-0.5-VLA
From Vision-Language-Action Models to a Real-World Robot Learning Stack
Tencent Robotics X × Tencent Hy Team
📖 Abstract
We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data.mllm-as-embodied-world-judge
MLLM-as-Embodied-World-Judge
Data for judging physical adherence and instruction alignment of generated
embodied-manipulation videos.
Start here
path
what it is
final/
the current release — train.jsonl (11,520), test.jsonl (802), and its README
data/
source and generated videos, referenced by video_url in the splits
Benchmark tooling
path
what it is
bench/LEADERBOARD.md
judge results table
bench/TESTSET.md
benchmark… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/mllm-as-embodied-world-judge.Embodied-Captioning
Embodied Image Captioning – Manually Annotated Test Set
Paper: Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions (ICCV 2025)Authors: Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo NataleAffiliations: Italian Institute of Technology (IIT), University of GenoaProject Website: https://hsp-iit.github.io/embodied-captioningCode: https://github.com/hsp-iit/embodied-captioning
📦… See the full description on the dataset page: https://huggingface.co/datasets/TommyBsk/Embodied-Captioning.PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes
PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card
Dataset Description
PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.Xlang-Embodied-datasetEmbodied-R1.5-SFT-Dataset
Embodied-R1.5-SFT-Dataset
🌐 Project Page |
📄 arXiv |
💻 Code |
🧰 EmbodiedEvalKit |
🤗 Models & Datasets
🗓️ Update — 2026-08-20 (20260820). All 34 Stage 1 SFT JSON annotation files have been uploaded to sft_datasets_json/. The complete JSON ↔ image/video data mapping is documented in the Dataset composition table below.
⚠️ Partial release. This repository currently contains only a subset of the full Stage 1 SFT… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1.5-SFT-Dataset.EmbodiedSplat
EmbodiedSplat 🛋️
Online Feed-Forward Semantic 3DGS
for Open-Vocabulary 3D Scene Understanding
Seungjun Lee ·
Zihan Wang ·
Yunsong Wang ·
Gim Hee Lee
National University of Singapore
CVPR 2026
Code | Paper | Project Page
Build and understand at Once! By taking over 300 streaming images, our EmbodiedSplat reconstructs whole-scene open-vocabulary 3DGS in online manner at up to 5-6 FPS per-frame processing time. Reconstructed scene… See the full description on the dataset page: https://huggingface.co/datasets/onandon/EmbodiedSplat.embodied_features_and_demos_liberoDataset for Embodied Chain-of-Thought Reasoning for LIBERO-90, as used by ECoT-Lite.
TFDS Demonstration Data
The TFDS dataset contains successful demonstration trajectories for LIBERO-90 (50 trajectories for each of 90 tasks). It was created by rolling out the actions provided in the original LIBERO release and filtering out all unsuccessful ones, leaving 3917 successful demo trajectories. This is done via a modified version of a script from the MiniVLA codebase. In addition to… See the full description on the dataset page: https://huggingface.co/datasets/Embodied-CoT/embodied_features_and_demos_libero.embodied_reasoner
Embodied-Reasoner Dataset
Dataset Overview
Embodied-Reasoner is a multimodal reasoning dataset designed for embodied interactive tasks. It contains 9,390 Observation-Thought-Action trajectories for training and evaluating multimodal models capable of performing complex embodied tasks in indoor environments.
Key Features
📸 Rich Visual Data: Contains 64,000 first-person perspective interaction images🤔 Deep Reasoning Capabilities: 8 million thought… See the full description on the dataset page: https://huggingface.co/datasets/zwq2018/embodied_reasoner.iWorld-Bench-Dataset
iWorld-Bench Simulation Archives
News: Congratulations to the iWorldBench team! iWorldBench has been accepted to ICML 2026.
Important Dates
Paper accepted: May 1, 2026
Code release: May 18, 2026
Dataset release: May 19, 2026
iWorldBench is a benchmark for evaluating camera-controllable video generation models and interactive world models. This dataset repository hosts the packaged simulation archives associated with iWorld-Bench. It provides rendered simulation videos and… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/iWorld-Bench-Dataset.UrbanVideo-Bench
[ACL'25 Oral] UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
This repository contains the dataset introduced in the paper, consisting of two parts: 5k+ multiple-choice question-answering (MCQ) data and 1k+ video clips.
Arxiv: https://arxiv.org/pdf/2503.06157
Project: https://embodiedcity.github.io/UrbanVideo-Bench/
Code: https://github.com/EmbodiedCity/UrbanVideo-Bench.code
Dataset Description
The… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/UrbanVideo-Bench.phyground
PhyGround: Benchmarking Physical Reasoning in Generative World Models
Project page ·
Paper ·
Evaluation code ·
PhyJudge-9B
PhyGround is a criteria-grounded benchmark for diagnosing physical failures in
generated video. It contains 250 prompts covering 13 observable physical
laws across solid-body mechanics, fluid dynamics, and optics. Each prompt is
paired with a first-frame image, 10 released generation configurations, and
applicable-law labels.
The Hub repository includes:… See the full description on the dataset page: https://huggingface.co/datasets/NU-World-Model-Embodied-AI/phyground.ATARAEmbodiedEvalThis repository contains the dataset of the paper EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents.
Github repository: https://github.com/thunlp/EmbodiedEval
Project Page: https://embodiedeval.github.io/
AirScape-Dataset
[ACM MM'25] AirScape: An Aerial Generative World Model with Motion Controllability
This repository contains the dataset introduced in the paper, consisting of two parts: 11k+ motion intention prompts and corresponding video clips.
Arxiv: https://arxiv.org/pdf/2507.08885
Project: https://embodiedcity.github.io/AirScape/
Code: https://github.com/EmbodiedCity/AirScape.code
Dataset Description
This dataset is proposed for training and testing of aerial world models.… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/AirScape-Dataset.ANWM-Dataset
ANWM-Dataset
Training / evaluation trajectories for ANWM (Aerial Navigation World Model),
released with the paper
Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space.
Code: https://github.com/EmbodiedCity/ANWM.code
Model: EmbodiedCity/ANWM
Contents
Sharded tar archives of AirVLN-16 style trajectories (airvln_16-*.tar).
Each archive preserves the original folder layout ({ID}_processed/... with
images and traj_data.pkl).… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/ANWM-Dataset.Embodied-R1.5-RFT-Dataset
Embodied-R1.5-RFT-Dataset
🌐 Project Page |
📄 arXiv |
💻 Code |
🧰 EmbodiedEvalKit |
🤗 Models & Datasets
🗓️ Update — 2026-08-20 (20260820). All 28 Stage 2 RFT JSON annotation files have been uploaded to rft_datasets_json/. The complete JSON ↔ media archive mapping is documented in the Dataset composition table below.
⚠️ Partial release. This repository currently contains only a subset of the full Stage 2 RFT data… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1.5-RFT-Dataset.EB-ManipulationThis repository contains the data of the paper EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.
📄paper
💻 Github
🏠Website
EmbodiedGenRLv2-BGembodied-spatial-reasoning
Embodied Spatial Reasoning Tasks
Dataset Description
This dataset is part of the embodied-spatial-reasoning project, where the agent has to actively explore the environment to determine if certain spatial relationships hold true. The tasks involve spatial reasoning with various objects and scenes. Each task includes a query about the spatial relationships between objects within a scene, which the agent must verify through exploration.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/thanhqt2002/embodied-spatial-reasoning.vsi-benchEB-Alfred_trajectory_dataset
EB-Alfred trajectory dataset
We release the trajectory dataset collected from EmbodiedBench using several closed-source and open-source models. We hope this dataset will support the development of more capable embodied agents with improved perception, reasoning, and planning abilities. When using the trajectories, we recommend separating the training and evaluation sets—for example, using the “base” subset for training and other EmbodiedBench subsets for evaluation.
📖… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedBench/EB-Alfred_trajectory_dataset.EmbodiedRestore
EmbodiedRestore
Paired robotic first-frame observations (low-quality / ground-truth) under 25 distortions from the TID2013 / KADID-10k taxonomy, evaluated by three policies (π0.5, π0, OpenVLA). Built for benchmarking image restoration / IQA on robot-observation distributions, with downstream policy success rates(SR) and steps to successas(StS) as secondary signals.
To promote the development of image restoration model for robot vision systems, we will continue to maintain this… See the full description on the dataset page: https://huggingface.co/datasets/qruisjtu/EmbodiedRestore.ActiveFly-BenchScanReQAOpen3DVQA-v2EB-ALFREDSelected dataset for paper "EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents". The dataset originates from ALFRED. Modifications to task descriptions are in the github repo 💻 Github.
📄paper
💻 Github
🏠Website
Embodied-R1-Dataset
Embodied-R1-3B-v1
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation (ICLR 2026)
[🌐 Project Website] [📄 Paper] [🏆 ICLR2026 Version] [🎯 Dataset] [📦 Code]
Model Details
Model Description
Embodied-R1 is a 3B vision-language model (VLM) for general robotic manipulation.
It introduces a Pointing mechanism and uses Reinforced Fine-tuning (RFT) to bridge perception and action, with strong zero-shot generalization in embodied… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1-Dataset.BasicSpatialAbility
[ACL'25 Main] Defining and Evaluating Visual Language Models’ Basic Spatial Abilities: A Perspective from Psychometrics
[!IMPORTANT]
You can find the sample testing code on GitHub!
This dataset is a benchmark designed for evaluating Multimodal Large Language Models' Basic Spatial Abilities based on authentic Psychometric theories. It is structured specifically to support both Zero-shot and Few-shot evaluation protocols.
Split Name
Role
Description
test
Query Set… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/BasicSpatialAbility.
