datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EmbodiedGenDatahttps://huggingface.co/spaces/HorizonRobotics/EmbodiedGen-Gallery-Explorer
Hy-Embodied-0.5-VLA-Data
Hy-Embodied-0.5-VLA
From Vision-Language-Action Models to a Real-World Robot Learning Stack
Tencent Robotics X × Tencent Hy Team
📖 Abstract
We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data.Embodied-Captioning
Embodied Image Captioning – Manually Annotated Test Set
Paper: Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions (ICCV 2025)Authors: Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo NataleAffiliations: Italian Institute of Technology (IIT), University of GenoaProject Website: https://hsp-iit.github.io/embodied-captioningCode: https://github.com/hsp-iit/embodied-captioning
📦… See the full description on the dataset page: https://huggingface.co/datasets/TommyBsk/Embodied-Captioning.UrbanVideo-Bench
[ACL'25 Oral] UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
This repository contains the dataset introduced in the paper, consisting of two parts: 5k+ multiple-choice question-answering (MCQ) data and 1k+ video clips.
Arxiv: https://arxiv.org/pdf/2503.06157
Project: https://embodiedcity.github.io/UrbanVideo-Bench/
Code: https://github.com/EmbodiedCity/UrbanVideo-Bench.code
Dataset Description
The… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/UrbanVideo-Bench.phyground
PhyGround: Benchmarking Physical Reasoning in Generative World Models
Project page ·
Paper ·
Evaluation code ·
PhyJudge-9B
PhyGround is a criteria-grounded benchmark for diagnosing physical failures in
generated video. It contains 250 prompts covering 13 observable physical
laws across solid-body mechanics, fluid dynamics, and optics. Each prompt is
paired with a first-frame image, 10 released generation configurations, and
applicable-law labels.
The Hub repository includes:… See the full description on the dataset page: https://huggingface.co/datasets/NU-World-Model-Embodied-AI/phyground.EmbodiedEvalThis repository contains the dataset of the paper EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents.
Github repository: https://github.com/thunlp/EmbodiedEval
Project Page: https://embodiedeval.github.io/
AirScape-Dataset
[ACM MM'25] AirScape: An Aerial Generative World Model with Motion Controllability
This repository contains the dataset introduced in the paper, consisting of two parts: 11k+ motion intention prompts and corresponding video clips.
Arxiv: https://arxiv.org/pdf/2507.08885
Project: https://embodiedcity.github.io/AirScape/
Code: https://github.com/EmbodiedCity/AirScape.code
Dataset Description
This dataset is proposed for training and testing of aerial world models.… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/AirScape-Dataset.EmbodiedGenRLv2-BGembodied-spatial-reasoning
Embodied Spatial Reasoning Tasks
Dataset Description
This dataset is part of the embodied-spatial-reasoning project, where the agent has to actively explore the environment to determine if certain spatial relationships hold true. The tasks involve spatial reasoning with various objects and scenes. Each task includes a query about the spatial relationships between objects within a scene, which the agent must verify through exploration.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/thanhqt2002/embodied-spatial-reasoning.vsi-benchEmbodiedRestore
EmbodiedRestore
Paired robotic first-frame observations (low-quality / ground-truth) under 25 distortions from the TID2013 / KADID-10k taxonomy, evaluated by three policies (π0.5, π0, OpenVLA). Built for benchmarking image restoration / IQA on robot-observation distributions, with downstream policy success rates(SR) and steps to successas(StS) as secondary signals.
To promote the development of image restoration model for robot vision systems, we will continue to maintain this… See the full description on the dataset page: https://huggingface.co/datasets/qruisjtu/EmbodiedRestore.ScanReQAEmbodied-R1-Dataset
Embodied-R1-3B-v1
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation (ICLR 2026)
[🌐 Project Website] [📄 Paper] [🏆 ICLR2026 Version] [🎯 Dataset] [📦 Code]
Model Details
Model Description
Embodied-R1 is a 3B vision-language model (VLM) for general robotic manipulation.
It introduces a Pointing mechanism and uses Reinforced Fine-tuning (RFT) to bridge perception and action, with strong zero-shot generalization in embodied… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1-Dataset.BasicSpatialAbility
[ACL'25 Main] Defining and Evaluating Visual Language Models’ Basic Spatial Abilities: A Perspective from Psychometrics
[!IMPORTANT]
You can find the sample testing code on GitHub!
This dataset is a benchmark designed for evaluating Multimodal Large Language Models' Basic Spatial Abilities based on authentic Psychometric theories. It is structured specifically to support both Zero-shot and Few-shot evaluation protocols.
Split Name
Role
Description
test
Query Set… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/BasicSpatialAbility.EmbodiedAgentInterface
Dataset Card for Embodied Agent Interface
This is the dataset used in our project Embodied Agent Interface.
Jai-World-VRM-3D-Embodied-AI
Jai World - VRM 3D Embodied AI
Interact with an AI powered 3D vrm avatar in a virtual world. A lightweight desktop folder-based Flask app that supports both Ollama and OpenRouter.
This project explores personalized 3D world generation, embodied AI, AI spatial navigation, and virtual character interaction. This code is a working example. It's a starting point that can be tuned and expanded.
Tech stack:
Three.js + HTML + CSS + JS + Flask + Ollama/OpenRouter (qwen3.5:9b/… See the full description on the dataset page: https://huggingface.co/datasets/vbookshelf/Jai-World-VRM-3D-Embodied-AI.PaperBench-X-Embodied-Manipulation-Pi05-LIBERO
PaperBench-X — pi05 / LIBERO
Standalone embodied manipulation public-demo package for the released openpi
pi05_libero checkpoint on all four LIBERO suites. This repository intentionally
contains no navigation tasks and is not the general manipulation bundle.
What is here
path
content
task/
complete Harbor task, evaluator, rubric, scoring regression tests and report builder
raw_task/
canonical source used by convert_raw_to_rp.py
demo/
static… See the full description on the dataset page: https://huggingface.co/datasets/SueMintony/PaperBench-X-Embodied-Manipulation-Pi05-LIBERO.EmbodiedNav-Bench
[KDD'26 Oral] How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace
EmbodiedNav-Bench is a goal-oriented embodied navigation benchmark for evaluating spatial action in urban 3D airspace. The benchmark contains 5,037 high-quality navigation trajectories with natural-language navigation goals, initial drone poses, target positions, and ground-truth 3D trajectories.
This Hugging Face repository… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/EmbodiedNav-Bench.embodied-spatial-resoningEmbodiedVerse-BenchEmbodied-CoTembodied-ai-conference-papers-2023-2025
Embodied AI Conference Papers 2023–2025
A dataset of 3,140 papers from 12 robotics conferences (CoRL 2023–2025, ICRA 2023–2025, IROS 2023–2025, RSS 2023–2025), processed with KNOWHERE.
Series
This collection of KNOWHERE-processed paper datasets is continuously expanding. New datasets and editions will be released — stay tuned!
Top ML Conference Papers
2023
https://huggingface.co/datasets/JensCS/top-ml-conference-papers-2023… See the full description on the dataset page: https://huggingface.co/datasets/JensCS/embodied-ai-conference-papers-2023-2025.ChenLong_Embodied_Intelligence_Dataset
ChenLong Embodied Intelligence Dataset
本仓库用于统一管理辰龙机器人实习中的数据集、模型权重、训练结果和说明文档。后续新增不同任务、采集批次、模型版本或实验资源时,都放在这里统一维护。
当前目录
embodied_dataset/:具身智能采集数据集,采用 LeRobot v3.0 结构,包含 data/、meta/、videos/。
yolo_dataset/:YOLO 目标检测数据、模型权重、训练参数和评估结果,当前包含 blue_bucket_yolov8/。
待新增新的数据集或模型。
新增数据集要求
具身数据优先采用 LeRobot v3.0 格式:meta/info.json、meta/stats.json、tasks、episodes、逐帧 Parquet 数据和按相机划分的视频。新增数据集至少写清:
任务:任务文本、目标物、成功标准、失败标准。
硬件:机器人型号、自由度、夹爪、相机位置、分辨率、FPS。… See the full description on the dataset page: https://huggingface.co/datasets/vvzc/ChenLong_Embodied_Intelligence_Dataset.robotics-embodied-ai-physical-control-2026
🤖 Robotics, Embodied AI & Physical World Control Dataset (2026 Edition)
A structured research dataset featuring 2,123 domain-verified research papers and official code repositories focused on Vision-Language-Action Models (VLA), Humanoid Robotics, Quadruped Locomotion, Diffusion Policies, Sim-to-Real Transfer, and Physics Simulation Environments (Isaac Sim, MuJoCo, Genesis).
Built with Universal Scientific Engine V16.1 Gold, providing 43 schema attributes with verified… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/robotics-embodied-ai-physical-control-2026.embodied-perceptual-state-integrity-v0.1Embodied Perceptual State Integrity v0.1
What this tests
Whether an embodied agent preserves a coherent internal world state across movement, delay, and perceptual absence.
Failure modes
state_driftThe response contradicts the true state at time t1
fabricated_updateThe response claims a state change when the world facts did not change
state_integrityThe response states the correct t1 state without contradiction
How it works
world_facts_t0 provides initial ground truth
actions_between… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/embodied-perceptual-state-integrity-v0.1.embodied-and-math-problemsembodied_web_agent_outdoor_trajectoryembodied-constraint-aware-recovery-v0.1Embodied Constraint-Aware Recovery v0.1
What this tests
Whether an embodied agent recovers from failed or blocked actions without inventing success and without looping blindly.
Failure modes
false_successResponse claims completion despite a failure outcome
non_adaptive_repeatResponse repeats the same failed strategy without change
recovery_okResponse proposes a constraint-aware alternative step
How it works
world_facts_t0 defines the initial state
goal defines intent
action_attempt is the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/embodied-constraint-aware-recovery-v0.1.embodied-decision-integrity-v01Embodied Decision Integrity v0.1
What this dataset is
This dataset evaluates decision quality before motion in embodied robotic systems.
You give the model a snapshot of the world.
Sensors.
Conflicts.
Constraints.
You ask it to decide what to do next.
Not how to move.
Whether to move at all.
Why this matters
Most robotics failures are not control failures.
They are judgment failures.
Robots fail when they:
Act while perception is unresolved
Commit under uncertainty
Ignore safety margins
Choose… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/embodied-decision-integrity-v01.embodied-action-outcome-coherence-v0.1Embodied Action–Outcome Coherence v0.1
What this tests
Whether an embodied agent updates world state from observed outcomes
Whether it avoids claiming success when the outcome says failure
Failure modes
outcome_ignoredResponse does not reflect the true post-action state
false_successResponse claims success despite an observed failure
causal_update_okResponse states the correct post-action state without contradiction
How it works
world_facts_t0 is the initial state
action_taken is what the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/embodied-action-outcome-coherence-v0.1.
