CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HorizonRobotics /EmbodiedGenDatahttps://huggingface.co/spaces/HorizonRobotics/EmbodiedGen-Gallery-Explorer 3d1K<n<10K7 likes149k downloads3mo agoHugging Face02tencent /Hy-Embodied-0.5-VLA-Data Hy-Embodied-0.5-VLA From Vision-Language-Action Models to a Real-World Robot Learning Stack Tencent Robotics X × Tencent Hy Team 📖 Abstract We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data.tabularroboticsn<1K23 likes81k downloads3mo agoHugging Face03HuggingFriends /mllm-as-embodied-world-judge MLLM-as-Embodied-World-Judge Data for judging physical adherence and instruction alignment of generated embodied-manipulation videos. Start here path what it is final/ the current release — train.jsonl (11,520), test.jsonl (802), and its README data/ source and generated videos, referenced by video_url in the splits Benchmark tooling path what it is bench/LEADERBOARD.md judge results table bench/TESTSET.md benchmark… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/mllm-as-embodied-world-judge.2 likes17k downloads15d agoHugging Face04TommyBsk /Embodied-Captioning Embodied Image Captioning – Manually Annotated Test Set Paper: Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions (ICCV 2025)Authors: Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo NataleAffiliations: Italian Institute of Technology (IIT), University of GenoaProject Website: https://hsp-iit.github.io/embodied-captioningCode: https://github.com/hsp-iit/embodied-captioning 📦… See the full description on the dataset page: https://huggingface.co/datasets/TommyBsk/Embodied-Captioning.tabularimage-to-text1K<n<10K0 likes8.8k downloads1y agoHugging Face05nvidia /PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card Dataset Description PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.video100K<n<1M25 likes7.3k downloads4mo agoHugging Face06EricsXt /Xlang-Embodied-dataset0 likes3.9k downloads1y agoHugging Face07IffYuan /Embodied-R1.5-SFT-Dataset Embodied-R1.5-SFT-Dataset 🌐 Project Page &nbsp;|&nbsp; 📄 arXiv &nbsp;|&nbsp; 💻 Code &nbsp;|&nbsp; 🧰 EmbodiedEvalKit &nbsp;|&nbsp; 🤗 Models & Datasets 🗓️ Update — 2026-08-20 (20260820). All 34 Stage 1 SFT JSON annotation files have been uploaded to sft_datasets_json/. The complete JSON ↔ image/video data mapping is documented in the Dataset composition table below. ⚠️ Partial release. This repository currently contains only a subset of the full Stage 1 SFT… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1.5-SFT-Dataset.imageimage-text-to-text10 likes3.2k downloads1mo agoHugging Face08onandon /EmbodiedSplat EmbodiedSplat 🛋️ Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding Seungjun Lee · Zihan Wang · Yunsong Wang · Gim Hee Lee National University of Singapore CVPR 2026 Code | Paper | Project Page Build and understand at Once! By taking over 300 streaming images, our EmbodiedSplat reconstructs whole-scene open-vocabulary 3DGS in online manner at up to 5-6 FPS per-frame processing time. Reconstructed scene… See the full description on the dataset page: https://huggingface.co/datasets/onandon/EmbodiedSplat.0 likes1.9k downloads4mo agoHugging Face09Embodied-CoT /embodied_features_and_demos_liberoDataset for Embodied Chain-of-Thought Reasoning for LIBERO-90, as used by ECoT-Lite. TFDS Demonstration Data The TFDS dataset contains successful demonstration trajectories for LIBERO-90 (50 trajectories for each of 90 tasks). It was created by rolling out the actions provided in the original LIBERO release and filtering out all unsuccessful ones, leaving 3917 successful demo trajectories. This is done via a modified version of a script from the MiniVLA codebase. In addition to… See the full description on the dataset page: https://huggingface.co/datasets/Embodied-CoT/embodied_features_and_demos_libero.robotics4 likes1.8k downloads6mo agoHugging Face10zwq2018 /embodied_reasoner Embodied-Reasoner Dataset Dataset Overview Embodied-Reasoner is a multimodal reasoning dataset designed for embodied interactive tasks. It contains 9,390 Observation-Thought-Action trajectories for training and evaluating multimodal models capable of performing complex embodied tasks in indoor environments. Key Features 📸 Rich Visual Data: Contains 64,000 first-person perspective interaction images🤔 Deep Reasoning Capabilities: 8 million thought… See the full description on the dataset page: https://huggingface.co/datasets/zwq2018/embodied_reasoner.imageimage-text-to-text10K<n<100K22 likes1.8k downloads1y agoHugging Face11EmbodiedCity /iWorld-Bench-Dataset iWorld-Bench Simulation Archives News: Congratulations to the iWorldBench team! iWorldBench has been accepted to ICML 2026. Important Dates Paper accepted: May 1, 2026 Code release: May 18, 2026 Dataset release: May 19, 2026 iWorldBench is a benchmark for evaluating camera-controllable video generation models and interactive world models. This dataset repository hosts the packaged simulation archives associated with iWorld-Bench. It provides rendered simulation videos and… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/iWorld-Bench-Dataset.1 likes1.7k downloads3mo agoHugging Face12EmbodiedCity /UrbanVideo-Bench [ACL'25 Oral] UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces This repository contains the dataset introduced in the paper, consisting of two parts: 5k+ multiple-choice question-answering (MCQ) data and 1k+ video clips. Arxiv: https://arxiv.org/pdf/2503.06157 Project: https://embodiedcity.github.io/UrbanVideo-Bench/ Code: https://github.com/EmbodiedCity/UrbanVideo-Bench.code Dataset Description The… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/UrbanVideo-Bench.textvisual-question-answering1K<n<10K9 likes1.6k downloads2mo agoHugging Face13NU-World-Model-Embodied-AI /phyground PhyGround: Benchmarking Physical Reasoning in Generative World Models Project page · Paper · Evaluation code · PhyJudge-9B PhyGround is a criteria-grounded benchmark for diagnosing physical failures in generated video. It contains 250 prompts covering 13 observable physical laws across solid-body mechanics, fluid dynamics, and optics. Each prompt is paired with a first-frame image, 10 released generation configurations, and applicable-law labels. The Hub repository includes:… See the full description on the dataset page: https://huggingface.co/datasets/NU-World-Model-Embodied-AI/phyground.texttext-to-videon<1K2 likes1.3k downloads2mo agoHugging Face14embodiedfoundation /ATARAvideon<1K1 likes1.3k downloads1y agoHugging Face15EmbodiedEval /EmbodiedEvalThis repository contains the dataset of the paper EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents. Github repository: https://github.com/thunlp/EmbodiedEval Project Page: https://embodiedeval.github.io/ 3droboticsn<1K3 likes922 downloads2y agoHugging Face16EmbodiedCity /AirScape-Dataset [ACM MM'25] AirScape: An Aerial Generative World Model with Motion Controllability This repository contains the dataset introduced in the paper, consisting of two parts: 11k+ motion intention prompts and corresponding video clips. Arxiv: https://arxiv.org/pdf/2507.08885 Project: https://embodiedcity.github.io/AirScape/ Code: https://github.com/EmbodiedCity/AirScape.code Dataset Description This dataset is proposed for training and testing of aerial world models.… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/AirScape-Dataset.textsummarization10K<n<100K6 likes834 downloads2mo agoHugging Face17EmbodiedCity /ANWM-Dataset ANWM-Dataset Training / evaluation trajectories for ANWM (Aerial Navigation World Model), released with the paper Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space. Code: https://github.com/EmbodiedCity/ANWM.code Model: EmbodiedCity/ANWM Contents Sharded tar archives of AirVLN-16 style trajectories (airvln_16-*.tar). Each archive preserves the original folder layout ({ID}_processed/... with images and traj_data.pkl).… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/ANWM-Dataset.robotics0 likes805 downloads1mo agoHugging Face18IffYuan /Embodied-R1.5-RFT-Dataset Embodied-R1.5-RFT-Dataset 🌐 Project Page &nbsp;|&nbsp; 📄 arXiv &nbsp;|&nbsp; 💻 Code &nbsp;|&nbsp; 🧰 EmbodiedEvalKit &nbsp;|&nbsp; 🤗 Models & Datasets 🗓️ Update — 2026-08-20 (20260820). All 28 Stage 2 RFT JSON annotation files have been uploaded to rft_datasets_json/. The complete JSON ↔ media archive mapping is documented in the Dataset composition table below. ⚠️ Partial release. This repository currently contains only a subset of the full Stage 2 RFT data… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1.5-RFT-Dataset.imageimage-text-to-text2 likes730 downloads1mo agoHugging Face19EmbodiedBench /EB-ManipulationThis repository contains the data of the paper EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents. 📄paper 💻 Github 🏠Website 4 likes700 downloads2y agoHugging Face20xinjjj /EmbodiedGenRLv2-BG3dn<1K0 likes630 downloads1y agoHugging Face21thanhqt2002 /embodied-spatial-reasoning Embodied Spatial Reasoning Tasks Dataset Description This dataset is part of the embodied-spatial-reasoning project, where the agent has to actively explore the environment to determine if certain spatial relationships hold true. The tasks involve spatial reasoning with various objects and scenes. Each task includes a query about the spatial relationships between objects within a scene, which the agent must verify through exploration. Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/thanhqt2002/embodied-spatial-reasoning.imagevisual-question-answering1K<n<10K1 likes591 downloads2y agoHugging Face22Embodied1 /vsi-benchimage1K<n<10K0 likes584 downloads9mo agoHugging Face23EmbodiedBench /EB-Alfred_trajectory_dataset EB-Alfred trajectory dataset We release the trajectory dataset collected from EmbodiedBench using several closed-source and open-source models. We hope this dataset will support the development of more capable embodied agents with improved perception, reasoning, and planning abilities. When using the trajectories, we recommend separating the training and evaluation sets—for example, using the “base” subset for training and other EmbodiedBench subsets for evaluation. 📖… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedBench/EB-Alfred_trajectory_dataset.2 likes504 downloads1y agoHugging Face24qruisjtu /EmbodiedRestore EmbodiedRestore Paired robotic first-frame observations (low-quality / ground-truth) under 25 distortions from the TID2013 / KADID-10k taxonomy, evaluated by three policies (π0.5, π0, OpenVLA). Built for benchmarking image restoration / IQA on robot-observation distributions, with downstream policy success rates(SR) and steps to successas(StS) as secondary signals. To promote the development of image restoration model for robot vision systems, we will continue to maintain this… See the full description on the dataset page: https://huggingface.co/datasets/qruisjtu/EmbodiedRestore.imageimage-to-image10K<n<100K0 likes460 downloads5mo agoHugging Face25EmbodiedCity /ActiveFly-Bench0 likes458 downloads2mo agoHugging Face26EmbodiedCity /ScanReQAtext10K<n<100K0 likes422 downloads4mo agoHugging Face27EmbodiedCity /Open3DVQA-v22 likes363 downloads2mo agoHugging Face28EmbodiedBench /EB-ALFREDSelected dataset for paper "EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents". The dataset originates from ALFRED. Modifications to task descriptions are in the github repo 💻 Github. 📄paper 💻 Github 🏠Website 2 likes308 downloads2y agoHugging Face29IffYuan /Embodied-R1-Dataset Embodied-R1-3B-v1 Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation (ICLR 2026) [🌐 Project Website] [📄 Paper] [🏆 ICLR2026 Version] [🎯 Dataset] [📦 Code] Model Details Model Description Embodied-R1 is a 3B vision-language model (VLM) for general robotic manipulation. It introduces a Pointing mechanism and uses Reinforced Fine-tuning (RFT) to bridge perception and action, with strong zero-shot generalization in embodied… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1-Dataset.text100K<n<1M0 likes302 downloads7mo agoHugging Face30EmbodiedCity /BasicSpatialAbility [ACL'25 Main] Defining and Evaluating Visual Language Models’ Basic Spatial Abilities: A Perspective from Psychometrics [!IMPORTANT] You can find the sample testing code on GitHub! This dataset is a benchmark designed for evaluating Multimodal Large Language Models' Basic Spatial Abilities based on authentic Psychometric theories. It is structured specifically to support both Zero-shot and Few-shot evaluation protocols. Split Name Role Description test Query Set… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/BasicSpatialAbility.imagevisual-question-answeringn<1K0 likes231 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.