CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tencent /Hy-Embodied-0.5-VLA-Data Hy-Embodied-0.5-VLA From Vision-Language-Action Models to a Real-World Robot Learning Stack Tencent Robotics X × Tencent Hy Team 📖 Abstract We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data.tabularroboticsn<1K23 likes81k downloads3mo agoHugging Face02depth2world /VLADBenchimage1K<n<10K3 likes51k downloads8mo agoHugging Face03vLAR /LavalObjaverseDataset Laval Objaverse Dataset vLAR Group | SIGGRAPH Asia 2026 A large-scale, high-quality dataset for multi-view relighting. 📖 Dataset Summary The Laval Objaverse Dataset is a comprehensive dataset designed for multi-view relighting and novel view synthesis tasks. It combines high-quality 3D assets from Objaverse with realistic, diverse illumination conditions… See the full description on the dataset page: https://huggingface.co/datasets/vLAR/LavalObjaverseDataset.3dimage-to-image10M<n<100M11 likes39k downloads0m agoHugging Face04vLAR /PhysInOnePhysInOne: Visual Physics Learning and Reasoning in One Suite vLAR Group | The Hong Kong Polytechnic University | Syai Singapore | Meta CVPR 2026 🧭 Navigation 📌 Summary 🚀 Release Timetable 📦 Repositories & Downloads 📊 Data Splits 🧱 3D Assets 🛠️ Data Processing 🏆 Leaderboard Evaluation Data 🎞️ Rendered Data &nbsp;&nbsp;&nbsp;1. Download Scripts &nbsp;&nbsp;&nbsp;2. Install Dependencies… See the full description on the dataset page: https://huggingface.co/datasets/vLAR/PhysInOne.videotext-to-video1M<n<10M20 likes33k downloads4d agoHugging Face05VLABench /vlabench_primitive_pretrain_lerobot Datacard This is the official VLABench primitive pretraining dataset converted to the LeRobot format. The dataset contains language-conditioned manipulation trajectories collected with a Franka Panda robot in VLABench simulation. This LeRobot version is hosted at: https://huggingface.co/datasets/VLABench/vlabench_primitive_pretrain_lerobot Source Project Page: https://vlabench.github.io/ Arxiv Paper: https://arxiv.org/abs/2412.18194 Code:… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlabench_primitive_pretrain_lerobot.0 likes30k downloads3mo agoHugging Face06mikasa-robo /mikasa-robo-vla-rlds0 likes17k downloads3mo agoHugging Face07UCSC-VLAA /gpt-edit-simplerimage1M<n<10M13 likes14k downloads1y agoHugging Face08mim-chess-vlas /eval-results0 likes12k downloads4m agoHugging Face09lerobot /vlabench-assets1 likes10k downloads5mo agoHugging Face10UCSC-VLAA /Recap-DataComp-1B Dataset Card for Recap-DataComp-1B Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions. Dataset Details Dataset Description Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.imagezero-shot-classification1B<n<10B205 likes9.8k downloads2y agoHugging Face11ricl-vla /collected_demos_trainingimage10K<n<100K0 likes8.2k downloads1y agoHugging Face12UCSC-VLAA /GPT-Image-Edit-1.5M GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset 📃Arxiv | 🌐 Project Page | 💻Github GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1. 📣 News [2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download. [2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.imageimage-to-image1M<n<10M90 likes7.2k downloads1y agoHugging Face13LejuRobotics /LET-KUAVO-VLA-1.0-Datasetgated LET-KUAVO-VLA-1.0-Dataset videon<1K3 likes7.1k downloads14d agoHugging Face14BiggerXu /rabench-vlabench-unified-libero-v1 VLABench Data Prep This directory contains an independent, non-Docker data conversion workflow for turning lerobot/libero into an episode-based HDF5 format that is easier for RABench agents to consume. Goal The source lerobot/libero dataset is distributed as: parquet tables for numeric columns mp4 video shards for image streams separate metadata parquet files for tasks and episode boundaries That structure is compact, but it is awkward for an agent to discover and use… See the full description on the dataset page: https://huggingface.co/datasets/BiggerXu/rabench-vlabench-unified-libero-v1.0 likes7.1k downloads6mo agoHugging Face15sam-guided-vlas /eval-resultsvideo100K<n<1M0 likes6k downloads9d agoHugging Face16InternRobotics /VLAC-Cut-Benchmark Video Progress Benchmark Paper · Code · Model · Benchmark Overview The Video Progress Benchmark (VPB) evaluates process-level task progress estimation for robot manipulation. It measures whether a model can capture advancement, stagnation, regression, and recovery throughout an execution video. VPB is built from the held-out portion of the Progress Annotation Dataset. Unlike endpoint-only evaluations, VPB focuses on temporal task progress and supports analysis… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/VLAC-Cut-Benchmark.1 likes5.9k downloads2mo agoHugging Face17ShuaiYang03 /VLA_Instruction_TuningThis repository contains the VLA-IT dataset, a curated 650K-sample Vision-Language-Action Instruction Tuning dataset, and the SimplerEnv-Instruct benchmark. These are presented in the paper InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation. The dataset is designed to enable robots to integrate multimodal reasoning with precise action generation, preserving the flexible reasoning of large vision-language models while delivering leading manipulation… See the full description on the dataset page: https://huggingface.co/datasets/ShuaiYang03/VLA_Instruction_Tuning.robotics4 likes5.3k downloads1y agoHugging Face18VLABench /vlabench_primitive_ft_lerobot_video VLABench Primitive Tasks Dataset - LeRobot v3.0 Dataset Description This dataset is organized in the LeRobot v3.0 format and is used for integrating VLABench into the LeRobot framework officially. Compared with the v2.0 version and the RLDS version of the dataset, this release stores visual observations in a video-compressed format rather than as individual image files. This design provides significant advantages in both storage efficiency and data loading performance.… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlabench_primitive_ft_lerobot_video.tabularrobotics100K<n<1M3 likes4.2k downloads5mo agoHugging Face19InternRobotics /VLAC-Cut-FullData VLAC-Cut-FullData VLAC-Cut-FullData is the full-data release for VLAC-Cut. It provides the complete raw-data archive set, benchmark-style JSON files, and a lightweight frame-extraction workflow for reproducing evaluation on the released benchmark protocol. Contents benchmark_style_all/ train/video_progress_benchmark_file.json test_expert_seen/video_progress_benchmark_file.json test_expert_unseen/video_progress_benchmark_file.json… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/VLAC-Cut-FullData.3 likes3.7k downloads2mo agoHugging Face20UCSC-VLAA /HQ-Edit Dataset Card for HQ-EDIT HQ-Edit, a high-quality instruction-based image editing dataset with total 197,350 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely GPT-4V and DALL-E 3. HQ-Edit’s high-resolution images, rich in detail and accompanied by comprehensive editing prompts, substantially enhance the capabilities of existing image editing… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/HQ-Edit.image1K<n<10K42 likes3.6k downloads2y agoHugging Face21VLABench /vlabench_composite_ft_lerobot_videotabular1M<n<10M0 likes3.5k downloads9mo agoHugging Face22VLABench /vlm_evaluation_v1.0 Datacard This dataset is the evaluation VLM dataset used in VLABench. It is designed to evaluate the planning capabilities of Vision-Language Models (VLMs) in embodied scenarios. Source Project Page: https://vlabench.github.io/ Arxiv Paper: https://arxiv.org/abs/2412.18194 Code: https://github.com/OpenMOSS/VLABench Uses The dataset structure is as follows: vlm_evaluation_v1.0/ ├── CommenSence/ ├── add_condiment_common_sense/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlm_evaluation_v1.0.image1K<n<10K0 likes3.5k downloads1y agoHugging Face23christian0420 /drifting-vla-v2-droid0 likes3.4k downloads6mo agoHugging Face24VladPyatov /CADFS CADFS Dataset CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models 🚀 Project Page 📃 Paper 💻 Code 🤗 Model A large-scale dataset for parametric CAD model generation from text descriptions and multi-view images. Models are represented as FeatureScript programs, enabling direct import into Onshape environment. This dataset was used to train and evaluate CADFS-2B, a fine-tuned Qwen2-VL-2B multimodal language model… See the full description on the dataset page: https://huggingface.co/datasets/VladPyatov/CADFS.3dtext-to-3d100K<n<1M6 likes3.2k downloads2mo agoHugging Face25VITRA-VLA /VITRA-1M VITRA-1M: Human Hand V-L-A Dataset Dataset Summary VITRA-1M is a large-scale Human Hand Visual-Language-Action (V-L-A) dataset constructed as described in the paper Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos. It contains 1.2 million short episodes with segmented language annotations, camera parameters (corrected intrinsics/extrinsics), and 3D hand reconstructions (left and right… See the full description on the dataset page: https://huggingface.co/datasets/VITRA-VLA/VITRA-1M.1M<n<10M29 likes2.9k downloads10mo agoHugging Face26SpencerSon2001 /AIR-VLA_hdf5_datasets0 likes2.7k downloads5mo agoHugging Face27VR-VLA /VR-egoverse-annotation-curated-v6.0 VR-egoverse-annotation-full (egoverse_v06) Egocentric VR hand-tracking corpus: per-frame hand pose, wrist/camera trajectories, and language narratives paired with head-mounted video clips. Schema 0.6.0, generated 2026-07-23, tooling commit e885040. Private, in-progress upload. This release is being synced from local storage in the background, incrementally and in chunks (see Layout below); clip counts on the Hub will grow until the sync catches up to the full local release.… See the full description on the dataset page: https://huggingface.co/datasets/VR-VLA/VR-egoverse-annotation-curated-v6.0.videorobotics10K<n<100K0 likes2.7k downloads18d agoHugging Face28VLyb /VLABench VLABench VLM Evaluation Dataset This dataset is the VLM evaluation split of VLABench, prepared for reproducible VLABench evaluation with PhysBrainEvalKit. Source Project page: https://vlabench.github.io/ Paper: https://arxiv.org/abs/2412.18194 Official code: https://github.com/OpenMOSS/VLABench Directory layout The dataset is organized by evaluation dimension and subtask: vlm_evaluation_v1.0/ ├── CommenSence/ ├── Complex/ ├── M&T/ ├── PhysicsLaw/… See the full description on the dataset page: https://huggingface.co/datasets/VLyb/VLABench.0 likes2.5k downloads12d agoHugging Face29VLABench /vlabench_primitive_ft_lerobotimage100K<n<1M2 likes2.4k downloads1y agoHugging Face30Linslab /VLA-OS-Dataset Dataset Card This is the training dataset used in the paper VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models. Source Project Page: https://nus-lins-lab.github.io/vlaos/ Paper: https://arxiv.org/abs/2506.17561 Code: https://github.com/HeegerGao/VLA-OS Model: https://huggingface.co/Linslab/VLA-OS Usage Ensure you have installed git lfs: curl -s… See the full description on the dataset page: https://huggingface.co/datasets/Linslab/VLA-OS-Dataset.4 likes2.3k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.