CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /PhysicalAI-SmartSpaces Physical AI Smart Spaces Dataset Overview Comprehensive, annotated dataset for multi-camera tracking and 2D/3D object detection. This dataset is synthetically generated with Omniverse and Cosmos Transfer. This dataset consists of over 280 hours of video from across nearly 1,800 cameras from indoor scenes in warehouses, hospitals, retail, and more. The dataset is time synchronized for tracking humans, forklifts, pallet trucks and Autonomous Mobile Robots (AMRs)… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SmartSpaces.video1K<n<10K88 likes91k downloads2mo agoHugging Face02cvml-nus /assembly101gated Assembly101 Assembly101 is a procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations in action ordering, mistakes, and corrections. Assembly101 is the first multi-view action dataset, with simultaneous static (8) and egocentric (4) recordings. Sequences are annotated with more than 100K coarse and 1M fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/cvml-nus/assembly101.textn<1K19 likes70k downloads3mo agoHugging Face03nvidia /PhysicalAI-Robotics-Locomanipulation-GRAIL 📢 News [2026-07-15] Released task-general tracking policy checkpoints trained on the released data. Follow the tracking doc to use them to track our released motion data. [2026-07-14] Updated data/pickup_table and data/pickup_ground. If you downloaded them before this date, please re-download. Dataset Overview Tabletop Pickup Ground Pickup Tabletop Manipulation Ground Manipulation Sitting Curb Slope… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Locomanipulation-GRAIL.image1K<n<10K29 likes69k downloads21d agoHugging Face04nvidia /PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios Dataset Description: PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.video100K<n<1M26 likes55k downloads4mo agoHugging Face05nkp37 /OpenVid-1M Summary This is the dataset proposed in our paper [ICLR 2025] OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation. OpenVid-1M is a high-quality text-to-video dataset designed for research institutions to enhance video quality, featuring high aesthetics, clarity, and resolution. It can be used for direct training or as a quality tuning complement to other video datasets. All videos in the OpenVid-1M dataset have resolutions of at least 512×512.… See the full description on the dataset page: https://huggingface.co/datasets/nkp37/OpenVid-1M.videotext-to-video1M<n<10M284 likes50k downloads6mo agoHugging Face06McGill-NLP /WebLINX-full WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WARNING: This is not the main WebLINX data card! You might want to use the main WebLINX data card instead: WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WebLINX: Real-World Website Navigation with Multi-Turn Dialogue Xing Han Lù*, Zdeněk Kasner*, Siva Reddy 💾Code 📄Paper 🌐Website 📓Colab 🤖Models 💻Explorer 🐦Tweets 🏆Leaderboard Your browser does not support the… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX-full.text10K<n<100K8 likes29k downloads1y agoHugging Face07nvidia /PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes PhysicalAI SDG-Warehouse PhysicalAI SDG-Warehouse is a synthetic, fully-annotated video dataset of staged industrial-safety events captured in a simulated warehouse environment. It contains approximately 123k video clips, totaling roughly 412 hours of footage at 1920x1080 resolution and 30 frames per second, organized across four scenarios: a forklift near-miss with a human worker, a warehouse fire with worker evacuation, a forklift collision with a storage shelf, and a routine… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes.videovideo-classification100K<n<1M20 likes24k downloads4mo agoHugging Face08nvidia /PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes Dataset Description: The SDG-SynHuman is a large-scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post-training of NVIDIA Cosmos world foundation models and related physical AI research. Each sample is a temporally coherent 60-120 second video clip rendered at 1080p and 30 fps. Clips contain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes.video41 likes23k downloads4mo agoHugging Face09MCG-NJU /VideoChat3-LV116k VideoChat3-LV116K VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video data with supervision over longer temporal contexts, where evidence can be sparse, delayed, and distributed across multiple video segments. The dataset is constructed through a long-video synthesis pipeline. Candidate long videos are filtered for visual quality, semantic content, and temporal coherence. Videos are then split into manageable… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k.textvideo-text-to-text1K<n<10K15 likes22k downloads2mo agoHugging Face10NTU-yiwen /code-world-model-project-page-videos Code World Model Project Page Videos Public research-demo video assets used by the Code World Model project page. The gallery/ directory contains aligned RGB and proxy videos for interactive comparison. videon<1K0 likes19k downloads1mo agoHugging Face11nvidia /PhysicalAI-Robotics-GR00T-Teleop-Sim Simulation GR1 Tabletop Task 1K Dataset Dataset Description: The PhysicalAI-Robotics-GR00T-Teleop-GR1 dataset consists of 1000 teleoperation trajectories in simulation using the GR1 robot with upper body control. The simulation setup mimics tabletop manipulation tasks and uses RGB observations with a virtual camera. The robot is equipped with simulated Fourier hands. This dataset is ready for non-commercial use. Dataset Owner(s): NVIDIA GEAR… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-Sim.tabular1M<n<10M19 likes19k downloads9mo agoHugging Face12nvidia /PhysicalAI-Robotics-GR00T-Teleop-GR1 Introduction TL;DR: DreamDojo is a generalist robot world model pretrained on 44k hours of human egocentric data, showing unprecedented generalization to diverse objects and environments. Project page: https://dreamdojo-world.github.io/ Paper: https://arxiv.org/abs/2602.06949 Code: https://github.com/NVIDIA/DreamDojo How to Use Check out https://github.com/NVIDIA/DreamDojo Citation @article{gao2026dreamdojo, title={DreamDojo: A Generalist Robot… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-GR1.tabular1M<n<10M28 likes18k downloads7mo agoHugging Face13netflix /Vera-Layered-Video-Dataset Dataset for Vera: A Layered Diffusion Model for Content-Preserving Video Editing Hongkai Zheng¹²* &nbsp;·&nbsp; Ta-Ying Cheng² &nbsp;·&nbsp; Benjamin Klein² &nbsp;·&nbsp; Yisong Yue¹ &nbsp;·&nbsp; Zhuoning Yuan²† ¹California Institute of Technology &nbsp;&nbsp; ²Netflix, Inc. *Work done during an internship at Netflix &nbsp; †Project Lead TL;DR: A layered diffusion framework for video editing. Vera jointly generates an edit layer, an alpha… See the full description on the dataset page: https://huggingface.co/datasets/netflix/Vera-Layered-Video-Dataset.videotext-to-video10K<n<100K59 likes17k downloads2mo agoHugging Face14nvidia /Cosmos3-DROID DROID: Distributed Robot Interaction Dataset Dataset Summary DROID (Distributed Robot Interaction Dataset) is a large-scale "in-the-wild" robot manipulation dataset containing 76K teleoperated demonstration trajectories — approximately 350 hours of interaction data — collected across 564 unique scenes, 86 tasks, and 52 buildings over the course of 12 months. The data was collected by 50 data collectors at 18 labs across 13 institutions in North America, Asia, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Cosmos3-DROID.video1K<n<10K47 likes17k downloads3mo agoHugging Face15nexar-ai /nexar_collision_predictiongated Nexar Collision Prediction Dataset This dataset is part of the Nexar Dashcam Crash Prediction Challenge on Kaggle. Dataset The Nexar collision prediction dataset comprises videos from Nexar dashcams. Videos have a resolution of 1280x720 at 30 frames per second and typically have about 40 seconds of duration. The dataset contains 1500 videos where half show events where there was a collision or a collision was eminent (positive cases), and the other half shows… See the full description on the dataset page: https://huggingface.co/datasets/nexar-ai/nexar_collision_prediction.tabularvideo-classification1K<n<10K21 likes12k downloads3d agoHugging Face16nyu-visionx /VSI-Bench Dataset arXiv Website Code VSI-Bench VSI-Bench-Debiased v1 [!IMPORTANT] [Aug. 9, 2026] PROVENANCE UPDATE: The existing "Debiased" subset is VSI-Bench-Debiased v1, a designer-in-the-loop manual pilot created with bespoke per-question-type filtering heuristics. It predates and was not generated by the automated Iterative Bias Pruning (IBP) algorithm. We retain v1 for reproducibility and will version any future automated subset separately.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/VSI-Bench.textvisual-question-answering10K<n<100K71 likes8k downloads2mo agoHugging Face17rahul-ai-01 /groot_n1.7_inference_on_diff_data GROOT Inference Analysis Log Evaluation records for a GR00T policy trained on the task "pick octopus and place inside brown basket", run on a Unitree G1 at 20 Hz with the ego_view stereo camera. Six training checkpoints (50, 100, 150, 200, 250, 300 demonstration episodes) were each evaluated on 50 inference episodes. Every episode is recorded here with video, per-tick state/action logs and run metadata. Success rate Checkpoint (training episodes) Success… See the full description on the dataset page: https://huggingface.co/datasets/rahul-ai-01/groot_n1.7_inference_on_diff_data.video0 likes7.3k downloads1mo agoHugging Face18Ahmed-Nasri /llava-video-178k-siglip-tokens-ftov-new LLaVA-Video-178K SigLIP Token Cache (LLaVA-OV fine-tuned vision tower) Derived data (vision-encoder features of video frames), not a redistribution of the source videos. Source: lmms-lab/LLaVA-Video-178K -- its card restricts use to academic research and education, and its annotations come from GPT-4-class models (see the OpenAI usage policy). Complete: 85000 clips. Subset Folders: 0_30_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_academic_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new.video5 likes7k downloads6d agoHugging Face19IPEC-COMMUNITY /libero_spatial_no_noops_1.0.0_lerobottabular10K<n<100K5 likes6.8k downloads11mo agoHugging Face20nvidia /PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card Dataset Description PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.video100K<n<1M27 likes6.2k downloads4mo agoHugging Face21IPEC-COMMUNITY /libero_goal_no_noops_1.0.0_lerobottabular10K<n<100K1 likes6.1k downloads11mo agoHugging Face22IPEC-COMMUNITY /libero_object_no_noops_1.0.0_lerobottabular10K<n<100K1 likes6k downloads11mo agoHugging Face23IPEC-COMMUNITY /libero_10_no_noops_1.0.0_lerobottabular100K<n<1M3 likes5.9k downloads11mo agoHugging Face24nvidia /BridgeData2_LeRobot_v3 BridgeData2 LeRobot v3 Dataset Summary nvidia/bridge_lerobot_v3 is a LeRobotDataset v3.0 conversion of the BridgeDataset / BridgeData V2 robot manipulation dataset. BridgeData V2 is a large-scale real-world robotics dataset collected to support scalable robot learning, including imitation learning, offline reinforcement learning, and open-vocabulary multi-task policies conditioned on goal images or natural-language instructions. This repository packages Bridge… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/BridgeData2_LeRobot_v3.videon<1K12 likes5.5k downloads4mo agoHugging Face25RUC-NLPIR /Omnimodal-Agent-SFT-2K OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.audioquestion-answering1K<n<10K9 likes5.2k downloads7mo agoHugging Face26Ayesha758 /Emotion_new_collected_datasetaudio10K<n<100K0 likes5.1k downloads4mo agoHugging Face27Joocjun /GR1-Tabletop-NextState-1000x24 GR1 Tabletop Merged LeRobot Datasets Merged and subsampled versions of the GR1 tabletop manipulation datasets from the NVIDIA PhysicalAI-Robotics-GR00T-X-Embodiment-Sim collection, formatted in LeRobot v2.0 format. Dataset Variants Variant Demos/Task Tasks Total Episodes Total Frames Approx Size 1000x24/ 1000 24 folders, 186 unique tasks 24,000 6,020,058 ~40 GB 300x24/ 300 24 folders, 186 unique tasks 7,200 1,803,236 ~12 GB 100x24/ 100 24 folders… See the full description on the dataset page: https://huggingface.co/datasets/Joocjun/GR1-Tabletop-NextState-1000x24.tabularrobotics1M<n<10M0 likes4.9k downloads6mo agoHugging Face28Yangyihui /awm-nav-pretrain AWM Navigation Pre-training Corpus (private) — uniform 30 fps Action-free video pre-training data for AWM. Egocentric go-to navigation clips, open-vocabulary route+landmark / object-level instructions (Gemini v3 harvest), across indoor (object-level) and outdoor (route+landmark) domains. All video is normalized to 30 fps (native 60fps down-sampled, native 30fps kept; 24/25fps clips dropped — no clean resample). fps_map.json records each clip's ORIGINAL fps for reference.… See the full description on the dataset page: https://huggingface.co/datasets/Yangyihui/awm-nav-pretrain.videorobotics10K<n<100K3 likes4.9k downloads2mo agoHugging Face29nkkbr /ViCA-322K ViCA-322K: A Dataset for Visuospatial Cognition in Real-World Indoor Videos Quickstart You can load our dataset using the following code: from datasets import load_dataset vica_322k_arkit_base = load_dataset("nkkbr/ViCA-322K", "arkitscenes_base") Replace "arkitscenes_base" with any of the following configuration names depending on your need: ["arkitscenes_base", "arkitscenes_complex", "scannet_base", "scannet_complex", "scannetpp_base", "scannetpp_complex"]… See the full description on the dataset page: https://huggingface.co/datasets/nkkbr/ViCA-322K.videovisual-question-answering3 likes3.9k downloads9mo agoHugging Face30lmms-eval /NExTQAtabular10K<n<100K6 likes3.9k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.