datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios
Dataset Description:
PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes
PhysicalAI SDG-Warehouse
PhysicalAI SDG-Warehouse is a synthetic, fully-annotated video dataset of staged industrial-safety events captured in a simulated warehouse environment. It contains approximately 123k video clips, totaling roughly 412 hours of footage at 1920x1080 resolution and 30 frames per second, organized across four scenarios: a forklift near-miss with a human worker, a warehouse fire with worker evacuation, a forklift collision with a storage shelf, and a routine… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes.PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes
Dataset Description:
The SDG-SynHuman is a large-scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post-training of NVIDIA Cosmos world foundation models and related physical AI research.
Each sample is a temporally coherent 60-120 second video clip rendered at 1080p and 30 fps. Clips contain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes.PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes
PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card
Dataset Description
PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.Synthetic_Visual_Genome2
Synthetic Visual Genome 2 (SVG2)
A large-scale panoptic video scene graph dataset containing object labels, attributes, relationships, and instance-level segmentation masks.
Paper: Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
Website: Synthetic Visual Genome 2
Versions
cleaned
The cleaned version has two sources: PVD (~593K videos) and SA-V (~43K videos).
SAM-3 outputs
We also provide the instance masks and… See the full description on the dataset page: https://huggingface.co/datasets/UWGZQ/Synthetic_Visual_Genome2.robocurate-synth100
synth100 — 100 generated clips for validating Pre-Contact Level Filtering
100 episodes drawn (seed 20260824) from the 952-episode multi-object generation set, packaged so
Stage-5 filtering can be run on them without re-deriving anything. Every input the filter needs
travels with the package, in the space it is consumed in.
Read section 1 before using this. The single most important fact about this data is not in the
file layout, and getting it wrong invalidates any score… See the full description on the dataset page: https://huggingface.co/datasets/glory-hyeok/robocurate-synth100.SynthSite
SynthSite
SynthSite is a curated benchmark of 227 synthetic construction site safety videos (115 unsafe, 112 safe) generated using four text-to-video models: Sora 2 Pro, Veo 3.1, Wan 2.2-14B, and Wan 2.6. Each video was independently labeled by 2–3 human reviewers for the presence of a Worker Under Suspended Load hazard, producing a binary classification: unsafe (True_Positive — worker remains in the suspended-load fall zone) or safe (False_Positive — no worker in fall zone or… See the full description on the dataset page: https://huggingface.co/datasets/govtech/SynthSite.ByteDance_Synthetic_Videos
Dataset Name
CGI synthetic videos generated in paper "Synthetic Video Enhances Physical Fidelity in Video Synthesis" (https://simulation.seaweed.video/)
Dataset Overview
Number of samples: [uploading...]
Annotations: [tags, captions]
License: [apache-2.0]
Citation: @article{zhao2025synthetic,
title={Synthetic Video Enhances Physical Fidelity in Video Synthesis},
author={Zhao, Qi and Ni, Xingyu and Wang, Ziyu and Cheng, Feng and Yang, Ziyan and Jiang, Lu and Wang, Bohan}… See the full description on the dataset page: https://huggingface.co/datasets/kevinzzz8866/ByteDance_Synthetic_Videos.EgoExo-Synthetic
Synchronized Egocentric–Exocentric Dataset for Synthetic Scenario
synthetic-emotions
Synthetic Emotions Dataset
Overview
Synthetic Emotions is a video dataset of AI-generated human emotions created using OpenAI Sora. It features short (5-sec, 480p, 9:16) videos depicting diverse individuals expressing emotions like happiness, sadness, anger, fear, surprise, and more.
This dataset is ideal for emotion recognition, facial expression analysis, affective computing, and AI-human interaction research.
Dataset Details
Total Videos: 100
Video Format:… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/synthetic-emotions.myarm-8-synthetic-cube-to-cup-largeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 874,
"total_frames": 421190,
"total_tasks": 1,
"total_videos": 1748,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:874"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/apockill/myarm-8-synthetic-cube-to-cup-large.SynthForensicsSynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes
Official Repository for the SynthForensics (SF) Benchmark
Abstract
Modern T2V/I2V generators synthesize people increasingly hard to distinguish from authentic footage, while current evaluation suites lag: legacy benchmarks target manipulation-based forgeries, and recent synthetic-video benchmarks prioritize scale over realistic human depiction. We introduce SynthForensics, a… See the full description on the dataset page: https://huggingface.co/datasets/SynthForensics/SynthForensics.SynthGait-19K
SynthGait-19K
SynthGait-19K contains 19,272 synthetic walking videos from
6,427 underlying motion sequences. Each motion is rendered from up to
three camera viewpoints. The dataset supports video-based gait parameter estimation.
Dataset structure
The release uses uncompressed WebDataset TAR shards of approximately 1 GB. The shared
vid_XXXXX base identifies one walking sequence and the suffix identifies its camera
view. Each full stem forms one… See the full description on the dataset page: https://huggingface.co/datasets/SoroushMehraban/SynthGait-19K.SynthForensics_sampleSynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes
Official Repository for the SynthForensics (SF) Benchmark
Note: This is the sample release of SynthForensics, comprising 10 videos per generator with their respective metadata in JSON format selected to broadly represent the diversity and characteristics of the full benchmark. It is intended for dataset preview, model selection, and preliminary evaluation purposes. The complete dataset… See the full description on the dataset page: https://huggingface.co/datasets/SynthForensics/SynthForensics_sample.myarm-7-synthetic-cube-to-cupThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 277,
"total_frames": 131449,
"total_tasks": 1,
"total_videos": 554,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:277"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/apockill/myarm-7-synthetic-cube-to-cup.SyntheticStreetScenes
Synthetic Street Scenes
Mehmet Kerem Turkcan
Columbia University, Center for Smart Streetscapes (CS3)
Synthetic Street Scenes is a collection of 22 video datasets of street situations that are rare, dangerous or impractical to film: street flooding, snow cover, traffic jams, collisions and near misses, storm damage, blocked bike lanes, tampering with roadside sensors, and civic maintenance problems seen from egocentric and fixed viewpoints. It holds 5… See the full description on the dataset page: https://huggingface.co/datasets/mehmetkeremturkcan/SyntheticStreetScenes.egoproactive-synth-annotations
EgoProactive synthetic proactive annotations
Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026 EgoProactive
Dense timestamped proactive walkthroughs generated with the ambient agent
(orchestrator deepseek/deepseek-v4-flash-0731 + vision Qwen3.6-27B),
using the held-out-validated dense policy (setup-phase coverage, repetition-collapse,
fire-at-onset) and a -0.5s onset correction at chunk-binning.
set
clips
median events/clip… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egoproactive-synth-annotations.so100_kitchen_synthetic_dataTOY-GOAT-synthetic-video-256-2000Cube_Stacking_synthsquishy_synthetic_dataset_40robocasa_30_demos_lerobot_5_chosen_tasks_v3_synthetic_left_right_allThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 104,
"total_frames": 26745,
"total_tasks": 45,
"total_videos": 624,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:104"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/robocasa_30_demos_lerobot_5_chosen_tasks_v3_synthetic_left_right_all.SimMotion-Synthetic
SimMotion-Synthetic Benchmark
Synthetic benchmark for evaluating motion representation invariance, introduced in:
"SemanticMoments: Training-Free Motion Similarity via Third Moment Features" (arXiv:2602.09146)
License: For research purposes only.
Dataset Description
250 triplets (750 videos) across 5 categories of visual variations:
Category
Description
Examples
static_object
Variations in surrounding static context
50
dynamic_attribute
Variations in moving… See the full description on the dataset page: https://huggingface.co/datasets/Shuberman/SimMotion-Synthetic.wantrack-synth-toy-720p
WanTrack synth toy — 720p
A 720p/24fps recreation of noctuashap/wantrack_synth_toy:
one shared seed image + 50 motion captions → 50 I2V clips, each varying only the motion.
Used as the overfit set for the bidirectional TrackWan teacher recipe in
FastVideo.
How it was built
Generation — Wan2.1-I2V-14B-720P, conditioned on the single synthetic_seed.png + each
line of captions.txt, 720×1280, 121 frames @ 24fps (gen_synth_i2v_worker.py).
Tracks — CoTracker3 on a… See the full description on the dataset page: https://huggingface.co/datasets/FastVideo/wantrack-synth-toy-720p.robocasa_30_demos_lerobot_5_chosen_tasks_v3_synthetic_left_right_all_correct_inferenceThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 104,
"total_frames": 26745,
"total_tasks": 45,
"total_videos": 624,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:104"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/robocasa_30_demos_lerobot_5_chosen_tasks_v3_synthetic_left_right_all_correct_inference.Syntheticsynthengine-cot-edge-case-v1
SynthEngine CoT Edge Case Dataset v1.0
Premium synthetic Chain-of-Thought reasoning data for autonomous driving, robotics, and embodied AI edge cases.
🔗 Full dataset (1000 records) available on Gumroad
This HuggingFace repo contains a free sample (10 records) under CC BY-NC-SA 4.0.
🎯 Why This Dataset?
In 2025, NVIDIA Alpamayo-R1 proved that Chain-of-Causation reasoning improves autonomous driving planning accuracy by +12% and reduces close encounters by -35%.… See the full description on the dataset page: https://huggingface.co/datasets/NeroSeungSan/synthengine-cot-edge-case-v1.nfa-inspection-synth
Nuclear Fuel Assembly (NFA) Inspection Videos
This repository contains inspection data from three nuclear fuel assemblies (NFAs) of varying ages and deformation levels. Each assembly is inspected on all six hexagonal faces, with corresponding videos and supplementary data files provided for each face.
Dataset Contents.
This dataset was generated by a simulator described in the publication Simulating Nuclear Fuel Inspections: Enhancing Reliability through Synthetic Data.
Each NFA… See the full description on the dataset page: https://huggingface.co/datasets/research-centre-rez/nfa-inspection-synth.pusht-synthetic-v3pusht-synthetic-v5
