datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes
PhysicalAI SDG-Warehouse
PhysicalAI SDG-Warehouse is a synthetic, fully-annotated video dataset of staged industrial-safety events captured in a simulated warehouse environment. It contains approximately 123k video clips, totaling roughly 412 hours of footage at 1920x1080 resolution and 30 frames per second, organized across four scenarios: a forklift near-miss with a human worker, a warehouse fire with worker evacuation, a forklift collision with a storage shelf, and a routine… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes.PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes
Dataset Description:
The SDG-SynHuman is a large-scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post-training of NVIDIA Cosmos world foundation models and related physical AI research.
Each sample is a temporally coherent 60-120 second video clip rendered at 1080p and 30 fps. Clips contain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes.anime-scenes-vidPhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes
PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card
Dataset Description
PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.incantation-elden-ring-scenes
Incantation Elden Ring Combat Captions
Paper | Project page | GitHub
Preview subset. This repository is a public preview and reference subset of the Incantation dataset. It documents the data format, annotation style, and initial training material used by the project. It should not be interpreted as the final dataset, the complete benchmark, or the full data scale used by the paper.
This dataset contains manually collected Elden Ring combat clips paired with structured… See the full description on the dataset page: https://huggingface.co/datasets/zhush/incantation-elden-ring-scenes.game-scenes-posed-rgbd
Origin Lab Game Scenes: Posed RGB-D Flythroughs of Game Worlds
Every frame carries the camera that rendered it and the depth the engine computed for it. Ten game worlds, with the camera released from the player for 60% of the footage: metric depth, world-space normals, 4x4 pose, and per-frame intrinsics on one frame index, plus hundreds of full in-place turns and long stretches in which the world is frozen and only the camera moves. Two trajectories per world and one whole… See the full description on the dataset page: https://huggingface.co/datasets/originlab/game-scenes-posed-rgbd.incantation-elden-ring-scenes
Incantation Elden Ring Combat Captions
Preview subset. This repository is an early public preview and reference subset of the Incantation dataset. It is provided to document the data format, annotation style, and initial training material ahead of the full paper/project release. It should not be interpreted as the final dataset, the complete benchmark, or the full data scale used by the paper.
This dataset contains manually collected Elden Ring combat clips paired with structured… See the full description on the dataset page: https://huggingface.co/datasets/MatrixTeam/incantation-elden-ring-scenes.humancentric-scenes-ai
HumanCentric-Scenes-AI
A multimodal benchmark of 296 AI-generated human-centric scenes across four domains:
CCTV / surveillance imagery (Set 2, 85 images). Midjourney-generated stills that mimic low-resolution security-camera footage — parking lots, building interiors, outdoor public spaces — designed to test whether detection cues survive heavy compression and low-light noise.
Occupation × gender portraits (Set 3, 128 images). A balanced 64-occupation × 2-gender paired design… See the full description on the dataset page: https://huggingface.co/datasets/rjmaftv33/humancentric-scenes-ai.cinepile-t2v-split_scenes_single_shot_uniform
Important Columns for Captioning
Caption_t2v_style: Expressive and long caption generated by Gemini Flash 2.5 for the extracted shot.
Caption_t2v_style_short: Short caption generated by Gemini Flash 2.5 for the extracted shot.
Avg-Aesthetic-Score-Laion-Aesthetics: Average (over frames) aesthetic score of the extracted shot from Laion Aesthetics.
Frame-Aesthetic-Scores-Laion-Aesthetics: Aesthetic scores of each frame of the extracted shot from Laion Aesthetics.… See the full description on the dataset page: https://huggingface.co/datasets/CinematicT2vData/cinepile-t2v-split_scenes_single_shot_uniform.scene-scanning-video
Environment Scanning Dataset
The dataset comprises 10,000 videos capturing humans performing natural human activity by scanning their surroundings in controlled settings. It features individuals interacting with various objects and natural environments, providing rich environmental scans for research applications.
By utilizing this data, researchers can identify areas for improving detection and recognition algorithms, including face detection, face recognition, and object… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/scene-scanning-video.my-movie-scenes
