CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios Dataset Description: PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.video100K<n<1M26 likes55k downloads4mo agoHugging Face02JZSG /synth_dataset0 likes40k downloads11mo agoHugging Face03efficient-deep-research /synthesized_datasettext10K<n<100K0 likes34k downloads11mo agoHugging Face04nvidia /PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card Dataset Description PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding. Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.image100M<n<1B41 likes29k downloads4mo agoHugging Face05nvidia /PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes PhysicalAI SDG-Warehouse PhysicalAI SDG-Warehouse is a synthetic, fully-annotated video dataset of staged industrial-safety events captured in a simulated warehouse environment. It contains approximately 123k video clips, totaling roughly 412 hours of footage at 1920x1080 resolution and 30 frames per second, organized across four scenarios: a forklift near-miss with a human worker, a warehouse fire with worker evacuation, a forklift collision with a storage shelf, and a routine… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes.videovideo-classification100K<n<1M20 likes23k downloads4mo agoHugging Face06nvidia /PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes Dataset Description: The SDG-SynHuman is a large-scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post-training of NVIDIA Cosmos world foundation models and related physical AI research. Each sample is a temporally coherent 60-120 second video clip rendered at 1080p and 30 fps. Clips contain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes.video41 likes21k downloads4mo agoHugging Face07fubel /synthehicle Dataset Card for Synthehicle Synthehicle is a massive CARLA-based synthehic multi-vehicle multi-camera tracking dataset and includes ground truth for 2D detection and tracking, 3D detection and tracking, depth estimation, and semantic, instance and panoptic segmentation. All details can be found in our paper and git repository. 1M<n<10M1 likes20k downloads3y agoHugging Face08princeton-vl /InFlux-Synth InFlux++ Synth InFlux++ Synth is a large-scale synthetic training dataset for the InFlux project, providing per-frame ground truth camera intrinsics and camera pose for videos with dynamic intrinsics. The dataset contains 441,840 annotated frames from 1,841 procedurally generated high-resolution videos. Every video contains 240 frames at a resolution of 1280 × 720. The dataset spans indoor and nature scenes and features changing zoom and focus, dynamic objects, and realistic… See the full description on the dataset page: https://huggingface.co/datasets/princeton-vl/InFlux-Synth.1 likes18k downloads10d agoHugging Face09tonyc54 /Total_Editing_Synthetic_Video_Albedo_Full0 likes18k downloads1y agoHugging Face10research-backup /qa_squadshifts_synthetic_randomTBA 0 likes17k downloads4y agoHugging Face11Arturito1 /OCR-Synthetic-Multilingual-v1 OCR-Synthetic-Multilingual-v1 Overview Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al. This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection. Languages… See the full description on the dataset page: https://huggingface.co/datasets/Arturito1/OCR-Synthetic-Multilingual-v1.object-detection10M<n<100M0 likes15k downloads5mo agoHugging Face12Nishant2414 /OCR-Synthetic-Multilingual-v1 OCR-Synthetic-Multilingual-v1 Overview Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al. This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection. Languages… See the full description on the dataset page: https://huggingface.co/datasets/Nishant2414/OCR-Synthetic-Multilingual-v1.object-detection10M<n<100M0 likes15k downloads5mo agoHugging Face13nvidia /OCR-Synthetic-Multilingual-v1 OCR-Synthetic-Multilingual-v1 Dataset Description Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al. This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OCR-Synthetic-Multilingual-v1.object-detection10M<n<100M53 likes14k downloads5mo agoHugging Face14PleIAs /SYNTH SYNTH Blog announcement SYNTH is the first open generalist synthetic dataset for training small reasoning model end-to-end, jointly released by Pleias and the AI Alliance. SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the Structured Wikipedia dataset from Wikimedia Enterprise. SYNTH differs… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/SYNTH.texttext-generation10M<n<100M277 likes13k downloads5mo agoHugging Face15YuanHo /OCR-Synthetic-Multilingual-v1 OCR-Synthetic-Multilingual-v1 Dataset Description Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al. This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection. This… See the full description on the dataset page: https://huggingface.co/datasets/YuanHo/OCR-Synthetic-Multilingual-v1.object-detection10M<n<100M0 likes11k downloads5mo agoHugging Face16TMaxxx /agent-task-recursive-task-synthesis Apptainer pool for hamishivi/agent-task-recursive-task-synthesis This repository hosts tmax-compatible SIF images and a unified download manifest. Training data and task archives are in hamishivi/agent-task-recursive-task-synthesis. The manifest includes earlier images hosted under hamishivi and new images hosted under TMaxxx; the downloader selects the correct repository and immutable commit for each image. Apptainer images The pool currently contains 29,501 / 29… See the full description on the dataset page: https://huggingface.co/datasets/TMaxxx/agent-task-recursive-task-synthesis.5 likes11k downloads8d agoHugging Face17allenai /olmOCR-synthmix-1025 olmOCR-synthmix-1025 olmOCR-synthmix-1025 is a dataset of 2,186 single PDF pages, that have been synthetically rerendered into HTML by claude-sonnet-4-20250514. In total, across these PDF pages, 30,381 synthetic benchmark cases have been created, following the format of olmOCR-bench. These documents contain no overlap with the original olmOCR-bench documents, and thus can be used as RLVR training data to improve the performance of OCR engines. Directory Structure… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-synthmix-1025.document1K<n<10K3 likes11k downloads11mo agoHugging Face18InternRobotics /SynthVerseViewer is explicitly configured to read the parquet split only. Citation If you find this dataset useful, please cite: @article{zhao2026SythnVerse, title={SynthVerse: A Large-Scale Diverse Synthetic Dataset for Point Tracking}, author={Weiguang Zhao and Haoran Xu and Xingyu Miao and Qin Zhao and Rui Zhang and Kaizhu Huang and Ning Gao and Peizhou Cao and Mingze Sun and Mulin Yu and Tao Lu and Linning Xu and Junting Dong and Jiangmiao Pang}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/SynthVerse.imagerobotics1K<n<10K15 likes9.5k downloads4mo agoHugging Face19hammh0a /SynthCLIPSynthCI-30M This repo contains SynthCI-30M which is the dataset proposed in "SynthCLIP: Are We Ready For a Fully Synthetic CLIP Training?". The dataset contains 30M synthetic text-image pairs covering a wide range of concepts. "We will reach a time where machines will create machines." Abstract We present SynthCLIP, a novel framework for training CLIP models with entirely synthetic text-image pairs, significantly departing from previous methods relying on real… See the full description on the dataset page: https://huggingface.co/datasets/hammh0a/SynthCLIP.12 likes8.8k downloads3y agoHugging Face20Felix92 /OCR-Synthetic-Multilingual-v1 OCR-Synthetic-Multilingual-v1 Dataset Description Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al. This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection.… See the full description on the dataset page: https://huggingface.co/datasets/Felix92/OCR-Synthetic-Multilingual-v1.object-detection10M<n<100M1 likes8.3k downloads5mo agoHugging Face21sophia1ch /zendo-synthetic-data Zendo Synthetic Visual Reasoning Dataset Synthetic Zendo-style scenes with associated rules and per-scene tensor representations. Each scene either follows ("positive", label=1) or violates ("negative", label=0) a rule that is given in natural language and as a Prolog query. Splits split scenes train 56475 test 3344 rules total 3439 Layout images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.imageimage-classification10K<n<100K1 likes7.7k downloads4mo agoHugging Face22pablovela5620 /nerf-synthetic-mirrorimage0 likes7.5k downloads6mo agoHugging Face23nvidia /PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card Dataset Description PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.video100K<n<1M27 likes7.3k downloads4mo agoHugging Face24weikaih /synthetic_data_v5_finegrain_layout_relight_with_our_synthetic_data_coco_l_full_500kimage1K<n<10K0 likes7.3k downloads2y agoHugging Face25vkehfdl1 /banana-vidorev3-synthetic-arms Banana ViDoRe v3 Synthetic Arms Domain-separated ViDoRe v3 synthetic training arms for finance and industrial adaptation. The Hub dataset uses finance and industrial as dataset configs/subsets. Within each config, splits separate vlm_in_batch, vlm_ocr_bm25, banana_fullpipe, and hybrid_vlm_ocr_bm25_banana_fullpipe. Generated at: 2026-06-29T11:49:45.670912+00:00 Total JSONL rows across configs/splits: 151691. Images are stored once per subset under… See the full description on the dataset page: https://huggingface.co/datasets/vkehfdl1/banana-vidorev3-synthetic-arms.imagevisual-document-retrieval10K<n<100K0 likes7k downloads2mo agoHugging Face26vctvct123 /MVS-synthimage1K<n<10K0 likes6.8k downloads6mo agoHugging Face27annahbanannah /synthetic-math-toolcall-deception Synthetic Math Tool-Call Deception 200 paired multi-turn math-assistant trajectories (400 rows) for evaluating deception detectors on mid-trajectory tool-call misreporting. Each trajectory: a system prompt instructs the model to compute via an execute_python tool under a stated tool-call limit, and requires every call to carry a running call_index argument (1 for the first call, 2 for the second, …). The platform enforcing the limit is said to only see the reported call_index… See the full description on the dataset page: https://huggingface.co/datasets/annahbanannah/synthetic-math-toolcall-deception.tabulartext-classificationn<1K0 likes6.5k downloads2mo agoHugging Face28SynthLabsAI /Big-Math-RL-Verifiedgated Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models Big-Math is the largest open-source dataset of high-quality mathematical problems, curated specifically for reinforcement learning (RL) training in language models. With over 250,000 rigorously filtered and verified problems, Big-Math bridges the gap between quality and quantity, establishing a robust foundation for advancing reasoning in LLMs. Request Early Access to Private… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/Big-Math-RL-Verified.textquestion-answering100K<n<1M243 likes5.3k downloads2y agoHugging Face29hf-internal-testing /transformers-synthetic-assets transformers-synthetic-assets Synthetic media fixtures for Transformers tests. These assets are generated from prompts or deterministic code and are not derived from third-party source files. imagen<1K0 likes5.1k downloads16d agoHugging Face30riotu-lab /Synthetic-UAV-Flight-Trajectories UAV Trajectory Dataset Summary This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training. Data Description The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Synthetic-UAV-Flight-Trajectories.tabular100K<n<1M17 likes5.1k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.