CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sophia1ch /zendo-synthetic-data Zendo Synthetic Visual Reasoning Dataset Synthetic Zendo-style scenes with associated rules and per-scene tensor representations. Each scene either follows ("positive", label=1) or violates ("negative", label=0) a rule that is given in natural language and as a Prolog query. Splits split scenes train 56475 test 3344 rules total 3439 Layout images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.imageimage-classification10K<n<100K1 likes4.7k downloads4mo agoHugging Face02openbmb /VisRAG-Ret-Train-Synthetic-data Dataset Description This dataset is the synthetic part of the training set of VisRAG it includes 239,358 Query-Document (Q-D) Pairs from a synthetic dataset made up of pages from web-crawled PDF documents and augmented with VLM-generated (GPT-4o) pseudo-queries. Our training data is organized with a batch size of 128, ensuring that all data within the same batch comes from the same dataset. Name Source Description # Pages Textbooks https://openstax.org/ College-level… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-Synthetic-data.image100K<n<1M20 likes3.4k downloads2y agoHugging Face03alpha-brain /ocr-synthetic-cheque-datatsetimage10K<n<100K1 likes2.4k downloads2y agoHugging Face04physicl /synthetic-bathroom-dataset-for-robotic-perception Synthetic Bathroom Dataset for Robotic Perception Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by the… See the full description on the dataset page: https://huggingface.co/datasets/physicl/synthetic-bathroom-dataset-for-robotic-perception.imagen<1K0 likes2k downloads3mo agoHugging Face05ProGamerGov /synthetic-dataset-1m-dalle3-high-quality-captions Dataset Card for Dalle3 1 Million+ High Quality Captions Alt name: Human Preference Synthetic Dataset Example grids for landscapes, cats, creatures, and fantasy are also available. Description: This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.imagetext-to-image1M<n<10M154 likes2k downloads2y agoHugging Face06physicl /synthetic-living-room-dataset-for-robotic-perception Synthetic Living Room Dataset for Robotic Perception Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/synthetic-living-room-dataset-for-robotic-perception.imagen<1K0 likes1.7k downloads3mo agoHugging Face07julioojalvo /synthetic_kidney_stone_segmentation_dataimage10K<n<100K1 likes1.4k downloads3mo agoHugging Face08myvision /yuanchuan-synthetic-dataset-finaltext100M<n<1B0 likes1.4k downloads4y agoHugging Face09catmint123 /HGGT-synthetic-data HGGT Synthetic Dataset This is the synthetic multi-view hand-object interaction dataset introduced in: HGGT: Robust and Flexible 3D Hand Mesh Reconstruction from Uncalibrated ImagesYumeng Liu, Xiao-Xiao Long, Marc Habermann, Xuanze Yang, Cheng Lin, Yuan Liu, Yuexin Ma, Wenping Wang, Ligang Liu[Paper] · [Project Page] · [Code] The dataset contains diverse photorealistic hand-object interactions rendered with randomized camera viewpoints, providing critical viewpoint diversity… See the full description on the dataset page: https://huggingface.co/datasets/catmint123/HGGT-synthetic-data.imageimage-to-3dn<1K3 likes688 downloads5mo agoHugging Face10cl-syn-data /longhealth-synthetic LongHealth Synthetic Internalization Corpora Synthetic training corpora generated for the Art of Scaling Continual Learning study of knowledge internalization: how well a model absorbs a small document collection into its weights (vs. reading it in-context) as a function of how much synthetic data you train on. Each corpus re-expresses the same source documents — the six study patients of the public LongHealth benchmark (fictional patient records) — through a different… See the full description on the dataset page: https://huggingface.co/datasets/cl-syn-data/longhealth-synthetic.text1M<n<10M0 likes643 downloads2mo agoHugging Face11kk456123 /VisRAG-Ret-Train-Synthetic-data Dataset Description This dataset is the synthetic part of the training set of VisRAG it includes 239,358 Query-Document (Q-D) Pairs from a synthetic dataset made up of pages from web-crawled PDF documents and augmented with VLM-generated (GPT-4o) pseudo-queries. Our training data is organized with a batch size of 128, ensuring that all data within the same batch comes from the same dataset. Name Source Description # Pages Textbooks https://openstax.org/ College-level… See the full description on the dataset page: https://huggingface.co/datasets/kk456123/VisRAG-Ret-Train-Synthetic-data.image100K<n<1M0 likes586 downloads6mo agoHugging Face12deepsense-ai /synthetic-rag-dataset_v1.0textn<1K0 likes584 downloads1y agoHugging Face13nomic-ai /VisRAG-Ret-Train-Synthetic-data-hn-mine-corpusimage100K<n<1M0 likes560 downloads1y agoHugging Face14PrimeIntellect /SYNTHETIC-1-SFT-Data SYNTHETIC-1: Two Million Crowdsourced Reasoning Traces from Deepseek-R1 SYNTHETIC-1 is a reasoning dataset obtained from Deepseek-R1, generated with crowdsourced compute and annotated with diverse verifiers such as LLM judges or symbolic mathematics verifiers. This is the SFT version of the dataset - the raw data and preference dataset can be found in our 🤗 SYNTHETIC-1 Collection. The dataset consists of the following tasks and verifiers that were implemented in our library… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SYNTHETIC-1-SFT-Data.text100K<n<1M34 likes555 downloads2y agoHugging Face15Kylan12 /Synthetic-AI-ML-Dataset Synthetic-AI-ML-Dataset Synthetic Q&A dataset on AI and Machine Learning Dataset Details Metric Value Topic AI and Machine Learning Total Q&A Pairs 14021 Valid Pairs 14021 Provider/Model ollama/gpt-oss:120b Generation Cost Metric Value Prompt Tokens 14,941,957 Completion Tokens 17,159,263 Total Tokens 32,101,220 GPU Energy 12.9628 kWh Sources This dataset was generated from 474 scholarly papers: #… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.textquestion-answering10K<n<100K2 likes473 downloads6mo agoHugging Face16Lego-X /SWE-Lego-Synthetic-Data Dataset Summary Paper | Github | HF Collection SWE-Lego-Synthetic-Data contains 11.5k synthetic github issues (Python language) and their multi-turn agent trajectories. The column named messages is collected using Qwen/Qwen3-Coder-480B-A35B-Instruct with OpenHands (v0.53.0) agent scaffolding, which can be directly used for SFT training. This dataset is part of the work presented in SWE-Lego, a supervised fine-tuning (SFT) recipe designed to achieve state-of-the-art performance… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/SWE-Lego-Synthetic-Data.texttext-generation10K<n<100K5 likes422 downloads9mo agoHugging Face17nomic-ai /VisRAG-Ret-Train-Synthetic-dataimage100K<n<1M0 likes405 downloads1y agoHugging Face18medieval-data /catmus-synthetic-v2image10K<n<100K1 likes369 downloads2y agoHugging Face19VRKomari /100K_tx_synthetic_patient_data Dataset Card for 100K Texas Synthetic Patient Dataset Dataset Details Dataset Description This dataset contains 112,413 synthetic patient records from Texas, generated using Synthea (version 3.2.0). The data is provided in two formats: FHIR R4 bundles (NDJSON) and flattened Parquet tables for analytics. Total dataset includes 261+ million records across 18 healthcare entity types. Curated by: VenkataRaghu Komari Language(s): English License: CC-BY-4.0 DOI:… See the full description on the dataset page: https://huggingface.co/datasets/VRKomari/100K_tx_synthetic_patient_data.tabulartext-classification100M<n<1B1 likes357 downloads8mo agoHugging Face20din0s /synthetic-beir-datatext1M<n<10M0 likes354 downloads3y agoHugging Face21open-r1 /SYNTHETIC-1-SFT-Data-Code_decontaminated Dataset description This dataset is the same as open-r1/SYNTHETIC-1-SFT-Data-Code decontaminated against the benchmark datasets. The decontamination has been run using the script in huggingface/open-r1: python scripts/decontaminate.py \ --dataset "open-r1/SYNTHETIC-1-SFT-Data-Code" \ -c ... Removed 5 samples from 'aime_2025' Removed 50 samples from 'math_500' Removed 13234 samples from 'lcb' Initial size: 62953, Final size: 49664 tabular10K<n<100K3 likes303 downloads2y agoHugging Face22Hani89 /Synthetic-Medical-Speech-Dataset Synthetic Medical Speech Dataset Overview Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.audioautomatic-speech-recognition10K<n<100K4 likes248 downloads11mo agoHugging Face23Stereotypes-in-LLMs /hiring-bias-mitigation-synthetic-data Hiring-bias mitigation — synthetic training data Semi-synthetic data for training LLMs to make hiring decisions that do not depend on a protected attribute (military status, gender, religion), in English and Ukrainian. Real inputs, synthetic labels. CVs and job descriptions are real, anonymised postings from the Djinni Recruitment Dataset (MIT). Decisions and rationales were written by the teacher model Qwen/Qwen3.5-122B-A10B-GPTQ-Int4. Code and results:… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-synthetic-data.tabulartext-generation100K<n<1M0 likes235 downloads2h agoHugging Face24electricsheepafrica /africa-synth-cerebral-palsy-synthetic-dataset-all African Cerebral Palsy Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cerebral-palsy-synthetic-dataset-all.tabulartabular-classification10K<n<100K0 likes231 downloads1mo agoHugging Face25ContextReq /Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58 items after the acceptance gates were strengthened (prompt-instruction leaks, markdown bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26. SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling. Metrics Value genres 26 stories per genre 1.153-1.154K stories total characters 38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.texttext-generation10K<n<100K1 likes210 downloads7d agoHugging Face26BAAI /OpenSeek-Synthetic-Reasoning-Data-Examples OpenSeek-Reasoning-Data OpenSeek [Github|Blog] Recent reseach has demonstrated that the reasoning ability of LLMs originates from the pre-training stage, activated by RL training. Massive raw corpus containing complex human reasoning process, but lack of generalized and effective synthesis method to extract these reasoning process. News 🔥🔥🔥[2025/02/25] We publish some math, code, and general knowledge domain reasoning data synthesized from the current pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples.text1M<n<10M27 likes208 downloads2y agoHugging Face27VoicAndrei /so100_kitchen_synthetic_datatext100K<n<1M0 likes191 downloads9mo agoHugging Face28VladS159 /romanian_speech_dataset_with_15_percent_6_speakers_synthetic_dataaudio10K<n<100K0 likes188 downloads7mo agoHugging Face29magwrap /synthetic_sami_ocr_data Synthetic text images for North, South, Lule and Inari Sámi This dataset contains synthetic line images meant for fitting OCR models for North, South, Lule and Inari Sámi. Clean line images are created using Pillow and they are subsequently distorted using Augraphy [1]. Text sources The text in this dataset comes from Giellatekno's corpus. Specifically, we used the data files of the converted/-directories of [2][3][4][5] (commit hashes… See the full description on the dataset page: https://huggingface.co/datasets/magwrap/synthetic_sami_ocr_data.imageimage-to-text100K<n<1M0 likes186 downloads7mo agoHugging Face30llm-jp /Synthetic-JP-EN-Coding-Dataset Synthetic-JP-EN-Coding-Dataset This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan. The dataset comprises a subset from Aratako/Synthetic-JP-EN-Coding-Dataset-801k. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order. Hirokazu Kiyomaru and Takashi Kodama. text100K<n<1M2 likes183 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.