CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sophia1ch /zendo-synthetic-data Zendo Synthetic Visual Reasoning Dataset Synthetic Zendo-style scenes with associated rules and per-scene tensor representations. Each scene either follows ("positive", label=1) or violates ("negative", label=0) a rule that is given in natural language and as a Prolog query. Splits split scenes train 56475 test 3344 rules total 3439 Layout images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.imageimage-classification10K<n<100K1 likes7.7k downloads4mo agoHugging Face02ProGamerGov /synthetic-dataset-1m-dalle3-high-quality-captions Dataset Card for Dalle3 1 Million+ High Quality Captions Alt name: Human Preference Synthetic Dataset Example grids for landscapes, cats, creatures, and fantasy are also available. Description: This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.imagetext-to-image1M<n<10M154 likes1.9k downloads2y agoHugging Face03Thermostatic /frontier-synthetic-images-2026 Frontier Synthetic Images — Deduplicated Research Corpus This is a training-only corpus of 40,290 exact-deduplicated AI-generated images from recent and frontier generators. It normalizes four provenance-pinned sources into one row-per-image schema for image-forensics research. It is not an evaluation benchmark and should not be used to report detector accuracy after training on it. Sources and licensing Qwen/Qwen-Image-Bench at… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/frontier-synthetic-images-2026.imageimage-classification10K<n<100K0 likes370 downloads1mo agoHugging Face04nadizik /synthetic-human-expressions-poses-3d 3D Synthetic Human Poses and FACS Expressions Dataset This is a high-fidelity synthetic dataset consisting of 10,075 pairs of 3D human character renders and detailed natural language annotations. Dataset Structure & Generation To ensure consistency, the dataset is generated using a single base 3D human model. The diversity of the dataset is achieved through a wide range of body poses, facial expressions, and camera angles: Character: 1 base human model. Camera… See the full description on the dataset page: https://huggingface.co/datasets/nadizik/synthetic-human-expressions-poses-3d.imagetext-to-image10K<n<100K0 likes246 downloads3mo agoHugging Face05tehnik-tehnolog /bilingual-ocr-ru-en-synthetic Bilingual OCR RU-EN Synthetic Dataset This synthetic dataset is designed for bilingual text recognition (OCR) and script classification tasks (cyrillic / latin) at the word and short-line level. Why are numbers, mathematical symbols, and the Greek alphabet included in the generation? When creating synthetic OCR datasets, including an expanded set of characters (digits, mathematical signs, and Greek letters) is a deliberate step aimed at two main goals:… See the full description on the dataset page: https://huggingface.co/datasets/tehnik-tehnolog/bilingual-ocr-ru-en-synthetic.imageimage-classification100K<n<1M1 likes236 downloads3d agoHugging Face06raman07 /SynthCheX-75K-v2 SynthCheX-75K SynthCheX-75K is released as a part of the CheXGenBench paper. It is a synthetic dataset generated using Sana (0.6B) [1] fine-tuned on chest radiographs. Sana (0.6B) establishes the SoTA performance on the CheXGenBench benchmark. The dataset contains 75,649 high-quality image-text samples along with the pathological annotations. Filtration Process for SynthCheX-75K Generative models can lead to both high and low-fidelity generations on different subsets… See the full description on the dataset page: https://huggingface.co/datasets/raman07/SynthCheX-75K-v2.texttext-to-image10K<n<100K1 likes233 downloads1y agoHugging Face07AbstractPhil /synthetic-characters Synthetic Characters Dataset A synthetic image dataset generated with Flux Schnell featuring structured character prompts designed for training character generation, fashion understanding, and portrait synthesis models. Recommended Filters: Age People Count Hair color Camera angle Anime/Realistic Nudity/Clothed I'll likely recaption everything with a list of classifications attached to them for easy filtering later. Dataset Description This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-characters.imagetext-to-image100K<n<1M0 likes217 downloads4mo agoHugging Face08AbstractPhil /synthetic-object-relations Synthetic Object Relations Dataset A synthetic image dataset generated with Flux Schnell featuring clean object-relation prompts designed for training spatial reasoning in vision and diffusion models. Dataset Description This dataset contains images generated from structured prompts describing spatial relationships between objects. Unlike typical caption datasets that use free-form text, our prompts follow consistent patterns that explicitly encode: Object identities… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-object-relations.imagetext-to-image100K<n<1M2 likes196 downloads8mo agoHugging Face09DigiGreen /Crop-Disease-Image-Eval-Synthetic Crop, Category, Disease and Pest Test Set 11,057 smallholder-farmer photographs sent to FarmerChat from Ethiopia, India, Kenya and Nigeria, each labelled with the crop, whether the problem is a disease or a pest, and which one. This is the held-out test split of a four-head classification benchmark, restricted to the rows whose labels came from an independent model council rather than from the production vendor. Why 11,057 and not 16,275 The full held-out split is… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Crop-Disease-Image-Eval-Synthetic.textimage-classification10K<n<100K0 likes104 downloads5d agoHugging Face10electricsheepafrica /africa-synth-aid-flows-brain-tumor-mri-colorized-ehr-all Brain Tumor (MRI) Detection Colourized with EHR | Africa (Electric Sheep Africa metadata inventory) Size category: n<1K - Formats: not declared - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-brain-tumor-mri-colorized-ehr-all.imageimage-classification1K<n<10K0 likes95 downloads2mo agoHugging Face11Zhincore /synthetic-human-portrait-attributes AI-generated potraits of people It's not perfect, but could be useful for training for example hairstyle and color classification! Includes 6 categories of labels: 6 ages: 'adult', 'elderly', 'mature', 'teenage', 'young', 'young_adult' 2 sexes: 'man', 'woman' 4 hair lengths: 'long', 'short', 'buzzcut', 'bald' 4 hair shapes: 'curly', 'straight', 'wavy', 'None' (for bald) 9 hair colors: 'black', 'blonde', 'platinum_blonde', 'brunette', 'ginger', 'multicolored', 'None' (for bald) 5… See the full description on the dataset page: https://huggingface.co/datasets/Zhincore/synthetic-human-portrait-attributes.imageimage-classification1K<n<10K2 likes87 downloads1y agoHugging Face12elischwartz /synthetic-watch-faces-dataset Synthetic Watch Faces Dataset A synthetic dataset of analog watch faces displaying various times for training vision models in time recognition tasks. Dataset Description This dataset consists of randomly generated analog watch faces showing different times. Each image contains a watch with hour and minute hands positioned to display a specific time. The dataset is designed to help train and evaluate computer vision models and Vision-Language Models (VLMs) for time… See the full description on the dataset page: https://huggingface.co/datasets/elischwartz/synthetic-watch-faces-dataset.imagevisual-question-answering10K<n<100K0 likes74 downloads1y agoHugging Face13AbdullahKhanSherwani /dubai-taxi-advertising-synthetic Dubai Taxi Advertising Compliance (Synthetic) 63 synthetic photorealistic images of Dubai RTA taxis carrying advertising, built to evaluate whether vision-language models can judge out-of-home (OOH) advertising compliance rules from a single photograph. Each image is generated to be an unambiguous pass or fail against one specific rule from a Dubai taxi advertising technical checklist. The dataset is an evaluation set — it is small, adversarially balanced, and deliberately… See the full description on the dataset page: https://huggingface.co/datasets/AbdullahKhanSherwani/dubai-taxi-advertising-synthetic.imageimage-classificationn<1K0 likes74 downloads2mo agoHugging Face14Nininkkka /Synth-Text-Eng-512x128 Synthetic Text Images (English) A synthetic dataset of rendered text images with rich per-sample annotations: the text itself, its rendering attributes, background description, applied post-processing, and a natural-language caption. Each image is generated by compositing English text over a procedurally generated background with random font, color, position, rotation, blur, brightness and noise. All samples are accompanied by a structured metadata.csv and a ready-to-use… See the full description on the dataset page: https://huggingface.co/datasets/Nininkkka/Synth-Text-Eng-512x128.imageimage-to-text10K<n<100K1 likes74 downloads3d agoHugging Face15mr3vial /paleo-hebrew-seals-synthetic PaleoHebrew-Seals Synthetic Corpus This repository hosts the synthetic corpus part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions. Why this dataset is needed Annotated real Paleo-Hebrew seal photographs are scarce. The synthetic corpus is designed to provide large-scale supervision for training and augmentation while preserving explicit structure at the character level. Overview The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-synthetic.imageobject-detection100K<n<1M0 likes44 downloads4mo agoHugging Face16BrundageLab /synthetic_wildlife_health Synthetic Wildlife Health: Camera Trap Imagery for Alopecia and Body Condition Screening Dataset Summary This dataset contains 553 synthetic camera trap images depicting alopecia (hair loss consistent with mange) and body condition deterioration in North American wildlife, along with paired visual question-answering annotations for health assessment tasks. All images are AI-generated edits of real camera trap photographs sourced from iWildCam 2022. The generative pipeline… See the full description on the dataset page: https://huggingface.co/datasets/BrundageLab/synthetic_wildlife_health.imageimage-classificationn<1K0 likes39 downloads6mo agoHugging Face17lingcarzy /synthetic-dataset-1m-dalle3-high-quality-captions Dataset Card for Dalle3 1 Million+ High Quality Captions Alt name: Human Preference Synthetic Dataset Example grids for landscapes, cats, creatures, and fantasy are also available. Description: This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/lingcarzy/synthetic-dataset-1m-dalle3-high-quality-captions.imagetext-to-image1M<n<10M0 likes39 downloads6mo agoHugging Face18Mattysmittttt /camonet-synthetic CamoNet Synthetic A procedurally-generated military camouflage pattern dataset — 40 historical and contemporary patterns × 200 samples each = 8,000 256×256 RGB images, each tagged with origin, era, and visual family. Sister project to the CamoNet model. Where that model is trained on real photographs scraped from the web, this dataset is fully synthetic — every image is generated from a small Python recipe per pattern family, so the data is reproducible from a seed and freely… See the full description on the dataset page: https://huggingface.co/datasets/Mattysmittttt/camonet-synthetic.imageimage-classification1K<n<10K0 likes38 downloads5mo agoHugging Face19saad02 /brain-tumor-single-slice-MRI-scan-with-synthetic-ehr-africa Dataset Card: Africa Brain Tumor Scans with Synthetic EHR (Bundled Parquet) This dataset bundles single-slice brain MRI scans and richly structured, synthetic EHR data into a single Parquet file suitable for multimodal ML research. Each row contains an image struct (bytes + path), a source label column, and an EHR payload with both a full JSON record and convenient summary columns. The synthetic EHRs are Africa-focused: they encode country, urban/rural, facility level, insurance… See the full description on the dataset page: https://huggingface.co/datasets/saad02/brain-tumor-single-slice-MRI-scan-with-synthetic-ehr-africa.imagetext-classification1K<n<10K0 likes33 downloads10mo agoHugging Face20DatarrX /myanmar-synthetic-syllable-glyphs 🇲🇲 Myanmar Synthetic Syllable Glyphs (MSSG) The Myanmar Synthetic Syllable Glyphs (MSSG) is a massive-scale, high-fidelity synthetic image dataset containing 14,295,552 heavily augmented glyph images (128x64 pixels, grayscale) representing the structural combinatorial matrix of the Burmese script. Developed and engineered by Khant Sint Heinn (Kalix Louis), this core foundational dataset is officially published and maintained under DatarrX (Myanmar Open Source Organization… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-synthetic-syllable-glyphs.imageimage-classification10M<n<100M6 likes27 downloads4mo agoHugging Face21amirali1985 /synthetic-shapes-3x6x7 Synthetic Shapes 3×6×7 A fully deterministic synthetic dataset of simple geometric shapes rendered as SVG images, with precomputed CLIP (ViT-B-32) embeddings for both text and images. Purpose This dataset is designed for controlled experiments in representation alignment and steering vector evaluation. Because images are generated deterministically from a known combinatorial space, it provides a clean testbed where ground-truth structure is fully known.… See the full description on the dataset page: https://huggingface.co/datasets/amirali1985/synthetic-shapes-3x6x7.imageimage-classification10K<n<100K1 likes24 downloads6mo agoHugging Face22marayagomez /synthetic-digital-signature-blocks Synthetic Digital Signature Blocks Synthetically generated images of digital signature appearance blocks — the visual stamp that PDF readers draw on a document when it is signed electronically. Every image is procedurally generated with PIL; no real document, no real person, and no real product branding is included. The dataset was built to train signature-presence detectors on scanned administrative forms, where handwritten signatures are only one of several valid visual… See the full description on the dataset page: https://huggingface.co/datasets/marayagomez/synthetic-digital-signature-blocks.imageimage-classification10K<n<100K0 likes16 downloads2mo agoHugging Face23arrmlet /car-ukraine-synth Ukraine Synthetic Vehicle Dataset — Toyota Corolla × BMW 3 Series Фотореалістичні синтетичні зображення седанів Toyota Corolla та BMW 3 Series в українських урбаністичних сценах, згенеровані через OpenAI gpt-image-2. Кожне зображення семплить з сітки 12 українських міст × 7 погодних умов × 5 часів доби × 6 типів камер (dashcam, CCTV, drone, smartphone, action cam, wall-mounted security). Зображень: 150 Джерело: OpenAI gpt-image-2 (reference-conditioned edits endpoint)… See the full description on the dataset page: https://huggingface.co/datasets/arrmlet/car-ukraine-synth.imageimage-classificationn<1K0 likes9 downloads5mo agoHugging Face24johnlockejrr /heb_synth_pangolinegated Dataset Card for Hebrew Synthetic Pangoline Dataset INFO: I'm not giving access to users with 0 models/0 datasets/0 activity - sharing is both ways Dataset Summary The Hebrew Synthetic Pangoline Dataset is a comprehensive collection of synthetic Hebrew document images generated using a custom implementation of Pangoline, a text-to-image synthesis tool. The dataset contains high-quality synthetic Hebrew text rendered as images, along with corresponding ground truth… See the full description on the dataset page: https://huggingface.co/datasets/johnlockejrr/heb_synth_pangoline.imagetext-to-image10K<n<100K1 likes8 downloads3mo agoHugging Face25SinKove /synthetic_brain_mrigatedThis dataset was obtained as part of the Generative Modelling project from the Artificial Medical Intelligence Group - AMIGO (https://amigos.ai/). It consists on of 1,000 synthetic T1w images sampled from generative models trained on data originally from the UK Biobank dataset (https://www.ukbiobank.ac.uk/).tabularimage-classificationn<1K7 likes7 downloads3y agoHugging Face26danb21 /synthetic-face-sdxl-instantid-benchgated Synthetic Face Detection Benchmark — SDXL+InstantID Version: v1.0.0 · Build date: 2026-05-16 · Rows: 26492 Evaluation benchmark for synthetic-face detection under platform-realistic conditions. Sampled to satisfy the ISO/IEC 19795 floor of 300 samples per demographic subgroup across a 6×2 (skin tone × gender) cell grid. Not training data; not licensed for commercial use. See release.json for build provenance, manifest.csv for per-row license attestation, LICENSES.csv for the… See the full description on the dataset page: https://huggingface.co/datasets/danb21/synthetic-face-sdxl-instantid-bench.imageimage-classification10K<n<100K2 likes4 downloads3mo agoHugging Face27johnlockejrr /yid_synth_pangolinegatedINFO: I'm not giving access to users with 0 models/0 datasets/0 activity - sharing is both ways Dataset Summary The Yiddish Synthetic Pangoline Dataset is a comprehensive collection of synthetic Yiddish document images generated using a custom implementation of Pangoline, a text-to-image synthesis tool. The dataset contains high-quality synthetic Yiddish text rendered as images, along with corresponding ground truth text and ALTO-XML layout annotations. This dataset is designed… See the full description on the dataset page: https://huggingface.co/datasets/johnlockejrr/yid_synth_pangoline.imagetext-to-image10K<n<100K3 likes3 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.