CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ComplexDataLab /OpenFake Dataset Card for OpenFake Known issues Prompt–image misalignment in the synthetic split (reported November 2025, fix pending) For five of the eighty generators, the prompt field attached to synthetic images does not correspond to the prompt actually used to generate that image. Affected generators: flux-realism sd-3.5 sdxl-realvis-v5 sd-1.5-dreamshaper sd-1.5-epicdream This affects approximately 19.77% of synthetic images. It was first reported in discussion… See the full description on the dataset page: https://huggingface.co/datasets/ComplexDataLab/OpenFake.imageimage-classification1M<n<10M33 likes19k downloads5d agoHugging Face02gatilin /open-vision-banana-snvc-train-full SNVC-50M v5_full — Multi-Task Vision Dataset Description This dataset is a curated subset of the SenseNova Vision Corpus 50M (SNVC-50M), containing 43,509 samples across 6 vision task families and 31 source datasets. Each sample follows a conversational format with interleaved <image> tokens, designed for training vision-language models (VLMs). Coverage: 43,509 / 57,878 (75.2%) of the original sampling plan. 23 datasets at 100%, 8 partial, 12 unrecoverable… See the full description on the dataset page: https://huggingface.co/datasets/gatilin/open-vision-banana-snvc-train-full.imageimage-segmentation10K<n<100K0 likes7.9k downloads2mo agoHugging Face03nebula /OpenSDI_trainThis repository contains the OpenSDI training dataset, presented in the paper OpenSDI: Spotting Diffusion-Generated Images in the Open World. Code: https://github.com/iamwangyabin/OpenSDI imageimage-classification100K<n<1M2 likes3.2k downloads1y agoHugging Face04jaddai /openbrush OpenBrush-75K A curated dataset of 75,313 public domain artworks with rich, structured VLM-generated captions designed for training image generation models, fine-tuning vision-language models, and art analysis research. Dataset Description OpenBrush-75K contains high-quality reproductions of paintings from the Western art canon, spanning from the Renaissance to the early 20th century. Each image is paired with a detailed structured caption generated by a… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush.tabularimage-to-text10K<n<100K3 likes3.1k downloads4mo agoHugging Face05nebula /OpenSDI_test OpenSDI: Spotting Diffusion-Generated Images in the Open World This dataset is designed to address the OpenSDI challenge: spotting diffusion-generated images in realistic, open-world scenarios. It is described in the paper: Project Page: https://iamwangyabin.github.io/OpenSDI/ OpenSDID Dataset Highlights: User Diversity: Simulates a wide range of user intentions and creative styles using diverse text prompts generated by VLMs. Model Innovation: Includes images from multiple… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDI_test.imageimage-classification100K<n<1M1 likes2.5k downloads1y agoHugging Face06nyuuzyou /OpenGameArt-CC0 Dataset Card for OpenGameArt-CC0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons 0 (CC0) license, making them effectively public domain works. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata. Languages The dataset is primarily monolingual:… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC0.audioimage-classification10K<n<100K10 likes1.5k downloads1y agoHugging Face07Rapidata /OpenAI-4o_t2i_human_preference Rapidata OpenAI 4o Preference This T2I dataset contains over 200'000 human responses from over ~45,000 individual annotators, collected in less than half a day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating OpenAI 4o (version from 26.3.2025) across three categories: preference, coherence, and alignment. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/OpenAI-4o_t2i_human_preference.imagetext-to-image10K<n<100K34 likes698 downloads1y agoHugging Face08nyuuzyou /OpenGameArt-OGA-BY-4.0 Dataset Card for OpenGameArt-OGA-BY-4.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the OpenGameArt Attribution 4.0 (OGA-BY-4.0) license. The dataset includes various types of game assets such as 2D art, music, sound effects, and associated metadata. Languages The dataset is primarily monolingual: English (en): All asset descriptions and metadata are in English Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-OGA-BY-4.0.audioimage-classificationn<1K2 likes687 downloads1y agoHugging Face09opendiffusionai /pexels-photos-janpf Images: There are approximately 130K images, borrowed from pexels.com. Thanks to those folks for curating a wonderful resource. There are millions more images on pexels. These particular ones were selected by the list of urls at https://github.com/janpf/self-supervised-multi-task-aesthetic-pretraining/blob/main/dataset/urls.txt . The filenames are based on the md5 hash of each image. Download From here or from pexels.com: You choose For those people who like… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf.text-to-image100K<n<1M45 likes654 downloads8mo agoHugging Face10nyuuzyou /openclipart Dataset Card for OpenClipart.org SVG Images Dataset Summary This dataset contains 178,604 public domain SVG vector clipart images collected from OpenClipart.org. OpenClipart.org is a community-driven platform where artists share vector clip art explicitly released into the public domain (CC0). The dataset includes the SVG content along with comprehensive metadata such as titles, descriptions, artist names, creation dates, tags, and image URLs. The SVG files in this… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/openclipart.textimage-classification100K<n<1M9 likes642 downloads1y agoHugging Face11Rapidata /OpenGVLab_Lumina_t2i_human_preference Rapidata Lumina Preference This T2I dataset contains over 400k human responses from over 86k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating Lumina across three categories: preference, coherence, and alignment. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/OpenGVLab_Lumina_t2i_human_preference.imagetext-to-image10K<n<100K13 likes604 downloads2y agoHugging Face12MichalMlodawski /closed-open-eyes 👀 Open and Closed Eyes Dataset Welcome to the Open and Closed Eyes Dataset! This dataset is designed to help researchers and developers in computer vision and machine learning tasks, particularly in recognizing and distinguishing between open and closed eyes in various contexts. Below, you'll find a detailed description of the dataset structure, categories, and how to interpret the data. 🌟 📁 Dataset Structure The dataset is stored in Parquet files, ensuring efficient… See the full description on the dataset page: https://huggingface.co/datasets/MichalMlodawski/closed-open-eyes.imageimage-classification100K<n<1M4 likes522 downloads2y agoHugging Face13nyuuzyou /OpenGameArt-CC-BY-SA-3.0 Dataset Card for OpenGameArt-CC-BY-SA-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution-ShareAlike 3.0 Unported (CC-BY-SA-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata. Languages The dataset is primarily… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-SA-3.0.audioimage-classification1K<n<10K1 likes481 downloads1y agoHugging Face14Trever896 /openbrush-75k OpenBrush-75K A curated dataset of 75,313 public domain artworks with rich, structured VLM-generated captions designed for training image generation models, fine-tuning vision-language models, and art analysis research. Dataset Description OpenBrush-75K contains high-quality reproductions of paintings from the Western art canon, spanning from the Renaissance to the early 20th century. Each image is paired with a detailed structured caption generated by a vision-language… See the full description on the dataset page: https://huggingface.co/datasets/Trever896/openbrush-75k.tabularimage-to-text10K<n<100K1 likes438 downloads7mo agoHugging Face15OpenSafetyLab /t2i_safety_dataset T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation This dataset, T2ISafety, is a comprehensive safety benchmark designed to evaluate Text-to-Image (T2I) models across three key domains: toxicity, fairness, and bias. It provides a detailed hierarchy of 12 tasks and 44 categories, built from meticulously collected 70K prompts. Based on this taxonomy and prompt set, T2ISafety includes 68K manually annotated images, serving as a robust resource for… See the full description on the dataset page: https://huggingface.co/datasets/OpenSafetyLab/t2i_safety_dataset.text-to-image10K<n<100K3 likes429 downloads1y agoHugging Face16nyuuzyou /OpenGameArt-OGA-BY-3.0 Dataset Card for OpenGameArt-OGA-BY-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the OpenGameArt Attribution (OGA-BY-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and associated metadata. Languages The dataset is primarily monolingual: English (en): All asset descriptions and metadata are in… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-OGA-BY-3.0.audioimage-classificationn<1K1 likes423 downloads1y agoHugging Face17jaddai /openbrush-landscapes OpenBrush Landscapes Every landscape painting from OpenBrush-75K — across all artists, movements, and centuries. Largest single-genre subset. Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 12,612 you actually want. Why this subset Every landscape across the parent dataset's full range — Romantic wildernesses, Impressionist… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-landscapes.tabularimage-to-text10K<n<100K1 likes414 downloads4mo agoHugging Face18nebula /OpenSDIDplus OpenSDID+ OpenSDID+ is an extended release of the OpenSDI dataset. It complements the original SD1.5 training split with large-scale images from the remaining OpenSDI generators: SD2, SD3, SDXL, and FLUX. The dataset follows the OpenSDI challenge introduced in "OpenSDI: Spotting Diffusion-Generated Images in the Open World". OpenSDI studies detection and localization of diffusion-generated images under realistic open-world settings, including diverse user intentions, evolving… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDIDplus.imageimage-classification100K<n<1M0 likes404 downloads3mo agoHugging Face19jaddai /openbrush-impressionism OpenBrush Impressionism Every Impressionist work from OpenBrush-75K — the largest movement subset. Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 12,798 you actually want. Why this subset Broad-coverage subset for training on the Impressionist visual language: broken brushwork, light-on-color theory, plein-air staging, atmospheric… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-impressionism.tabularimage-to-text10K<n<100K1 likes361 downloads4mo agoHugging Face20nyuuzyou /OpenGameArt-CC-BY-3.0 Dataset Card for OpenGameArt-CC-BY-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution 3.0 (CC-BY-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata. Languages The dataset is primarily monolingual: English (en): All… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-3.0.textimage-classification1K<n<10K1 likes349 downloads1y agoHugging Face21jaddai /openbrush-baroque OpenBrush Baroque Baroque works from OpenBrush-75K (~1600–1750) — chiaroscuro, religious painting, dramatic light. Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 4,240 you actually want. Why this subset The canonical Baroque visual language — Caravaggio, Rembrandt, Vermeer, Velázquez, Rubens. Useful for models learning dramatic… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-baroque.tabularimage-to-text1K<n<10K1 likes308 downloads4mo agoHugging Face22SkyWhal3 /SNAP25_ESM2_OpenFold3_Structural_Analysis 🧬 SNAP25 OpenFold3 Structural Analysis Dataset Comprehensive structural predictions and therapeutic discovery data for 677 SNAP25 missense variants 🎯 Overview This dataset provides the first comprehensive structural analysis of SNAP25 (Synaptosomal-Associated Protein 25 kDa) missense variants, generated to support therapeutic discovery for SNAP25-related developmental and epileptic encephalopathy (DEE-SNAP25). SNAP25 is a critical component of the neuronal… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/SNAP25_ESM2_OpenFold3_Structural_Analysis.imagetabular-classification1K<n<10K0 likes287 downloads8mo agoHugging Face23opendiffusionai /cc12m-cleaned CC12m-cleaned This dataset builds on two others: The Conceptual Captions 12million dataset, which lead to the LLaVa captioned subset done by CaptionEmporium (The latter is the same set, but swaps out the (Conceptual Captions 12million) often-useless alt-text captioning for decent ones_ I have then used the llava captions as a base, and used the detailed descrptions to filter out images with things like watermarks, artist signatures, etc. I have also manually thrown out all… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned.imagetext-to-image1M<n<10M13 likes278 downloads2y agoHugging Face24LAYEK-143 /Open-Pixel-1T 🌌 Open-Pixel-1T (Visual Atlas) A Large-Scale, High-Entropy Synthetic Image Dataset for Foundational Pre-Training 📑 Dataset Summary Open-Pixel-1T is a monumental open-source initiative designed to create a "Visual Atlas" of stochastic imagery. Unlike traditional datasets scraped from social media which contain inherent human bias, Open-Pixel-1T is constructed using high-entropy random seeds to generate unique, diverse visual signals. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LAYEK-143/Open-Pixel-1T.imagetext-to-image10M<n<100M8 likes251 downloads6mo agoHugging Face25nyuuzyou /OpenGameArt-Mixed-Licenses Dataset Card for OpenGameArt-Mixed-Licenses Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are available under multiple licenses simultaneously. This dataset includes assets where creators have made their work available under two or more license options. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata, all… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-Mixed-Licenses.audioimage-classification1K<n<10K0 likes226 downloads1y agoHugging Face26UniDataPro /open-palm-hand-images Palm dataset - 500,000 images Dataset containing 500,000 human palm images. Designed for hand detection, palm recognition, and gesture analysis, this palm dataset provides diverse training data with metadata on age, gender, and ethnicity for accurate computer vision model training. By leveraging this dataset, researchers and developers can advance computer vision models for highly accurate hand detection, palm recognition, and gesture analysis. - Get the data… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/open-palm-hand-images.imageimage-classificationn<1K0 likes219 downloads1mo agoHugging Face27OpenMed /multicare-images MultiCaRe: Open-Source Clinical Case Dataset MultiCaRe is an open-source, multimodal clinical case dataset built from the PubMed Central Open Access (OA) Case Report articles. It aggregates de-identified, open-access case narratives, figure images, captions, and rich article metadata across diverse specialties (radiology, pathology, surgery, ophthalmology, etc.). The data is normalized so images, cases, and articles can be joined via stable IDs. Source and process: OA case reports… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/multicare-images.imageimage-classification100K<n<1M6 likes205 downloads1y agoHugging Face28ductai199x /open-set-synth-img-attributionThis dataset contains synthetic images generated by different architectures trained on different datasets. There are two types of labels: architecture and generator. The architecture labels are the architectures of the generators used to synthesize the images. The generator labels are a tuple of (architecture, training_dataset) used to synthesize the images. The dataset is split into train, validation, and test sets. For more information, see the [BMVC 2023 paper](https://papers.bmvc2023.org/0659.pdf).image-classification0 likes181 downloads2y agoHugging Face29VasilyLoginov /closed-open-eyes 👀 Open and Closed Eyes Dataset Welcome to the Open and Closed Eyes Dataset! This dataset is designed to help researchers and developers in computer vision and machine learning tasks, particularly in recognizing and distinguishing between open and closed eyes in various contexts. Below, you'll find a detailed description of the dataset structure, categories, and how to interpret the data. 🌟 📁 Dataset Structure The dataset is stored in Parquet files, ensuring efficient… See the full description on the dataset page: https://huggingface.co/datasets/VasilyLoginov/closed-open-eyes.imageimage-classification100K<n<1M1 likes169 downloads9mo agoHugging Face30Aignostics /OpenTMEgated OpenTME: Open-Access Tumor Microenvironment Profiles from TCGA OpenTME is an open-access project by Aignostics for academic researchers. It provides comprehensive spatial outputs for whole slide images (WSIs) of H&E-stained, formalin-fixed, paraffin-embedded slides from The Cancer Genome Atlas (TCGA). OpenTME is powered by Atlas H&E-TME – a computational pathology application developed by Aignostics. Atlas H&E-TME Atlas H&E-TME is a foundation model-based… See the full description on the dataset page: https://huggingface.co/datasets/Aignostics/OpenTME.imageimage-classification10K<n<100K41 likes169 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.