datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenFake
Dataset Card for OpenFake
Known issues
Prompt–image misalignment in the synthetic split (reported November 2025, fix pending)
For five of the eighty generators, the prompt field attached to synthetic
images does not correspond to the prompt actually used to generate that image.
Affected generators:
flux-realism
sd-3.5
sdxl-realvis-v5
sd-1.5-dreamshaper
sd-1.5-epicdream
This affects approximately 19.77% of synthetic images. It was first reported in
discussion… See the full description on the dataset page: https://huggingface.co/datasets/ComplexDataLab/OpenFake.open-vision-banana-snvc-train-full
SNVC-50M v5_full — Multi-Task Vision Dataset
Description
This dataset is a curated subset of the SenseNova Vision Corpus 50M (SNVC-50M), containing 43,509 samples across 6 vision task families and 31 source datasets. Each sample follows a conversational format with interleaved <image> tokens, designed for training vision-language models (VLMs).
Coverage: 43,509 / 57,878 (75.2%) of the original sampling plan. 23 datasets at 100%, 8 partial, 12 unrecoverable… See the full description on the dataset page: https://huggingface.co/datasets/gatilin/open-vision-banana-snvc-train-full.OpenSDI_trainThis repository contains the OpenSDI training dataset, presented in the paper OpenSDI: Spotting Diffusion-Generated Images in the Open World.
Code: https://github.com/iamwangyabin/OpenSDI
OpenSDI_test
OpenSDI: Spotting Diffusion-Generated Images in the Open World
This dataset is designed to address the OpenSDI challenge: spotting diffusion-generated images in realistic, open-world scenarios. It is described in the paper:
Project Page: https://iamwangyabin.github.io/OpenSDI/
OpenSDID Dataset Highlights:
User Diversity: Simulates a wide range of user intentions and creative styles using diverse text prompts generated by VLMs.
Model Innovation: Includes images from multiple… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDI_test.OpenGameArt-CC0
Dataset Card for OpenGameArt-CC0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons 0 (CC0) license, making them effectively public domain works. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC0.OpenGameArt-OGA-BY-4.0
Dataset Card for OpenGameArt-OGA-BY-4.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the OpenGameArt Attribution 4.0 (OGA-BY-4.0) license. The dataset includes various types of game assets such as 2D art, music, sound effects, and associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions and metadata are in English
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-OGA-BY-4.0.OpenAI-4o_t2i_human_preference
Rapidata OpenAI 4o Preference
This T2I dataset contains over 200'000 human responses from over ~45,000 individual annotators, collected in less than half a day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating OpenAI 4o (version from 26.3.2025) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/OpenAI-4o_t2i_human_preference.closed-open-eyes
👀 Open and Closed Eyes Dataset
Welcome to the Open and Closed Eyes Dataset! This dataset is designed to help researchers and developers in computer vision and machine learning tasks, particularly in recognizing and distinguishing between open and closed eyes in various contexts. Below, you'll find a detailed description of the dataset structure, categories, and how to interpret the data. 🌟
📁 Dataset Structure
The dataset is stored in Parquet files, ensuring efficient… See the full description on the dataset page: https://huggingface.co/datasets/MichalMlodawski/closed-open-eyes.openclipart
Dataset Card for OpenClipart.org SVG Images
Dataset Summary
This dataset contains 178,604 public domain SVG vector clipart images collected from OpenClipart.org. OpenClipart.org is a community-driven platform where artists share vector clip art explicitly released into the public domain (CC0). The dataset includes the SVG content along with comprehensive metadata such as titles, descriptions, artist names, creation dates, tags, and image URLs. The SVG files in this… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/openclipart.OpenGameArt-CC-BY-SA-3.0
Dataset Card for OpenGameArt-CC-BY-SA-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution-ShareAlike 3.0 Unported (CC-BY-SA-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-SA-3.0.OpenGVLab_Lumina_t2i_human_preference
Rapidata Lumina Preference
This T2I dataset contains over 400k human responses from over 86k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Lumina across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/OpenGVLab_Lumina_t2i_human_preference.OpenGameArt-OGA-BY-3.0
Dataset Card for OpenGameArt-OGA-BY-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the OpenGameArt Attribution (OGA-BY-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions and metadata are in… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-OGA-BY-3.0.OpenGameArt-CC-BY-3.0
Dataset Card for OpenGameArt-CC-BY-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution 3.0 (CC-BY-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-3.0.OpenSDIDplus
OpenSDID+
OpenSDID+ is an extended release of the OpenSDI dataset. It complements the original SD1.5 training split with large-scale images from the remaining OpenSDI generators: SD2, SD3, SDXL, and FLUX.
The dataset follows the OpenSDI challenge introduced in "OpenSDI: Spotting Diffusion-Generated Images in the Open World". OpenSDI studies detection and localization of diffusion-generated images under realistic open-world settings, including diverse user intentions, evolving… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDIDplus.OpenGameArt-Mixed-Licenses
Dataset Card for OpenGameArt-Mixed-Licenses
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are available under multiple licenses simultaneously. This dataset includes assets where creators have made their work available under two or more license options. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata, all… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-Mixed-Licenses.Open-Pixel-1T
🌌 Open-Pixel-1T (Visual Atlas)
A Large-Scale, High-Entropy Synthetic Image Dataset for Foundational Pre-Training
📑 Dataset Summary
Open-Pixel-1T is a monumental open-source initiative designed to create a "Visual Atlas" of stochastic imagery. Unlike traditional datasets scraped from social media which contain inherent human bias, Open-Pixel-1T is constructed using high-entropy random seeds to generate unique, diverse visual signals.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LAYEK-143/Open-Pixel-1T.SNAP25_ESM2_OpenFold3_Structural_Analysis
🧬 SNAP25 OpenFold3 Structural Analysis Dataset
Comprehensive structural predictions and therapeutic discovery data for 677 SNAP25 missense variants
🎯 Overview
This dataset provides the first comprehensive structural analysis of SNAP25 (Synaptosomal-Associated Protein 25 kDa) missense variants, generated to support therapeutic discovery for SNAP25-related developmental and epileptic encephalopathy (DEE-SNAP25).
SNAP25 is a critical component of the neuronal… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/SNAP25_ESM2_OpenFold3_Structural_Analysis.cc12m-cleaned
CC12m-cleaned
This dataset builds on two others: The Conceptual Captions 12million dataset, which lead to the LLaVa captioned subset done by
CaptionEmporium
(The latter is the same set, but swaps out the (Conceptual Captions 12million) often-useless alt-text captioning for decent ones_
I have then used the llava captions as a base, and used the detailed descrptions to filter out
images with things like watermarks, artist signatures, etc.
I have also manually thrown out all… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned.multicare-images
MultiCaRe: Open-Source Clinical Case Dataset
MultiCaRe is an open-source, multimodal clinical case dataset built from the PubMed Central Open Access (OA) Case Report articles. It aggregates de-identified, open-access case narratives, figure images, captions, and rich article metadata across diverse specialties (radiology, pathology, surgery, ophthalmology, etc.). The data is normalized so images, cases, and articles can be joined via stable IDs.
Source and process: OA case reports… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/multicare-images.open-palm-hand-images
Palm dataset - 500,000 images
Dataset containing 500,000 human palm images. Designed for hand detection, palm recognition, and gesture analysis, this palm dataset provides diverse training data with metadata on age, gender, and ethnicity for accurate computer vision model training.
By leveraging this dataset, researchers and developers can advance computer vision models for highly accurate hand detection, palm recognition, and gesture analysis. - Get the data… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/open-palm-hand-images.closed-open-eyes
👀 Open and Closed Eyes Dataset
Welcome to the Open and Closed Eyes Dataset! This dataset is designed to help researchers and developers in computer vision and machine learning tasks, particularly in recognizing and distinguishing between open and closed eyes in various contexts. Below, you'll find a detailed description of the dataset structure, categories, and how to interpret the data. 🌟
📁 Dataset Structure
The dataset is stored in Parquet files, ensuring efficient… See the full description on the dataset page: https://huggingface.co/datasets/VasilyLoginov/closed-open-eyes.OpenTME
OpenTME: Open-Access Tumor Microenvironment Profiles from TCGA
OpenTME is an open-access project by Aignostics for academic researchers. It provides comprehensive spatial outputs for whole slide images (WSIs) of H&E-stained, formalin-fixed, paraffin-embedded slides from The Cancer Genome Atlas (TCGA). OpenTME is powered by Atlas H&E-TME – a computational pathology application developed by Aignostics.
Atlas H&E-TME
Atlas H&E-TME is a foundation model-based… See the full description on the dataset page: https://huggingface.co/datasets/Aignostics/OpenTME.amfitrite-open-waters-hab-sentinel2
Dataset Card for Amfitrite-Open-Waters-HAB-Sentinel2
This dataset contains multispectral Sentinel-2 satellite imagery tiles focused on open water and coastal marine environments, classified by the potential presence of Harmful Algal Blooms (HABs).
It is designed to train Machine or Deep Learning models (like CNNs) for large-scale environmental monitoring and ocean anomaly detection.
Dataset Details
Dataset Description
Amfitrite-Open-Waters-HAB-Sentinel2 is a… See the full description on the dataset page: https://huggingface.co/datasets/kostaspic/amfitrite-open-waters-hab-sentinel2.OpenGameArt-GPL-2.0
Dataset Card for OpenGameArt-GPL-2.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 2.0 (GPL-2.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-2.0.uchen_ume_classification_dataset
Uchen–Ume Classification Benchmark
A binary image classification dataset for distinguishing two fundamental categories of Tibetan script: Uchen (དབུ་ཅན།, headed script with a horizontal top stroke) and Ume (དབུ་མེད།, headless script without a top stroke). All images are raw, unprocessed manuscript scans from the Buddhist Digital Resource Center (BDRC).
Model: openpecha/uchen-ume-classifier
Dataset summary
Split
Examples
Uchen
Ume
Train
9,110
~3,124
~5,986… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/uchen_ume_classification_dataset.OpenDeepfake-Preview
OpenDeepfake-Preview Dataset
OpenDeepfake-Preview is a dataset curated for the purpose of training and evaluating machine learning models for deepfake detection. It contains approximately 20,000 labeled image samples with a binary classification: real or fake.
Dataset Details
Task: Image Classification (Deepfake Detection)
Modalities: Image, Video
Format: Parquet
Languages: English
Total Rows: 19,999
File Size: 4.77 GB
License: Apache 2.0
Features
image: The… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenDeepfake-Preview.OpenFake
Dataset Card for OpenFake
Dataset Details
Dataset Description
OpenFake is a dataset designed for evaluating deepfake detection and misinformation mitigation in the context of politically relevant media. It includes high-resolution real and synthetic images generated from prompts with political relevance, including faces of public figures, events (e.g., disasters, protests), and multimodal meme-style images with text overlays. Each image includes structured… See the full description on the dataset page: https://huggingface.co/datasets/karthik-2905/OpenFake.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-3.0.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/irfankabir02/OpenGameArt-GPL-3.0.OpenGameArt-CC-BY-4.0
Dataset Card for OpenGameArt-CC-BY-4.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution 4.0 International (CC-BY-4.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:
English… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-4.0.
