datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenFake
Dataset Card for OpenFake
Known issues
Prompt–image misalignment in the synthetic split (reported November 2025, fix pending)
For five of the eighty generators, the prompt field attached to synthetic
images does not correspond to the prompt actually used to generate that image.
Affected generators:
flux-realism
sd-3.5
sdxl-realvis-v5
sd-1.5-dreamshaper
sd-1.5-epicdream
This affects approximately 19.77% of synthetic images. It was first reported in
discussion… See the full description on the dataset page: https://huggingface.co/datasets/ComplexDataLab/OpenFake.OpenSDI_trainThis repository contains the OpenSDI training dataset, presented in the paper OpenSDI: Spotting Diffusion-Generated Images in the Open World.
Code: https://github.com/iamwangyabin/OpenSDI
openbrush
OpenBrush-75K
A curated dataset of 75,313 public domain artworks with rich, structured VLM-generated captions designed for training image generation models, fine-tuning vision-language models, and art analysis research.
Dataset Description
OpenBrush-75K contains high-quality reproductions of paintings from the Western art canon, spanning from the Renaissance to the early 20th century. Each image is paired with a detailed structured caption generated by a… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush.OpenSDI_test
OpenSDI: Spotting Diffusion-Generated Images in the Open World
This dataset is designed to address the OpenSDI challenge: spotting diffusion-generated images in realistic, open-world scenarios. It is described in the paper:
Project Page: https://iamwangyabin.github.io/OpenSDI/
OpenSDID Dataset Highlights:
User Diversity: Simulates a wide range of user intentions and creative styles using diverse text prompts generated by VLMs.
Model Innovation: Includes images from multiple… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDI_test.Open-Pixel-1T
🌌 Open-Pixel-1T (Visual Atlas)
A Large-Scale, High-Entropy Synthetic Image Dataset for Foundational Pre-Training
📑 Dataset Summary
Open-Pixel-1T is a monumental open-source initiative designed to create a "Visual Atlas" of stochastic imagery. Unlike traditional datasets scraped from social media which contain inherent human bias, Open-Pixel-1T is constructed using high-entropy random seeds to generate unique, diverse visual signals.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LAYEK-143/Open-Pixel-1T.OpenAI-4o_t2i_human_preference
Rapidata OpenAI 4o Preference
This T2I dataset contains over 200'000 human responses from over ~45,000 individual annotators, collected in less than half a day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating OpenAI 4o (version from 26.3.2025) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/OpenAI-4o_t2i_human_preference.closed-open-eyes
👀 Open and Closed Eyes Dataset
Welcome to the Open and Closed Eyes Dataset! This dataset is designed to help researchers and developers in computer vision and machine learning tasks, particularly in recognizing and distinguishing between open and closed eyes in various contexts. Below, you'll find a detailed description of the dataset structure, categories, and how to interpret the data. 🌟
📁 Dataset Structure
The dataset is stored in Parquet files, ensuring efficient… See the full description on the dataset page: https://huggingface.co/datasets/MichalMlodawski/closed-open-eyes.openbrush-75k
OpenBrush-75K
A curated dataset of 75,313 public domain artworks with rich, structured VLM-generated captions designed for training image generation models, fine-tuning vision-language models, and art analysis research.
Dataset Description
OpenBrush-75K contains high-quality reproductions of paintings from the Western art canon, spanning from the Renaissance to the early 20th century. Each image is paired with a detailed structured caption generated by a vision-language… See the full description on the dataset page: https://huggingface.co/datasets/Trever896/openbrush-75k.OpenGVLab_Lumina_t2i_human_preference
Rapidata Lumina Preference
This T2I dataset contains over 400k human responses from over 86k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Lumina across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/OpenGVLab_Lumina_t2i_human_preference.openbrush-landscapes
OpenBrush Landscapes
Every landscape painting from OpenBrush-75K — across all artists, movements, and centuries. Largest single-genre subset.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 12,612 you actually want.
Why this subset
Every landscape across the parent dataset's full range — Romantic wildernesses, Impressionist… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-landscapes.OpenSDIDplus
OpenSDID+
OpenSDID+ is an extended release of the OpenSDI dataset. It complements the original SD1.5 training split with large-scale images from the remaining OpenSDI generators: SD2, SD3, SDXL, and FLUX.
The dataset follows the OpenSDI challenge introduced in "OpenSDI: Spotting Diffusion-Generated Images in the Open World". OpenSDI studies detection and localization of diffusion-generated images under realistic open-world settings, including diverse user intentions, evolving… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDIDplus.closed-open-eyes
👀 Open and Closed Eyes Dataset
Welcome to the Open and Closed Eyes Dataset! This dataset is designed to help researchers and developers in computer vision and machine learning tasks, particularly in recognizing and distinguishing between open and closed eyes in various contexts. Below, you'll find a detailed description of the dataset structure, categories, and how to interpret the data. 🌟
📁 Dataset Structure
The dataset is stored in Parquet files, ensuring efficient… See the full description on the dataset page: https://huggingface.co/datasets/VasilyLoginov/closed-open-eyes.openbrush-impressionism
OpenBrush Impressionism
Every Impressionist work from OpenBrush-75K — the largest movement subset.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 12,798 you actually want.
Why this subset
Broad-coverage subset for training on the Impressionist visual language: broken brushwork, light-on-color theory, plein-air staging, atmospheric… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-impressionism.multicare-images
MultiCaRe: Open-Source Clinical Case Dataset
MultiCaRe is an open-source, multimodal clinical case dataset built from the PubMed Central Open Access (OA) Case Report articles. It aggregates de-identified, open-access case narratives, figure images, captions, and rich article metadata across diverse specialties (radiology, pathology, surgery, ophthalmology, etc.). The data is normalized so images, cases, and articles can be joined via stable IDs.
Source and process: OA case reports… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/multicare-images.OpenDeepfake-Preview
OpenDeepfake-Preview Dataset
OpenDeepfake-Preview is a dataset curated for the purpose of training and evaluating machine learning models for deepfake detection. It contains approximately 20,000 labeled image samples with a binary classification: real or fake.
Dataset Details
Task: Image Classification (Deepfake Detection)
Modalities: Image, Video
Format: Parquet
Languages: English
Total Rows: 19,999
File Size: 4.77 GB
License: Apache 2.0
Features
image: The… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenDeepfake-Preview.openbrush-impressionist-landscapes
OpenBrush Impressionist Landscapes
Cross-cut subset: Impressionist landscape paintings from OpenBrush-75K. The most-targeted style+genre combination for Impressionist landscape LoRA training.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 4,308 you actually want.
Why this subset
The intersection of the largest movement… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-impressionist-landscapes.OpenFake
Dataset Card for OpenFake
Dataset Details
Dataset Description
OpenFake is a dataset designed for evaluating deepfake detection and misinformation mitigation in the context of politically relevant media. It includes high-resolution real and synthetic images generated from prompts with political relevance, including faces of public figures, events (e.g., disasters, protests), and multimodal meme-style images with text overlays. Each image includes structured… See the full description on the dataset page: https://huggingface.co/datasets/karthik-2905/OpenFake.OpenJev-Vision-Research-v0.1
OpenJev Vision Research v0.1
12,832 image records, with public provenance, original synthetic scenes,
and programmatically derived decision questions.
This is an experimental research dataset for visual posterior learning and
compositional decisions, released with OpenJev.
It is not a reproduction of TypeSafe's proprietary Jev model or training method.
Three separate configurations
Config
Images
What the labels mean
License
synthetic
8,192
Exact… See the full description on the dataset page: https://huggingface.co/datasets/IamBusy/OpenJev-Vision-Research-v0.1.uchen_ume_classification_dataset
Uchen–Ume Classification Benchmark
A binary image classification dataset for distinguishing two fundamental categories of Tibetan script: Uchen (དབུ་ཅན།, headed script with a horizontal top stroke) and Ume (དབུ་མེད།, headless script without a top stroke). All images are raw, unprocessed manuscript scans from the Buddhist Digital Resource Center (BDRC).
Model: openpecha/uchen-ume-classifier
Dataset summary
Split
Examples
Uchen
Ume
Train
9,110
~3,124
~5,986… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/uchen_ume_classification_dataset.openbrush-religious-art
OpenBrush Religious Art
Religious paintings from OpenBrush-75K — saints, biblical scenes, devotional works.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 6,119 you actually want.
Why this subset
A coherent visual genre: religious narrative painting from medieval through early modern. Heavy on Renaissance and Baroque eras. Common… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-religious-art.openbrush-baroque
OpenBrush Baroque
Baroque works from OpenBrush-75K (~1600–1750) — chiaroscuro, religious painting, dramatic light.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 4,240 you actually want.
Why this subset
The canonical Baroque visual language — Caravaggio, Rembrandt, Vermeer, Velázquez, Rubens. Useful for models learning dramatic… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-baroque.openart-items-artifacts
OpenArt — Items & Artifacts
openart-items-artifacts is the items artifacts subject collection of the OpenArt family
of open, public-domain art datasets: 25,750 works (11,317 paintings/illustrations · 14,216
photographed objects · 217 unclassified), each paired with a structured VLM caption plus
medium, attribution and inscription metadata.
Human-made objects and the decorative arts — vessels, tools, arms and armor, textiles, furniture and ornament — both as physical artifacts… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openart-items-artifacts.nationalmuseet-open-images
Nationalmuseet Open Images
This dataset is an independently harvested research dataset from Nationalmuseet Samlinger Online.
It contains metadata and optionally WebDataset image shards for Nationalmuseet asset records whose
rights.license is one of:
Public Domain
CC-BY
No known rights
Public Domain and CC-BY are the strict open-license subset. No known rights is kept as a
separate license bucket because Nationalmuseet says this label means that, to their best assessment,
the… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/nationalmuseet-open-images.openbrush-anonymous-masters
OpenBrush Anonymous Masters
Unattributed works from OpenBrush-75K — anonymous old masters across centuries and styles. Useful for broad-style training without artist-specific bias.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 41,914 you actually want.
Why this subset
37% of the parent dataset is unattributed — a substantial… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-anonymous-masters.openbrush-renaissance
OpenBrush Renaissance
Renaissance works from OpenBrush-75K, combining Northern, Early, High, and Mannerism Late Renaissance into one period subset.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 6,565 you actually want.
Why this subset
Combined Renaissance period subset spanning ~1300–1600. Heavy on religious painting, portraits… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-renaissance.openbrush-ukiyo-e
OpenBrush Ukiyo-e
Japanese Ukiyo-e woodblock prints from OpenBrush-75K — the only non-Western style in the parent dataset, separated here for trainers who want a dedicated Japanese woodblock corpus without downloading 75K Western paintings to filter.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 1,167 you actually want.
Why this… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-ukiyo-e.openbrush-portraits
OpenBrush Portraits
Every portrait painting from OpenBrush-75K — across all artists, movements, and centuries.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 13,059 you actually want.
Why this subset
Portraits across the full historical range — Renaissance bust portraits, Baroque chiaroscuro, Rococo society, Romantic, Realist… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-portraits.OpenEar_developmental_status_classification
Openear Developmental Status Classification
This dataset provides real-world RGB images of maize ears collected in a field environment at Hongqi Base, Hainan, China, for developmental status classification. Images were captured using a ground-based Raspberry Pi HQ camera system with a Sony IMX477R sensor during the 2025-2026 growing season, offering high-resolution visual data for distinguishing between abnormal and normal developmental stages. The dataset contains 6,435 images… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/OpenEar_developmental_status_classification.openbrush-van-gogh
OpenBrush Van Gogh
All Vincent van Gogh works from OpenBrush-75K, with structured VLM captions.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 1,889 you actually want.
Why this subset
Van Gogh's catalog spans his Realism period (Dutch landscapes, peasant scenes) through his Post-Impressionist breakthroughs (Arles, Saint-Rémy… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-van-gogh.openm3chest-labels-v2
OpenM3Chest Labels v2
Stratified label subset of OpenM3Chest for LoRA fine-tuning of MedGemma 1.5-4B-IT.
CT scan images (NPY): UngLong/openm3chest-npy-v2
Dataset Summary
Agent
Tasks
Train samples
Test samples
Radiology
chest_abn_54–61, nodule_presence/location/attenuation/margin/size
~350/task
full
Cardiology
CVD_diagnosis, CVD_mortality
1,500/task
3,759/task
Oncology
lung_cancer_risk
2,000
10,308
Training subset: stratified sampling — binary tasks… See the full description on the dataset page: https://huggingface.co/datasets/UngLong/openm3chest-labels-v2.
