datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.beans
Dataset Card for Beans
Dataset Summary
Beans leaf dataset with images of diseased and health leaves.
Supported Tasks and Leaderboards
image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any.
Languages
English
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/AI-Lab-Makerere/beans.3d-defectbench
3D-DefectBench
A benchmark and controlled study of vision-language models (VLMs) as judges for
fine-grained defect detection in text-to-3D generated assets. Each asset carries
a 9-dimensional binary defect vector over five geometry and four texture defect
categories, three of which are prompt-conditioned.
Version 1.1 adds the complete set of 1,000 benchmark GLB assets, the cell-level
VLM prediction table, and the TRELLIS cross-generator prompts and prediction
tables. TRELLIS… See the full description on the dataset page: https://huggingface.co/datasets/aieval2026/3d-defectbench.AID
Aerial Image Dataset (AID)
Description
The Aerial Image Dataset (AID) is a scene classification dataset consisting of 10,000 RGB images, each with a resolution of 600x600 pixels. These images have been extracted using Google Earth and cover various scenes from regions and countries around the world. AID comprises 30 different scene categories, with several hundred images per class.
The new dataset is made up of the following 30 aerial scene types: airport, bare… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/AID.brain-structureA collection of T1-weighted .nii.gz structural MRI scans in a BIDS-like arrangement,
with JSON sidecar metadata indicating train/validation/test splits.RAIDThis dataset is for testing the adversarial robustness of AI-Generated Image Detectors, as described in the paper RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors.
aidovecl-vehicle-detection-classification-localization
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models.
Citation Notice
Please ensure that all publications and presentations using this data reference the following paper:
Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.AIGIBench
Is Artificial Intelligence Generated Image Detection a Solved Problem?
Ziqiang Li1, Jiazhen Yan1, Ziwen He1, Kai Zeng2, Weiwei Jiang1, Lizhi Xiong1, Zhangjie Fu1‡
‡Corresponding author
1Nanjing University of Information Science and Technology 2University of Siena
Paper | GitHub Repository
This repository is the official dataset of the AIGIBench.
AIGIBench dataset contains two types of training and 25 test subsets. This dataset has the following advantages:
Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/HorizonTEL/AIGIBench.military-aircraft-detection-dataset
Military Aircraft Detection Dataset
Military aircraft detection dataset in COCO and YOLO format.
The dataset contains 103 different military aircraft types.
['A10', 'A400M', 'AG600', 'AH64', 'AKINCI', 'AV8B', 'An124', 'An22', 'An225', 'An72', 'B1', 'B2', 'B21', 'B52', 'Be200', 'C1', 'C130', 'C17', 'C2', 'C390', 'C5', 'CH47', 'CH53', 'CL415', 'E2', 'E7', 'EF2000', 'EMB314', 'F117', 'F14', 'F15', 'F16', 'F18', 'F2', 'F22', 'F35', 'F4', 'FCK1', 'H6', 'Il76', 'J10', 'J20', 'J35'… See the full description on the dataset page: https://huggingface.co/datasets/a2015003713/military-aircraft-detection-dataset.anything-v3.0-glazed
Dataset Card for Anything v3.0 Glazed Samples
Dataset Description
Dataset Summary
This dataset contains image samples originally generated by Linaqruf/anything-v3.0
and subsequently processed by Glaze tool.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/anything-v3.0-glazed.stable-diffusion-v1-5-glazed
Dataset Card for Stable Diffusion v1.5 Glazed Samples
Dataset Description
Dataset Summary
This dataset contains image samples originally generated by runwayml/stable-diffusion-v1-5
and subsequently processed by Glaze tool.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/stable-diffusion-v1-5-glazed.FGVC-Aircraft
Dataset Card for FGVC-Aircraft
This is a FiftyOne dataset with 10000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = fouh.load_from_hub("Voxel51/FGVC-Aircraft")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/FGVC-Aircraft.Kuro-Siwo-Webdataset
Kuro Siwo webdatasets
Paper | GitHub |
Dataset Details
Dataset Description
Kuro Siwo is a global multi-temporal SAR dataset for rapid flood mapping. It contains 43 flood events in 6 continents and 3 climate zones, over the period 2015-2022. The annotations have been produced through meticulous photointerpretation by a team of experts, at 10m spatial resolution. For each flood event, we provide one Sentinel-1 post-flood and two Sentinel-1 pre-flood… See the full description on the dataset page: https://huggingface.co/datasets/orion-ai-lab/Kuro-Siwo-Webdataset.visual_ai_at_neurips2025
Dataset Card for neurips-2025-vision-papers
This is a FiftyOne dataset with 1134 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/visual_ai_at_neurips2025")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/visual_ai_at_neurips2025.AIGC-Detection-Benchmark
AIGC Detection Benchmark Dataset
📝 Dataset Description
Dataset Summary
The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.AI_REAL
AI vs REAL Image Dataset
Ce dataset contient deux classes d’images :
AI : images générées par intelligence artificielle
REAL : images réelles
Les fichiers sont organisés par batchs pour respecter les limites de Hugging Face.
gpt-image-2
GPT-Image-2 Twitter Dataset
10,217 confirmed GPT-image-2.0 generated images collected from Twitter/XCollection window: April 21 – April 28, 2026 (first week post-launch)Paper: GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment
Overview
This dataset contains 10,217 images confirmed to be GPT-image-2.0 outputs, sourced from public Twitter/X posts in the immediate aftermath of the model's April 21, 2026 release… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/gpt-image-2.AI-GenBench-fake_part
AI-GenBench: A New Ongoing Benchmark for AI-Generated Image Detection
Important: this is the fake part of the AI-GenBench dataset. To re-create the original benchmark, which includes real images, please check the official repository.
Important 2: before using, please check the licensing terms of the images included!
Details
The rapid advancement of generative AI has revolutionized image creation, enabling high-quality synthesis from text prompts while raising critical… See the full description on the dataset page: https://huggingface.co/datasets/lrzpellegrini/AI-GenBench-fake_part.real-fake-ai-generated-art-images
🎨 Real and Fake (AI-Generated) Art Images Dataset
21,642 balanced images — 10,821 real artworks and 10,821 AI-generated
images — for training models to distinguish authentic art from GAN-generated fakes.
🧭 Overview
This dataset is part of the FauxFinder project, designed to build
advanced models capable of distinguishing between authentic artworks
and AI-generated images. Ideal for binary classification, GAN research,
and computer vision benchmarking.… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/real-fake-ai-generated-art-images.AIForge-Doc-v1
AIForge-Doc: A Benchmark of AI-Forged Document Images
AIForge-Doc is the first large-scale benchmark of AI-forged document images, targeting
financial and identity document fraud. Every tampered image was produced by a
diffusion-model inpainting pipeline — a threat model that existing forgery detectors
cannot reliably handle.
At a Glance
Attribute
Value
Total forged images
4,061
Training split
3,249 (80 %)
Testing split
812 (20 %)
Authentic… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v1.ai-image-detector-dataset
AI Image Detector Dataset (training-v1)
wkaandemir/ai-image-detector modelini eğitmek için kullanılan, 20.000 normalize edilmiş görselden oluşan dengeli ve kaynak-farkında (source-aware) bir görüntü sınıflandırma veri kümesi. Görev, görselleri gerçek (real) ve yapay (fake) olarak ikili sınıflandırmaktır.
Bu veri kümesi, modelin kalibrasyon ve eşik seçiminde kullanılmayan, genelleme ölçümü için kaynak bazlı ayrı tutulmuş bir external_test split'i de içerir.
📌… See the full description on the dataset page: https://huggingface.co/datasets/wkaandemir/ai-image-detector-dataset.AIForge-Doc-v2
AIForge-Doc v2: A Paired Benchmark of GPT-Image-2 Document Forgeries
AIForge-Doc v2 is the first paired benchmark of document forgeries produced by
OpenAI's GPT-Image-2 (released April 2026). Every forged image is accompanied by
its authentic source image and a pixel-precise tampered-region mask in
DocTamper-compatible format. v2 reuses the forgery specifications of
AIForge-Doc v1 spec-for-spec and swaps only
the generator, so any difference in detector behaviour between v1… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v2.aircraft-images
Dataset Card for High-Resolution Aircraft Images
Dataset Summary
This dataset contains 165,340 high-resolution aircraft images collected from the internet, along with machine-generated captions. The captions were generated using Gemini Flash 1.5 AI model and are stored in separate text files matching the image filenames.
Languages
The dataset is monolingual:
English (en): All image captions are in English
Dataset Structure
Data Files
The… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/aircraft-images.FGVC-Aircraft
Dataset Card for "FGVC-Aircraft"
This is a non-official FGVC-Aircraft dataset for fine-grained Image Classification.
If you want to download the official dataset, please refer to the here.
Military-Aircraft-Recognition-datasetThis is a remote sensing image Military Aircraft Recognition dataset that include 3842 images, 20 types, and 22341 instances annotated with horizontal bounding boxes and oriented bounding boxes.
ROBIN-ImagesGT-Merged-Flanora-AI-v1
Flanora AI/ROBIN-ImagesGT-Merged-Flanora-AI-v1
ROBIN-ImagesGT-Merged-Flanora-AI-v1 is a curated collection of 622 floor-plan images created by merging floor-plan data from the ROBIN dataset and the CVC-FP / ImagesGT dataset.
The dataset is organized into six bedroom-count categories: 0 bedroom, 1 bedroom, 2 bedroom, 3 bedroom, 4 bedroom, and 5 bedroom.
The dataset contains the original, unprocessed floor-plan images. No image preprocessing or transformation was applied to the… See the full description on the dataset page: https://huggingface.co/datasets/BJyotibrat/ROBIN-ImagesGT-Merged-Flanora-AI-v1.SA-BENCH
SA-BENCH
SA-BENCH is the benchmark dataset released with “Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics.”
Accepted to CVPRW 2026.
GitHub | CVF Open Access | arXiv | Model
It evaluates the spatial aesthetics of interior images along four dimensions:
distortion
harmony
layout
lighting
SA-BENCH contains 17,768 annotated examples across four spatial-aesthetic dimensions, with image assets and human annotations for training and… See the full description on the dataset page: https://huggingface.co/datasets/gaoyuan-ai/SA-BENCH.aidetector-data
Companion dataset for the aidetector
project (live demo: https://humanorai.online). Code, trained model, and the
reproducibility record live in that GitHub repo.
Dataset Card — AI Image Detector
This card documents the data the v2 models were trained and evaluated on. All
counts are taken from the reproducibility record
experiment_v1.json (dataset content hash
72b88efc0497...). The image files themselves are not redistributed in this
repository (see Access below).
The intent… See the full description on the dataset page: https://huggingface.co/datasets/aman213/aidetector-data.kriyam-tamperflow
Kriyam TamperFlow
The first document tampering detection benchmark built specifically for Indian documents, with a built-in compression stress-test that exposes how quickly forensic models degrade on real-world scanned material.
Dataset Summary
State-of-the-art document forgery detectors — CAT-Net, DTD, MVSS-Net, CAFTB, and others — rely on JPEG compression artifacts as their primary forensic signal: inconsistencies in DCT coefficients, block boundaries, and… See the full description on the dataset page: https://huggingface.co/datasets/kriyam-ai/kriyam-tamperflow.AI-vs-Real
🖼️ AI-vs-Real Dataset
A balanced dataset for AI-generated vs Real image classification.This dataset is designed to help researchers, developers, and practitioners build and evaluate models that can distinguish between synthetic (AI-generated) and authentic (human-captured) images.
📊 Dataset Overview
Classes:
0 → AI-generated images
1 → Real (human-captured) images
Balance:The dataset is properly balanced across both classes.This ensures that models… See the full description on the dataset page: https://huggingface.co/datasets/Parveshiiii/AI-vs-Real.
