datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
game_character_skins
Game Character Skins Dataset
Summary
This comprehensive dataset contains game character skins and artwork from multiple popular mobile and PC games, providing a rich collection of character visual assets for computer vision research and game development applications. The dataset spans eight major game titles including Arknights, Azur Lane, Blue Archive, Fate/Grand Order, Genshin Impact, Girls' Frontline, Neural Cloud, Nikke, Path to Nowhere, and Honkai: Star Rail… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/game_character_skins.DeepTumorVQA_2.0
DeepTumorVQA v2
3D abdominal-CT diagnostic Visual Question Answering benchmark with 42
clinical subtypes and 438K total QA pairs (10K curated benchmark + 428K
training pool). Includes pre-extracted 2D and video modalities, 20K agent
training trajectories with tool-use traces, and a paper-locked leaderboard.
Resources
📄 Paper (arXiv)
https://arxiv.org/abs/2605.09679
💻 Code (GitHub)
https://github.com/Schuture/DeepTumorVQA
🤗 Dataset (this… See the full description on the dataset page: https://huggingface.co/datasets/tumor-vqa/DeepTumorVQA_2.0.california-flourishing-pollination
California Flourishing & Pollination
DeepEarth × UC Berkeley QED Lab — a self-supervised spatial-feature dataset of every iNaturalist Research-grade observation of every California-native plant and every California-observed flying pollinator, encoded with DINOv3 ViT-L/16 plus PhenoVision flowering/fruiting probabilities.
Maintained by Ecological Intelligence, Inc. (Lance Legel, PI) in collaboration with the Quantitative Ecosystem Dynamics Lab at UC Berkeley (Trevor Keenan, PI).… See the full description on the dataset page: https://huggingface.co/datasets/deepearth/california-flourishing-pollination.generic_character_skins
Generic Character Skins Dataset
Summary
This comprehensive dataset provides an extensive collection of character images sourced from Zerochan across multiple popular anime, game, and manga franchises. The dataset contains meticulously organized character artwork spanning diverse genres including gacha games, idol franchises, fantasy series, and action RPGs. With over 2,000 character folders and thousands of high-quality images, this repository serves as a valuable… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/generic_character_skins.anime_dbrating
Anime Danbooru Rating Dataset
Summary
This dataset provides comprehensive danbooru rating classifications for anime-style images, organized into four distinct categories based on content safety levels. The dataset contains over 1.2 million images distributed across explicit, general, questionable, and sensitive rating classes, making it ideal for training content moderation systems and image classification models specifically tailored for anime and manga-style… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/anime_dbrating.NTIRE-RobustAIGenDetection-train
Training set for NTIRE 2026 Robust AI-Generated Image Detection in the Wild
Robust AI-Generated Image Detection in the Wild Challenge is organized as a part of the New Trends in Image Restoration and Enhancement Workshop in conjunction with CVPR 2026.
Challenge overview
Text-to-image (T2I) models have made synthetic images nearly indistinguishable from real photos in many cases, which creates serious challenges for trust, authenticity, forensics, and content… See the full description on the dataset page: https://huggingface.co/datasets/deepfakesMSU/NTIRE-RobustAIGenDetection-train.deBary
Dataset Card for Dataset Name
Overview
Repository for the deBary plant disease dataset with train/test splits (85%/15%) proposed in [TBD].
Summary
Images: 9,454 JPG files
Train images: 6,065
Test images: 3,389
Species folders: 81
Health statuses: healthy, diseased
Condition labels: 1,265
Climate CSV rows: 70,999
Image Layout
deBary/
├── data/
│ ├── train/
│ │ └── <species>/<health_status>/<condition>/*.jpg
│ └── test/
│… See the full description on the dataset page: https://huggingface.co/datasets/deep-plants/deBary.DeepLearningProject
Deep Learning Project
Dataset Summary
This repository contains the datasets, trained models, notebooks, experiments, feature-extraction outputs, and supporting resources developed for a deep learning project focused on fire detection, fire severity classification, and related computer vision tasks.
The project covers multiple stages of a deep learning workflow, including binary fire classification, three-class fire severity classification, feature extraction… See the full description on the dataset page: https://huggingface.co/datasets/AbdullahImran/DeepLearningProject.20K_real_and_deepfake_images_PCAThis dataset contains the test images used to evaluate our deepfake detection framework. It originally contained 20,000 real and deepfake images, but as some 2600 files are protected by the UK Crown and we do not have a permission to reproduced them, so these files were removed.
Our framework contains 4 machine learning models, which feed in the original images, error-level analysis (ELA) images, noise analysis (NA) images and Principal Component Analysis (PCA) images.
The models were created… See the full description on the dataset page: https://huggingface.co/datasets/ts0pwo/20K_real_and_deepfake_images_PCA.deepfake-celeba
MetaFLOS Deepfake Dataset (CelebA Face Deepfake Generated Images)
Flux2-Klein generated deepfake images from CelebA face descriptions, for deepfake detection / comparison research.
Content
19,867 generated images (full CelebA validation split)
Resolution: 256×256
Each corresponds to a CelebA real face (see real_orig field in manifest)
Generation style: snapshot/crop candid-photo look (not studio portrait) — off-center framing, subject possibly touching or cut by… See the full description on the dataset page: https://huggingface.co/datasets/tjw/deepfake-celeba.NTIRE-RobustAIGenDetection-val
Validation set for NTIRE 2026 Robust AI-Generated Image Detection in the Wild (updated)
Note: This is an updated version of the dataset. For challenge submissions, please make sure you use this version.
Robust AI-Generated Image Detection in the Wild Challenge is organized as a part of the New Trends in Image Restoration and Enhancement Workshop in conjunction with CVPR 2026.
Challenge overview
Text-to-image (T2I) models have made synthetic images nearly… See the full description on the dataset page: https://huggingface.co/datasets/deepfakesMSU/NTIRE-RobustAIGenDetection-val.nsfw_detect
NSFW Detection Dataset
Summary
This dataset is specifically designed for training NSFW (Not Safe For Work) detection models in the context of artistic content and image classification. The collection follows the established categorization format from popular NSFW detection implementations, providing a comprehensive benchmark for content moderation systems. The dataset contains images organized into five distinct classes that represent different levels of appropriateness… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/nsfw_detect.Deepfake-Identity-Isolated-Dataset-PreP
Identity-Isolated Deepfake Face Images Dataset.
A rigorously preprocessed, identity-aware deepfake detection dataset of 142,837 face images, constructed for training generalized deepfake detectors across Face Swap and Entire Face Synthesis manipulation categories.
Built as part of the DFDS project - an open-source deepfake detection API for FinTech KYC identity verification. The full system is available on GitHub at DFDS-XAI
Dataset Summary.
Property
Value… See the full description on the dataset page: https://huggingface.co/datasets/ThinothW/Deepfake-Identity-Isolated-Dataset-PreP.Deepface_Annotated_3K
Deepface Annotated 3K Dataset
Deepface_Annotated_3K is a synthetic facial image dataset containing 3K AI-generated faces from StyleGAN2.Each image is automatically annotated with demographic attributes like:
Age (in years)
Gender (with prediction confidence)
Dominant Race (White, Asian, Latino Hispanic, Indian, etc.)
The dataset is designed for research on fairness, bias detection, demographic classification, and synthetic face representation.
Structure information… See the full description on the dataset page: https://huggingface.co/datasets/Subh775/Deepface_Annotated_3K.AGM
Dataset Card for AGM Dataset
Dataset Summary
The AGM (AGricolaModerna) Dataset is a comprehensive collection of high-resolution RGB images capturing harvest-ready plants in a vertical farm setting. This dataset consists of 972,858 images, each with a resolution of 120x120 pixels, covering 18 different plant crops. In the context of this dataset, a crop refers to a plant species or a mix of plant species.
Supported Tasks
Image classification: plant phenotyping… See the full description on the dataset page: https://huggingface.co/datasets/deep-plants/AGM.Deepfake-vs-Real-v2
Deepfake-vs-Real-v2
Deepfake-vs-Real-v2 is a dataset designed for image classification, distinguishing between deepfake and real images. This dataset includes a diverse collection of high-quality deepfake images to enhance classification accuracy and improve the model’s overall efficiency. By providing a well-balanced dataset, it aims to support the development of more robust deepfake detection models.
Label Mappings
Mapping of IDs to Labels: {0: 'Deepfake', 1:… See the full description on the dataset page: https://huggingface.co/datasets/dappai/Deepfake-vs-Real-v2.NTIRE-RobustAIGenDetection-test-public
Test set for NTIRE 2026 Robust AI-Generated Image Detection in the Wild
Robust AI-Generated Image Detection in the Wild Challenge is organized as a part of the New Trends in Image Restoration and Enhancement Workshop in conjunction with CVPR 2026.
Challenge overview
Text-to-image (T2I) models have made synthetic images nearly indistinguishable from real photos in many cases, which creates serious challenges for trust, authenticity, forensics, and content safety. At… See the full description on the dataset page: https://huggingface.co/datasets/deepfakesMSU/NTIRE-RobustAIGenDetection-test-public.GAMMA
GAMMA — Glaucoma grading from Multi-Modality imAges (Challenge dataset)
Image: Dataset Samples.
Short description
GAMMA is the first public multi-modality glaucoma grading dataset that pairs 2D color fundus photographs with 3D OCT volumes for each sample. It was released as part of the GAMMA challenge (OMIA8 / MICCAI 2021) to encourage algorithms that combine fundus and OCT information for… See the full description on the dataset page: https://huggingface.co/datasets/DeepuJames/GAMMA.20K_real_and_deepfake_imagesThis dataset contains the test images used to evaluate our deepfake detection framework. It originally contained 20,000 real and deepfake images, but as some 2600 files are protected by the UK Crown and we do not have a permission to reproduced them, so these files were removed.
Our framework contains 4 machine learning models, which feed in the original images, error-level analysis (ELA) images, noise analysis (NA) images and Principal Component Analysis (PCA) images.
The models were created… See the full description on the dataset page: https://huggingface.co/datasets/ts0pwo/20K_real_and_deepfake_images.MAHE_deepfake_recognition_datasetThis dataset is for different educational experiments of deepfake images classification.
It consists of the small datasets with various original and deepfake scenes (faces, animation, urban scenes and others):
Columbia Uncompressed Image Splicing Dataset
Kaggle Deepfake Dataset Challenge
CommunityForensics-Small
AI-vs-Deepfake-vs-Real-Resized-Aug
🧠 AI vs Deepfake vs Real — Processed Version
This dataset is the result of preprocessing and augmentation applied to the original datasetprithivMLmods/AI-vs-Deepfake-vs-Real.
📘 Overview
This dataset contains a collection of images categorized into three main classes:
🟩 AI-generated
🟥 Deepfake
🟦 Real (authentic human faces)
It is designed for image classification tasks that aim to distinguish between AI-generated, deepfake, and real faces.… See the full description on the dataset page: https://huggingface.co/datasets/chintalaswathi/AI-vs-Deepfake-vs-Real-Resized-Aug.deepscan-dataset
Tropical Reef Fish Dataset
A dataset of tropical reef fish images for species-level image classification, with additional negative classes for out-of-distribution rejection.
Dataset Description
Images were scraped from iNaturalist using research-grade observations with CC0 or CC-BY licenses. The dataset covers 12 species across 9 families, targeting fish commonly encountered by snorkelers at depths of 0–10 meters in tropical reef environments.
Two negative classes are… See the full description on the dataset page: https://huggingface.co/datasets/fish-gang/deepscan-dataset.20K_real_and_deepfake_images_NAThis dataset contains the test images used to evaluate our deepfake detection framework. It originally contained 20,000 real and deepfake images, but as some 2600 files are protected by the UK Crown and we do not have a permission to reproduced them, so these files were removed.
Our framework contains 4 machine learning models, which feed in the original images, error-level analysis (ELA) images, noise analysis (NA) images and Principal Component Analysis (PCA) images.
The models were created… See the full description on the dataset page: https://huggingface.co/datasets/ts0pwo/20K_real_and_deepfake_images_NA.Deepfake_dataset_cybersentinal
Dataset Card for OpenFake
OpenFake is a dataset and benchmark for detecting AI-generated images, with a focus on politically and socially salient content where misinformation risk is highest. It pairs real photographs with synthetic counterparts produced by a wide range of frontier proprietary generators, open-source diffusion models, and community fine-tunes. A separate in-the-wild test set is sourced from Reddit to evaluate detector performance on naturally circulated… See the full description on the dataset page: https://huggingface.co/datasets/vjjoshi23/Deepfake_dataset_cybersentinal.20K_real_and_deepfake_images_ELAThis dataset contains the test images used to evaluate our deepfake detection framework. It originally contained 20,000 real and deepfake images, but as some 2600 files are protected by the UK Crown and we do not have a permission to reproduced them, so these files were removed.
Our framework contains 4 machine learning models, which feed in the original images, error-level analysis (ELA) images, noise analysis (NA) images and Principal Component Analysis (PCA) images.
The models were created… See the full description on the dataset page: https://huggingface.co/datasets/ts0pwo/20K_real_and_deepfake_images_ELA.DeepTrees_Halle
DeepTrees Halle DOP20 labels + imagery
This is a sample dataset for model training and fine-tuning in tree crown segmentation tasks using the DeepTrees Python package.
Overview of subtiles with sample of labels:
Dataset Details
We have taken a single Multispectral (RGBi) 2x2 km DOP20 image tile for Halle, Sachsen-Anhalt, from LVermGeo ST for the year of 2022. TileID from source: 32_704_5708_2
We then sliced the tiles into subtiles of 100x100m, resulting in 400… See the full description on the dataset page: https://huggingface.co/datasets/thisistaimur/DeepTrees_Halle.Deepfakes-QA-15K
Deepfake Quality Assessment
Deepfake QA is a Deepfake Quality Assessment model designed to analyze the quality of deepfake images & videos. It evaluates whether a deepfake is of good or bad quality, where:
0 represents a bad-quality deepfake
1 represents a good-quality deepfake
This classification serves as the foundation for training models on deepfake quality assessment, helping improve deepfake detection and enhancement techniques.
Citation
If you use our… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Deepfakes-QA-15K.test_nested_dataset
Image Classification - Monochrome Or Not
2 labels, 260 samples in total, listed as the following:
Label
Samples
Sample #0
Sample #1
Sample #2
Sample #3
Sample #4
Sample #5
Sample #6
Sample #7
monochrome
4 (1.5%)
N/A
N/A
N/A
N/A
colored
256 (98.5%)
evaluation-dataset
DeepSafe Evaluation Dataset
Evaluation set for DeepSafe,
a deepfake detection benchmark.
Tiers
Tier
Samples
Generators
Size
Use
master_eval_small/
198
116
1.7 GB
smoke test, under 2 min
master_eval/
15,454
411
10 GB
the standard benchmark
master_eval_full/
45,954
411
25 GB
complete set
Medium tier composition: 9,954 image, 3,500 audio, 2,000 video.
from huggingface_hub import snapshot_download
snapshot_download("deepsafe/evaluation-dataset"… See the full description on the dataset page: https://huggingface.co/datasets/deepsafe/evaluation-dataset.deepfake-detection-dataset-v3
Deepfake Detection Dataset V3
This dataset contains images and detailed explanations for training and evaluating deepfake detection models. It includes original images, manipulated images, confidence scores, and comprehensive technical and non-technical explanations.
Dataset Structure
The dataset consists of:
Original images (image)
CAM visualization images (cam_image)
CAM overlay images (cam_overlay)
Comparison images (comparison_image)
Labels (label): Binary… See the full description on the dataset page: https://huggingface.co/datasets/saakshigupta/deepfake-detection-dataset-v3.
