datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AIGC-Detection-Benchmark
AIGC Detection Benchmark Dataset
📝 Dataset Description
Dataset Summary
The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.detection-images
Hassan881/detection-images
Staging dataset for Berhan XAI (BurhanXAI) AI-generated / edited-media detection.
Field
Value
Training contract
v0.3.0 (firefly top-up applied; 6/6 M2 providers)
Held-out eval (M2)
v0.3.0-eval
Held-out eval (legacy)
v0.2.1-eval
Status
published
Classes
real, ai_generated, ai_edited
Seed
42
from datasets import load_dataset
ds = load_dataset("Hassan881/detection-images", name="v0.3.0")
ev =… See the full description on the dataset page: https://huggingface.co/datasets/Hassan881/detection-images.early_printed_books_font_detection
Early Printed Books Font Detection
Photographs of 35,623 pages from books printed between the mid-15th and the end of the 18th century, each labelled by experts with the font group or groups used on the page. This is a mirror of Dataset of Pages from Early Printed Books with Multiple Font Groups by Mathias Seuret, Saskia Limbach, Nikolaus Weichselbaumer, Andreas Maier and Vincent Christlein, deposited on Zenodo in August 2019 and described in their HIP'19 paper.
The page images… See the full description on the dataset page: https://huggingface.co/datasets/biglam/early_printed_books_font_detection.xai-attack-detection-cifar10
XAI Attack Detection — CIFAR-10 PGD
This private research dataset contains balanced, paired clean and adversarial images for
studying whether an attack can be detected from a classifier explanation map.
Dataset construction
The source is the CIFAR-10 test split. A fine-tuned OpenCLIP ViT-B/16 classifies each
clean image. Clean-correct examples are attacked with untargeted L-infinity PGD using
epsilon 8/255, step size 2/255, 10 steps, and deterministic random… See the full description on the dataset page: https://huggingface.co/datasets/nimaeb/xai-attack-detection-cifar10.strawberry_growth_detection
Strawberry Growth Detection
A dataset for detection of strawberry growth stages. The dataset contains 1,477 images with 3,997 bounding box annotations across 7 categories. The dataset also contains ground truth data
related to the size of the strawberries from tagged leaves, as well as a decimal based growth stage.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{yang2024predicting… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/strawberry_growth_detection.Document-Type-Detection
Document-Type-Detection
Dataset Summary
The Document-Type-Detection dataset is a large-scale image classification dataset consisting of scanned or photographed document images. Each image is categorized into one of nine document types. This dataset is ideal for training document classification models in finance, administration, OCR, and automation workflows.
Supported Tasks
Multiclass Document Classification
Classify an input document image into one of the… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Document-Type-Detection.deepfake-detection-dataset-v3
Deepfake Detection Dataset V3
This dataset contains images and detailed explanations for training and evaluating deepfake detection models. It includes original images, manipulated images, confidence scores, and comprehensive technical and non-technical explanations.
Dataset Structure
The dataset consists of:
Original images (image)
CAM visualization images (cam_image)
CAM overlay images (cam_overlay)
Comparison images (comparison_image)
Labels (label): Binary… See the full description on the dataset page: https://huggingface.co/datasets/saakshigupta/deepfake-detection-dataset-v3.garbage-image-classification-detection
Garbage Image dataset
Dataset consists of images, bounding-boxes and segmentations for each elements.
@misc{
garbage-classifier-oehkt_dataset,
title = { Garbage Classifier Dataset },
type = { Open Source Dataset },
author = { Student },
howpublished = { \url{ https://universe.roboflow.com/student-utr07/garbage-classifier-oehkt } },
url = { https://universe.roboflow.com/student-utr07/garbage-classifier-oehkt },
journal = { Roboflow Universe },
publisher = { Roboflow… See the full description on the dataset page: https://huggingface.co/datasets/dmedhi/garbage-image-classification-detection.glaucoma-detection
Glaucoma Detection
Retinal fundus images for glaucoma stage classification.
Splits
Split
Samples
train
2847
validation
1259
test
1272
Columns
image: retinal image
class: glaucoma stage label
Classes
Class
Description
normal
No glaucoma visible in the image.
early
Early-stage glaucoma findings.
advanced
Advanced glaucoma findings.
fruit-ripeness-detection-dataset
Dataset Card for Fruit-Ripeness-Classification dataset
This is a collection of ripe and unripe fruits (mangoes and bananas) in outside lighting and outside conditions.
Train - 80% (4k images)
Test - 20% (1k images)
Dimensions of image : 640 x 480
The dataset has been collected from Mendeley data: https://data.mendeley.com/datasets/y3649cmgg6/3 (Mango and Banana Dataset (Ripe Unripe) : Indian RGB image datasets for YOLO object detection)
Initially the data was for training YOLO… See the full description on the dataset page: https://huggingface.co/datasets/darthraider/fruit-ripeness-detection-dataset.crop-burn-detection-raw
Crop Burn Detection — Raw Sentinel-2 (India, 2025)
Paired RGB + SWIR Sentinel-2 satellite image tiles across agricultural districts of northern India, capturing the paddy (Oct–Nov 2025) and wheat (Mar–May 2025) burning seasons. Built to train and benchmark vision models for real-time crop residue burn detection — including models designed to run directly on satellites.
Why We Built This
Every October and November, farmers across Punjab, Haryana, Uttar Pradesh, Rajasthan… See the full description on the dataset page: https://huggingface.co/datasets/munish0838/crop-burn-detection-raw.face-detection
Face Detection Dataset
Small grayscale face crops + a large pool of natural-image negatives, packaged
for classical face detectors (Viola-Jones, Haar cascades, sliding-window
classifiers).
At a glance
Split
Rows
Faces / Non-faces
train
106,977
102,429 / 4,548
test
24,045
472 / 23,573 (CBCL benchmark)
negatives
29,879
— / 29,879 (Caltech-256)
Same schema across all splits: image, label (0/1), source, category.
Quick start
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/salvacarrion/face-detection.xai-attack-detection-imagenette
XAI Attack Detection: Imagenette targeted BIM/PGD on ViT-B/16
Private research dataset of paired clean and targeted adversarial Imagenette images. It is
built to study how adversarial attacks change a Vision Transformer's explanation maps and to
support later work on attack detection. Each row is one source image with its clean and its
attacked version.
Summary
Pairs
12,420 (train 8,690 · validation 1,860 · test 1,870)
Source images
Imagenette v2… See the full description on the dataset page: https://huggingface.co/datasets/nimaeb/xai-attack-detection-imagenette.p3hw1-pen-detection
P3HW1 — Pen Detection (Binary) — 224×224
Purpose
Binary image classification: does the image contain a pen?
A compact dataset for course assignments and demos.
Dataset Composition
Original split (normalized to 224×224): 34 images
pen: 17
no pen: 17
Augmented split (also 224×224): 306 images
Contains the 34 resized originals plus 272 label‑preserving augmentations.
All 34 original images were student‑captured (17 pen, 17 no pen) under varied… See the full description on the dataset page: https://huggingface.co/datasets/0408happyfeet/p3hw1-pen-detection.crop-burn-detection-labeled
Crop Burn Detection — Labeled Sentinel-2 (India, 2025)
Paired RGB + SWIR Sentinel-2 satellite image tiles across agricultural districts of northern India, with per-tile burn annotations. Covers the paddy (Oct–Nov 2025) and wheat (Mar–May 2025) stubble burning seasons across 68 districts. Built to train and benchmark vision models for real-time crop residue burn detection.
The raw (unlabeled) version is available as munish0838/crop-burn-detection-raw.
Why We Built This… See the full description on the dataset page: https://huggingface.co/datasets/munish0838/crop-burn-detection-labeled.TealeafAgeQuality_detection
TeaLeafAgeQuality Detection
A dataset for detection of tea leaves. The dataset contains raw and augmented versions.
The raw split contains 2,208 images with 2,207 bounding box annotations across 4 categories.
The augmented split contains 5,740 images with 5,741 bounding box annotations across 4 categories.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{kabir2024tea,
title={Tea leaf age… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/TealeafAgeQuality_detection.deepfake-detection-dataset-v3
Deepfake Detection Dataset V3
This dataset contains images and detailed explanations for training and evaluating deepfake detection models. It includes original images, manipulated images, confidence scores, and comprehensive technical and non-technical explanations.
Dataset Structure
The dataset consists of:
Original images (image)
CAM visualization images (cam_image)
CAM overlay images (cam_overlay)
Comparison images (comparison_image)
Labels (label): Binary… See the full description on the dataset page: https://huggingface.co/datasets/KubasadNisha/deepfake-detection-dataset-v3.papaya_leaf_disease_detection
Papaya Leaf Disease Detection
A dataset for disease detection of Papaya leaves. The dataset contains 1,050 images with 7,616 bounding box annotations across 5 categories.The dataset can be used as a classification dataset based on the label column, which contains integer based labels for the following classes:Anthracnose: 0Bacterial Spot: 1Curl: 2Ring Spot: 3Healthy: 4
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/papaya_leaf_disease_detection.deepfake-detection-dataset-v3
Deepfake Detection Dataset V3
This dataset contains images and detailed explanations for training and evaluating deepfake detection models. It includes original images, manipulated images, confidence scores, and comprehensive technical and non-technical explanations.
Dataset Structure
The dataset consists of:
Original images (image)
CAM visualization images (cam_image)
CAM overlay images (cam_overlay)
Comparison images (comparison_image)
Labels (label): Binary… See the full description on the dataset page: https://huggingface.co/datasets/guglothmahipal007/deepfake-detection-dataset-v3.llm_pack_detection
LLM-Pack: Grocery Detection Dataset
A small object detection and scene understanding dataset containing tabletop grocery scenes with annotated item names and object locations.
The dataset consists of 40 images with varying object counts, designed for evaluating object detection, counting, and multimodal reasoning systems in cluttered grocery scenarios.
Dataset Overview
Total scenes: 40
Object counts per scene: 6, 8, 10, 12, 14, 16, 18, or 20 items
Samples per… See the full description on the dataset page: https://huggingface.co/datasets/Yannik019/llm_pack_detection.vegetable_crop_early_detection
Vegetable Crop Early Detection
A dataset for early stage object detection of vegetable crops. The dataset contains 2,801 images with 17,387 bounding box annotations across 6 categories.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{lac2022annotated,
title={An annotated image dataset of vegetable crops at an early stage of growth for proximal sensing applications},
author={Lac, Louis and… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/vegetable_crop_early_detection.deepfake-detection-dataset-v3
Deepfake Detection Dataset V3
This dataset contains images and detailed explanations for training and evaluating deepfake detection models. It includes original images, manipulated images, confidence scores, and comprehensive technical and non-technical explanations.
Dataset Structure
The dataset consists of:
Original images (image)
CAM visualization images (cam_image)
CAM overlay images (cam_overlay)
Comparison images (comparison_image)
Labels (label): Binary… See the full description on the dataset page: https://huggingface.co/datasets/ARPAN2026/deepfake-detection-dataset-v3.deepfake-detection-dataset-v3
Deepfake Detection Dataset V3
This dataset contains images and detailed explanations for training and evaluating deepfake detection models. It includes original images, manipulated images, confidence scores, and comprehensive technical and non-technical explanations.
Dataset Structure
The dataset consists of:
Original images (image)
CAM visualization images (cam_image)
CAM overlay images (cam_overlay)
Comparison images (comparison_image)
Labels (label): Binary… See the full description on the dataset page: https://huggingface.co/datasets/akahana/deepfake-detection-dataset-v3.deepfake-detection-dataset-v3
Deepfake Detection Dataset V3
This dataset contains images and detailed explanations for training and evaluating deepfake detection models. It includes original images, manipulated images, confidence scores, and comprehensive technical and non-technical explanations.
Dataset Structure
The dataset consists of:
Original images (image)
CAM visualization images (cam_image)
CAM overlay images (cam_overlay)
Comparison images (comparison_image)
Labels (label): Binary… See the full description on the dataset page: https://huggingface.co/datasets/arjunborra123/deepfake-detection-dataset-v3.interaction-readiness-detection-dataset-frame-path-version
Interaction Readiness Dataset (Public version)
This repository contains person-level interaction-readiness annotations derived from the AVIDAR, JRDB, and SSUP-HRI datasets.
Each row represents one target person in one video frame. The row contains the complete frame, the target person's bounding box, and the corresponding interaction-readiness label.
This public repository contains the processed interaction-readiness annotations and does not redistribute the original source… See the full description on the dataset page: https://huggingface.co/datasets/marilynchahine1/interaction-readiness-detection-dataset-frame-path-version.
