datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rsna-2023-abdominal-trauma-detectionThis dataset is the preprocessed version of the dataset from RSNA 2023 Abdominal Trauma Detection Kaggle Competition.
It is tailored for segmentation and classification tasks. It contains 3 different configs as described below:
- segmentation: 206 instances where each instance includes a CT scan in NIfTI format, a segmentation mask in NIfTI format, and its relevant metadata (e.g., patient_id, series_id, incomplete_organ, aortic_hu, pixel_representation, bits_allocated, bits_stored)
- classification: 4711 instances where each instance includes a CT scan in NIfTI format, target labels (e.g., extravasation, bowel, kidney, liver, spleen, any_injury), and its relevant metadata (e.g., patient_id, series_id, incomplete_organ, aortic_hu, pixel_representation, bits_allocated, bits_stored)
- classification-with-mask: 206 instances where each instance includes a CT scan in NIfTI format, a segmentation mask in NIfTI format, target labels (e.g., extravasation, bowel, kidney, liver, spleen, any_injury), and its relevant metadata (e.g., patient_id, series_id, incomplete_organ, aortic_hu, pixel_representation, bits_allocated, bits_stored)
All CT scans and segmentation masks had already been resampled with voxel spacing (2.0, 2.0, 3.0) and thus its reduced file size.visa-anomaly-detection
VisA — Visual Anomaly Dataset
Mirror of the VisA (Visual Anomaly) dataset for research use. Staged as a proxy/pretraining
dataset for CoRe's Situational Control paint-inspection work (core-lab/situational-control).
Source
Original repo: https://github.com/amazon-science/spot-diff
Original download: https://amazon-visual-anomaly.s3.us-west-2.amazonaws.com/VisA_20220922.tar
License: CC BY 4.0 (confirmed in the source repo README)
Contents
10,821… See the full description on the dataset page: https://huggingface.co/datasets/imaadd05/visa-anomaly-detection.aidovecl-vehicle-detection-classification-localization
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models.
Citation Notice
Please ensure that all publications and presentations using this data reference the following paper:
Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.military-aircraft-detection-dataset
Military Aircraft Detection Dataset
Military aircraft detection dataset in COCO and YOLO format.
The dataset contains 103 different military aircraft types.
['A10', 'A400M', 'AG600', 'AH64', 'AKINCI', 'AV8B', 'An124', 'An22', 'An225', 'An72', 'B1', 'B2', 'B21', 'B52', 'Be200', 'C1', 'C130', 'C17', 'C2', 'C390', 'C5', 'CH47', 'CH53', 'CL415', 'E2', 'E7', 'EF2000', 'EMB314', 'F117', 'F14', 'F15', 'F16', 'F18', 'F2', 'F22', 'F35', 'F4', 'FCK1', 'H6', 'Il76', 'J10', 'J20', 'J35'… See the full description on the dataset page: https://huggingface.co/datasets/a2015003713/military-aircraft-detection-dataset.MLLM-Generated-Image-Detection-Dataset
MLLM-Generated Image Dataset
This dataset contains real and AI-generated image samples organized for binary MLLM-generated image detection.
Paper | Code
Dataset Summary
We construct an MLLM-generated image detection benchmark from GPT Image2 and Nano Banana2. This benchmark covers texture-dominated, structure-dominated, and hybrid-dominated. It is designed to evaluate detector performance under the new challenges introduced by large-scale image generation models.… See the full description on the dataset page: https://huggingface.co/datasets/zr-zhang/MLLM-Generated-Image-Detection-Dataset.AIGC-Detection-Benchmark
AIGC Detection Benchmark Dataset
📝 Dataset Description
Dataset Summary
The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.cpsc5800-hand-detection
Training Datasets and Model Weights for CPSC 5800 Final Project
Project repository: https://github.com/rohanphanse/CPSC5800-Final
We provide all training datasets created in Step 1 and weights for the YOLO and ResNet models trained during Steps 2-4 in our Hugging Face repository: https://huggingface.co/datasets/rohanphanse/cpsc5800-hand-detection
# Recommended: download dataset using git-xet (https://hf.co/docs/hub/git-xet)
brew install git-xet
git xet install
# Download datasets… See the full description on the dataset page: https://huggingface.co/datasets/rohanphanse/cpsc5800-hand-detection.license-plate-detection
Licensed Plate - Character Recognition for LPR, ALPR and ANPR
The dataset features license plates from 86 countries and includes 1,940,000+ images with OCR. It focuses on plate recognitions and related detection systems, providing detailed information on plate numbers, country, bbox labeling and other data as well as corresponding masks for recognition tasks - Get the data
The dataset encompasses plate detection systems, cameras, and character recognition for accurate… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/license-plate-detection.ear-detection-dataset
Ear Detection - 14,000+ Images
The dataset comprises 14,000+ ear images from 2,000 unique individuals, paired with reference face photos and demographic labels. Designed for ear recognition and biometric identification, it helps research in human ear detection, recognition accuracy, and biometric systems.
By leveraging this dataset, researchers can improve identification systems, train neural networks, and develop robust recognition algorithms for person identification. - Get… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/ear-detection-dataset.fall-detection-posture-classification
AI-Driven Posture Analysis & Fall Detection Dataset (Elderly Care)
Overview
This dataset supports the dissertation project "AI-Driven Posture Analysis Fall Detection System for the Elderly", completed by Patrick O. Ogbuitepu for the degree of MSc Artificial Intelligence and its Applications (CE901 MSc Project and Dissertation), School of Computer Science and Electronic Engineering (CSEE), University of Essex (2024). Supervisor: Dr Adrian Clark.
The project uses… See the full description on the dataset page: https://huggingface.co/datasets/pat2echo/fall-detection-posture-classification.nut_defect_detection
Nut Defect Classification
Synthetic Industrial Quality Inspection Dataset
Nut Defect Classification (Synthetic Dataset)
This dataset is a synthetic collection of industrial nut images designed for image classification tasks, specifically focusing on defect detection in manufacturing pipelines. It serves as a benchmark and training resource for computer vision algorithms used in quality assurance.… See the full description on the dataset page: https://huggingface.co/datasets/Kinzaaa/nut_defect_detection.face-mask-detection
😷 Face Mask Detection
853 images across 3 classes, with bounding box annotations in PASCAL VOC format — for detecting whether a person is wearing a mask, not wearing one, or wearing one incorrectly.
🧭 Overview
Masks play a crucial role in protecting individuals against respiratory diseases and were one of the key precautions against COVID-19 in the absence of immunization. This dataset enables training object detection models to classify mask usage in images.… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/face-mask-detection.detection-dataset
Detection Dataset
Dataset for detection and visual inspection tasks. Contains items with structured inspection parameters, defect verification, and multi-angle high-resolution imagery.
📊 Structure
Each entry provides:
id: Unique item identifier
title: Item category and description
gallery: URLs of high-resolution images from multiple angles
specifications: Technical parameters and attributes
reports_overview: Summary metrics and history
inspection: Detailed… See the full description on the dataset page: https://huggingface.co/datasets/Alexander123q/detection-dataset.detection-images
Hassan881/detection-images
Staging dataset for Berhan XAI (BurhanXAI) AI-generated / edited-media detection.
Field
Value
Training contract
v0.3.0 (firefly top-up applied; 6/6 M2 providers)
Held-out eval (M2)
v0.3.0-eval
Held-out eval (legacy)
v0.2.1-eval
Status
published
Classes
real, ai_generated, ai_edited
Seed
42
from datasets import load_dataset
ds = load_dataset("Hassan881/detection-images", name="v0.3.0")
ev =… See the full description on the dataset page: https://huggingface.co/datasets/Hassan881/detection-images.Military-Aircraft-DetectionDataset for object detection of military aircraft
bounding box in PASCAL VOC format (xmin, ymin, xmax, ymax)
43 aircraft types
(A-10, A-400M, AG-600, AV-8B, B-1, B-2, B-52 Be-200, C-130, C-17, C-2, C-5, E-2, E-7, EF-2000, F-117, F-14, F-15, F-16, F/A-18, F-22, F-35, F-4, J-20, JAS-39, MQ-9, Mig-31, Mirage2000, P-3(CP-140), RQ-4, Rafale, SR-71(may contain A-12), Su-34, Su-57, Tornado, Tu-160, Tu-95(Tu-142), U-2, US-2(US-1A Kai), V-22, Vulcan, XB-70, YF-23)
Please let me know if you find wrong… See the full description on the dataset page: https://huggingface.co/datasets/Illia56/Military-Aircraft-Detection.early_printed_books_font_detection
Early Printed Books Font Detection
Photographs of 35,623 pages from books printed between the mid-15th and the end of the 18th century, each labelled by experts with the font group or groups used on the page. This is a mirror of Dataset of Pages from Early Printed Books with Multiple Font Groups by Mathias Seuret, Saskia Limbach, Nikolaus Weichselbaumer, Andreas Maier and Vincent Christlein, deposited on Zenodo in August 2019 and described in their HIP'19 paper.
The page images… See the full description on the dataset page: https://huggingface.co/datasets/biglam/early_printed_books_font_detection.spark-plug-anomaly-detection
Spark Plug Anomaly Detection
Images of spark plugs for use in visual anomaly detection with Edge Impulse’s FOMO-AD learning block. The model is trained only on normal samples and flags any deviation as an anomaly.
Input: Images (96x96)
Classes: normal (train/test), anomaly (test only)
Use case: Embedded anomaly detection in predictive maintenance
Model trained and demonstrated on Edge Impulse.
Looking for multi-class condition labels?See the companion dataset: Spark Plug… See the full description on the dataset page: https://huggingface.co/datasets/eoinedge/spark-plug-anomaly-detection.ai-image-detection-dataset
AI-Image Detection Dataset
Paired real / AI images, with shared image-grounded captions, for training and
evaluating AI-generated-image detectors.
Each of 10,000 real photos is captioned once (BLIP-2) and paired with one synthetic
partner per generator (6 generators → 60,000 AI images). A real image and all of its
AI partners share the same prompt, so the only systematic difference between the
classes is the generative process itself. A detector trained here is pushed toward the… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/ai-image-detection-dataset.blood-cell-detection
Blood Cell Detection Dataset
Introduction
This dataset contains annotated red blood cells(RBC) and white blood cells(WBC) from peripheral blood smear taken from a light microscope.
If you use this dataset in your research, please cite it. See the Citation section below for the correct reference. Please do not cite the auto-generated Hugging Face reference, as it lists an incorrect title and year.
About Blood Cell Detection Dataset
Images are… See the full description on the dataset page: https://huggingface.co/datasets/draaslan/blood-cell-detection.cpsc5800-hand-detection-test
Test Datasets and Model Weights for CPSC 5800 Final Project
Project repository: https://github.com/rohanphanse/CPSC5800-Final
We provide all test datasets created in Step 1 and weights for the YOLO and ResNet models trained during Steps 2-4 in our Hugging Face repository: https://huggingface.co/datasets/rohanphanse/cpsc5800-hand-detection-test
# Recommended: download dataset using git-xet (https://hf.co/docs/hub/git-xet)
brew install git-xet
git xet install
# Download datasets and… See the full description on the dataset page: https://huggingface.co/datasets/rohanphanse/cpsc5800-hand-detection-test.Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform
Person Detection and Re-Identification from Low Altitude UAV-based Platform
Dataset Description
This dataset was collected as part of a master's thesis on person detection and re-identification using low-altitude UAV (drone) footage. It contains labeled aerial images captured from a DJI Mini drone, annotated in YOLOv8 format.
The dataset supports two tasks:
Person Detection — detecting people in aerial drone footage
Person Re-Identification (Re-ID) — recognizing and… See the full description on the dataset page: https://huggingface.co/datasets/Mikiee/Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform.weapon-detection-dataset
Weapon Detection and Classification Dataset
This repository contains a large-scale collection of weapon images and annotations for object detection and image classification tasks, covering knives, pistols, and other weapons.
Dataset Structure
The dataset is partitioned into categories, each chunked into smaller parts (max 2,500 files per part) to keep commit sizes and downloads manageable on Hugging Face:
Total Files: 63,128 (46,241 images + 16,887 annotations)… See the full description on the dataset page: https://huggingface.co/datasets/shravya11/weapon-detection-dataset.liveness-detection-dataset
Face Liveness Detection Dataset for Anti-Spoofing & PAD Certification
100,000+ spoofing videos for liveness detection
A comprehensive face liveness detection dataset for face anti-spoofing, biometric face recognition, and presentation attack detection (PAD) systems. Unlike narrow public benchmarks that cover only one or two attack types, this dataset combines all major presentation attack categories in a single resource: paper attacks, replay attacks, and 3D mask attacks (silicone… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/liveness-detection-dataset.facial-keypoint-detection
Facial Keypoint Detection Dataset
Dataset comprises 5,000+ close-up images of human faces captured against various backgrounds. It is designed to facilitate keypoints detection and improve the accuracy of facial recognition systems. The dataset includes either presumed or accurately defined keypoint positions, allowing for comprehensive analysis and training of deep learning models.
By utilizing this dataset, practitioners can explore various applications in computer vision… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/facial-keypoint-detection.road-issues-detection-dataset
Road Issues Detection Dataset
Dataset Summary
This comprehensive dataset contains 9,660 high-resolution RGB images categorized for road infrastructure issues detection. The dataset focuses on identifying critical urban infrastructure problems including potholes, damaged roads, broken road signs, illegal parking violations, and environmental cleanliness issues. It has been specifically organized and curated for computer vision and machine learning applications in smart… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/road-issues-detection-dataset.XMR_Demo_Industrial_Foreign_Object_Detection_Lentils
Demo for Hyperspectral Foreign-Object Detection in Lentils
Video spectroscopy beyond the visible spectrum, applied to foreign-object detection on a sliding lentil conveyor. Captured with a Cubert Ultris XMR camera — 61 bands per pixel, 430–910 nm, 1080 × 1000 pixels at 4 fps.
Foreign-object detection in food sorting is a general industrial-inspection problem — the rejected target could be a stone, a stem, a piece of packaging, a metal shard, or an insect. In this… See the full description on the dataset page: https://huggingface.co/datasets/cubert-gmbh/XMR_Demo_Industrial_Foreign_Object_Detection_Lentils.web-camera-face-liveness-detection
Web Camera Face Liveness Detection
The dataset consists of videos featuring individuals wearing various types of masks. Videos are recorded under different lighting conditions and with different attributes (glasses, masks, hats, hoods, wigs, and mustaches for men).
The dataset is created on the basis of iBeta Level 1 Dataset
In the dataset, there are 7 types of videos filmed on a web camera:
Silicone Mask - demonstration of a silicone mask attack (silicone)
2D mask with… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/web-camera-face-liveness-detection.sar-ship-detectionTo cite the dataset please reference it as
@INPROCEEDINGS{8124934,
author={Li, Jianwei and Qu, Changwen and Shao, Jiaqi},
booktitle={2017 SAR in Big Data Era: Models, Methods and Applications (BIGSARDATA)},
title={Ship detection in SAR images based on an improved faster R-CNN},
year={2017},
volume={},
number={},
pages={1-6},
keywords={Marine vehicles;Feature extraction;Synthetic aperture radar;Proposals;Detectors;Image resolution;Deep learning;SAR;ship detection;Faster R-CNN}… See the full description on the dataset page: https://huggingface.co/datasets/agungpambudi/sar-ship-detection.Military_Aircraft_Detection_Classification_Image_Dataset
Military Aircraft Detection & Classification Dataset
88 Classes with Advanced Background Suppression
Overview
This dataset is a professionally curated resource for training high-performance object detection and image classification models such as YOLOv11.It contains 88 distinct military aircraft classes and is explicitly designed for real-world deployment, where false positives from civilian aircraft, birds, and small drones are common.
To address this, the… See the full description on the dataset page: https://huggingface.co/datasets/Ahnuf/Military_Aircraft_Detection_Classification_Image_Dataset.xai-attack-detection-cifar10
XAI Attack Detection — CIFAR-10 PGD
This private research dataset contains balanced, paired clean and adversarial images for
studying whether an attack can be detected from a classifier explanation map.
Dataset construction
The source is the CIFAR-10 test split. A fine-tuned OpenCLIP ViT-B/16 classifies each
clean image. Clean-correct examples are attacked with untargeted L-infinity PGD using
epsilon 8/255, step size 2/255, 10 steps, and deterministic random… See the full description on the dataset page: https://huggingface.co/datasets/nimaeb/xai-attack-detection-cifar10.
