datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cocoKAIST-Multispectral-Pedestrian-Detection-Datasettraffic-vehicle-detection
Edge-AI Traffic Vehicle Detection (UA-DETRAC CCTV)
Part of the Edge-AI Traffic & Vehicle Analytics System repository by thundarstrom.
Dataset Summary
Curated and normalized 23,319 CCTV traffic images from fixed intersection surveillance cameras (UA-DETRAC benchmark). Contains 215,109 annotated bounding boxes in standard YOLO format across 4 vehicle classes: car, bus, truck, and van.
Class Mapping
Class 0 (car): 177,403 bboxes (82.5%)
Class 1 (bus):… See the full description on the dataset page: https://huggingface.co/datasets/PRAS4NTH/traffic-vehicle-detection.Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/Abtinzandi/Obstacle-Detection-Dataset-YOLO.ritvij-saxena-iris-detection-pythonHere is the IRIS dataset for the project iris-detection-python.
Official Statement
I hereby declare that I do not own the rights to the dataset used in this project. This dataset was provided by the faculty and utilized solely for educational purposes as part of an assignment for the Biometrics course (CS 559) at the Illinois Institute of Technology.
The dataset is provided for academic and research purposes only, and I encourage others to use it responsibly for similar educational… See the full description on the dataset page: https://huggingface.co/datasets/saxenaritvij/ritvij-saxena-iris-detection-python.mvtec_anomaly_detectioncsgo-player-detection
CS2 player detection - native 640x640 centre crops
5473 labelled frames from Counter-Strike 2 with 10643 boxes for
player bodies and heads, split by team.
A YOLO26 detector trained on this data is at
fvossel/csgo-player-detection. Its training set is not quite
identical: it also contains the images from the source dataset named below, which
are not redistributed here.
Everything that produced this data is open: capture, demo rendering, the labelling
loop and the training and… See the full description on the dataset page: https://huggingface.co/datasets/fvossel/csgo-player-detection.visa-anomaly-detection
VisA — Visual Anomaly Dataset
Mirror of the VisA (Visual Anomaly) dataset for research use. Staged as a proxy/pretraining
dataset for CoRe's Situational Control paint-inspection work (core-lab/situational-control).
Source
Original repo: https://github.com/amazon-science/spot-diff
Original download: https://amazon-visual-anomaly.s3.us-west-2.amazonaws.com/VisA_20220922.tar
License: CC BY 4.0 (confirmed in the source repo README)
Contents
10,821… See the full description on the dataset page: https://huggingface.co/datasets/imaadd05/visa-anomaly-detection.indoor-safety-hazard-detection-and-work-zone-monitoring
Indoor Safety Hazard Detection & Work-Zone Monitoring
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-safety-hazard-detection-and-work-zone-monitoring.aidovecl-vehicle-detection-classification-localization
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models.
Citation Notice
Please ensure that all publications and presentations using this data reference the following paper:
Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.military-aircraft-detection-dataset
Military Aircraft Detection Dataset
Military aircraft detection dataset in COCO and YOLO format.
The dataset contains 103 different military aircraft types.
['A10', 'A400M', 'AG600', 'AH64', 'AKINCI', 'AV8B', 'An124', 'An22', 'An225', 'An72', 'B1', 'B2', 'B21', 'B52', 'Be200', 'C1', 'C130', 'C17', 'C2', 'C390', 'C5', 'CH47', 'CH53', 'CL415', 'E2', 'E7', 'EF2000', 'EMB314', 'F117', 'F14', 'F15', 'F16', 'F18', 'F2', 'F22', 'F35', 'F4', 'FCK1', 'H6', 'Il76', 'J10', 'J20', 'J35'… See the full description on the dataset page: https://huggingface.co/datasets/a2015003713/military-aircraft-detection-dataset.hard-hat-detection
Dataset Card for hard-hat-detection
This dataset, contains 5000 images with bounding box annotations in the PASCAL VOC format for these 3 classes:
Helmet
Person
Head
This is a FiftyOne dataset with 5000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'split', 'max_samples', etc… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/hard-hat-detection.MLLM-Generated-Image-Detection-Dataset
MLLM-Generated Image Dataset
This dataset contains real and AI-generated image samples organized for binary MLLM-generated image detection.
Paper | Code
Dataset Summary
We construct an MLLM-generated image detection benchmark from GPT Image2 and Nano Banana2. This benchmark covers texture-dominated, structure-dominated, and hybrid-dominated. It is designed to evaluate detector performance under the new challenges introduced by large-scale image generation models.… See the full description on the dataset page: https://huggingface.co/datasets/zr-zhang/MLLM-Generated-Image-Detection-Dataset.bedroom-change-detection-state-monitoring
Bedroom Change Detection & State Monitoring
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by the sync… See the full description on the dataset page: https://huggingface.co/datasets/physicl/bedroom-change-detection-state-monitoring.haris-weapon-detection-dataset-curatedsss-crab-pot-detection-ds
🦀 Ghost Pot Side‑Scan Sonar Detection Dataset
Side‑Scan Sonar Imagery & Annotations for Derelict Crab Pot Detection
This dataset contains manually annotated side‑scan sonar (SSS) imagery collected across Delaware’s Inland Bays and Delaware Bay to support research on automated detection of derelict crab pots (“ghost pots”). It is designed for training and evaluating object‑detection models in turbid, shallow‑water environments where visual surveys are limited and acoustic… See the full description on the dataset page: https://huggingface.co/datasets/PINGEcosystem/sss-crab-pot-detection-ds.Drone-DetectionDartboard-Detection-Dataset
Dartboard Detection Dataset
A curated dartboard image dataset for computer vision tasks such as detection, recognition, localization, and model training.
This dataset is used in my dartboard AI projects built with Rust and PyTorch. Anyone can use this dataset to train, test, or improve their own models for dartboard-related computer vision tasks.
About
This dataset contains cropped dartboard images organized in folders by capture sessions and dates. It is intended for… See the full description on the dataset page: https://huggingface.co/datasets/bhabha-kapil/Dartboard-Detection-Dataset.weapon-detection-runs-backup-2026-07-24CCTV-Smoke-Fire-Emergency-Detection-Dataset
CCTV Smoke & Fire Emergency Detection Dataset
Early-stage fire detection dataset featuring small ignition points, bin fires, and smoldering debris from a surveillance perspective.
🧐 Overview
CCTV Smoke & Fire is a specialized open-source synthetic dataset for Computer Vision (CV) tasks focused on Emergency Response, Smart City Safety, and Incident Monitoring.
The most critical fires are the ones detected in their first 60 seconds. While most fire datasets… See the full description on the dataset page: https://huggingface.co/datasets/Simuletic/CCTV-Smoke-Fire-Emergency-Detection-Dataset.drone-detectionweapon-detection-workerssd-backup-2026-07-24airport-security-detectionindoor-anomaly-detection-path-obstruction-monitoring
Indoor Anomaly Detection & Path Obstruction Monitoring
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-anomaly-detection-path-obstruction-monitoring.traffic-vehicle-detection
Edge-AI Traffic Vehicle Detection (UA-DETRAC CCTV)
Part of the Edge-AI Traffic & Vehicle Analytics System repository by thundarstrom.
Dataset Summary
Curated and normalized 23,319 CCTV traffic images from fixed intersection surveillance cameras (UA-DETRAC benchmark). Contains 215,109 annotated bounding boxes in standard YOLO format across 4 vehicle classes: car, bus, truck, and van.
Class Mapping
Class 0 (car): 177,403 bboxes (82.5%)
Class 1 (bus):… See the full description on the dataset page: https://huggingface.co/datasets/thundarstrom/traffic-vehicle-detection.fire-smoke-detection-corpus-v1
FireViewer Fire/Smoke Detection Corpus v1
Status
Active strict-clean detection corpus. Current catalogue state: 102,257 rows, split 60,981 train / 19,209 validation / 22,067 test.
The corpus stores source-specific provenance, hashes, grouping/de-duplication information, validation status and annotation metadata. It is the current training reference for the strict FireViewer detector releases.
Rights
There is no single licence covering all source… See the full description on the dataset page: https://huggingface.co/datasets/fireviewer/fire-smoke-detection-corpus-v1.ppe-detectionwelding-defect-object-detection
Welding Defect Object Detection
2,028 annotated images of welds for defect detection, in both YOLO and COCO
formats. Three classes:
id (YOLO / COCO)
name
0 / 1
Bad Weld
1 / 2
Good Weld
2 / 3
Defect
Splits
split
images
annotations
train
1,619
4,583
valid
283
802
test
126
301
Layout
├── data.yaml # YOLO class names + split paths
├── train|valid|test/
│ ├── images/ # .jpg
│ └── labels/… See the full description on the dataset page: https://huggingface.co/datasets/rikkarth/welding-defect-object-detection.synthetic-driver-monitoring-detection
Synthetic DMS – Driver Monitoring System Dataset by AnywayLabs.ai
Need a custom synthetic dataset for your own road safety detection use case?
This dataset is an open-source sample of our synthetic data generation work at AnywayLabs.
If you're working on:
industrial defect detection
visual inspection
supervised anomaly detection
hard-to-collect defect classes
synthetic data for computer vision training
You can request a custom synthetic dataset here, or email:… See the full description on the dataset page: https://huggingface.co/datasets/1841819173liu/synthetic-driver-monitoring-detection.PPE_Detection
