datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indoor-safety-hazard-detection-and-work-zone-monitoring
Indoor Safety Hazard Detection & Work-Zone Monitoring
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-safety-hazard-detection-and-work-zone-monitoring.Safety-helmet-datasetMM-SafetyBenchWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements.
Usage and License Notices: The dataset is intended and licensed for research use only. They are also restricted to uses that follow the license agreement GPT-4 and Stable Diffusion. The dataset is CC BY NC 4.0 (allowing only non-commercial use).
Data Source: For more information about the dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/MM-SafetyBench.Safety-helmet-datasetMMA-SafetyBenchMM-SafetyBench-plus-plus
MM-SafetyBench++
Project Page | Paper | Code
MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent.
Dataset Summary
For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.construction-safety-object-detection
Dataset Labels
['barricade', 'dumpster', 'excavators', 'gloves', 'hardhat', 'mask', 'no-hardhat', 'no-mask', 'no-safety vest', 'person', 'safety net', 'safety shoes', 'safety vest', 'dump truck', 'mini-van', 'truck', 'wheel loader']
Number of Images
{'train': 307, 'valid': 57, 'test': 34}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/construction-safety-object-detection.sigma_aldrich_safety_data-multimodalSafety-helmet-datasetsimworld-urban-transit-safety-benchMM-SafetyBenchSafety-helmet-datasetsafety-datasetSafety-helmet-datasetIITU_Safety-Helmet_Dataset_v1.0
IITU Safety-Helmet Dataset v1.0
Overview:
This dataset contains annotated images of safety helmets captured both by drone and at ground level, designed for helmet detection and color classification tasks in computer vision.
📖 Dataset Summary
This dataset contains 1,664 images annotated for safety-helmet detection and color classification.
6,473 helmet instances
Captured by drone (3–5 m, 10–15 m; angles 0°, 45°, 90°) and at ground level… See the full description on the dataset page: https://huggingface.co/datasets/ersace/IITU_Safety-Helmet_Dataset_v1.0.Airport_Drone_Threat_And_Safety_Dataset
✈️ Simuletic Airport Drone Threat & Safety Dataset
Synthetic Benchmark for Aviation Security & Rogue Drone Detection
Overview
This is an open-source synthetic dataset designed to solve a critical issue in Aviation Security: Runway Incursions by Rogue Drones.
Detecting a small drone against the complex background of a busy airport (moving planes, flashing lights, tarmac texture) is a massive challenge for standard AI. Real-world training data is nearly… See the full description on the dataset page: https://huggingface.co/datasets/Simuletic/Airport_Drone_Threat_And_Safety_Dataset.warehouse-safety-hazard-dataset
Warehouse Safety Hazard Detection Dataset
A synthetic dataset of 913 photorealistic warehouse images captured from a drone/overhead perspective, labeled across 5 safety hazard categories. Designed for training and benchmarking vision-language models (VLMs) on industrial safety inspection tasks.
Categories
Category
Train
Test
Total
Description
spill
170
43
213
Liquid spills on warehouse floors (oil, water, chemicals)
forklift_violation
108
30
138
Unsafe… See the full description on the dataset page: https://huggingface.co/datasets/Jsohal174/warehouse-safety-hazard-dataset.safetyhelmetconstruction-safety-gsnvb
Dataset Card for construction-safety-gsnvb
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
construction-safety-gsnvb
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image':… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/construction-safety-gsnvb.IITU_Safety-Helmet_Dataset_v1.0_Demo
IITU Safety-Helmet Dataset v1.0 Demo
Overview:
This dataset contains annotated images of safety helmets captured both by drone and at ground level, designed for helmet detection and color classification tasks in computer vision.
This is the DEMO version of the dataset, now it contains only 14 images and annotations.
📖 Dataset Summary
This dataset contains 1,664 images annotated for safety-helmet detection and color classification.
6,473 helmet instances… See the full description on the dataset page: https://huggingface.co/datasets/ersace/IITU_Safety-Helmet_Dataset_v1.0_Demo.bd-arsa-road-safety-visual-audit
BD-ARSA: Road Safety Visual Audit Dataset
A multi-task vision-language dataset for visual road-safety auditing in Bangladesh,
following the LGED (Local Government Engineering Department) audit methodology. Each record
pairs a road image with a structured safety audit. All 12 LGED hazard categories were
assessed visually in the field by the expert auditors and all 12 appear in the schema and
in evaluation. For two of them — skid_resistance (surface friction) and drainage… See the full description on the dataset page: https://huggingface.co/datasets/Thamed-Chowdhury/bd-arsa-road-safety-visual-audit.sem-seg-country-safety-binstool-safety-dataset
Tool Safety Dataset
Dataset Description
The Tool Safety Dataset is a specialized collection of tool images with detailed safety and usage information. It combines visual data with comprehensive metadata about various hand tools, making it valuable for both computer vision tasks and safety training applications.
Dataset Summary
Type: Image dataset with bounding boxes and detailed tool information
Size: Multiple splits (train/test/validation)
Format: Images with… See the full description on the dataset page: https://huggingface.co/datasets/akameswa/tool-safety-dataset.road-safety-vulnerable-road-users
Road Safety & Vulnerable Road Users Visual Dataset
Rows: 484
Dataset Description
Road Safety & Vulnerable Road Users Visual Dataset is a Global street-level imagery and geospatial computer vision dataset for image classification, visual search, and mobility research. The labeled features in the dataset are Crosswalk, Cyclist, School Bus, Traffic Cone, and Wheelchair, and each row records a matched visual feature label derived from the Outerview internal visual… See the full description on the dataset page: https://huggingface.co/datasets/Outerview/road-safety-vulnerable-road-users.safety-quant-phase0
Safety-Aware Configuration-Conditioned LoRA — Phase 0 baseline
Baseline table, frozen evaluation sets, and per-prompt judge verdicts for
meta-llama/Llama-3.2-1B-Instruct under five quantization configurations.
id
scheme
native?
c0
W16A16 bf16 (reference)
yes
c1
W8A8 (SmoothQuant + GPTQ)
yes
c2
W4A16 g128 (GPTQ)
yes
c3
W4A16 g128 (AWQ)
yes
c4
NF4 (bitsandbytes)
yes
c1sim
W8A16 g128 (GPTQ)
simulated
c2sim
W4A16 g128 (GPTQ)
simulated — the simulation control… See the full description on the dataset page: https://huggingface.co/datasets/Jeesup/safety-quant-phase0.vlm-safety-inspector-dataset
Safety Inspector V2 LoRA Training Dataset & Hyperparameter Specification
This dataset repository contains the offline warm-up SFT dataset and standardized LoRA training configuration for training the Vision-Language Model (VLM) Safety Inspector on the 50 tabletop manipulation scenes (Split into 45 Train + 5 Validation).
1. Dataset Overview
Source Scenes: 50 Tabletop Scenes (45 Train, 5 Val, 0 Test)
Task Levels: single_step (1 primitive), safety_two_step (2… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/vlm-safety-inspector-dataset.ambiguity-filtered-mm-safetybenchlibero_safety_v1
LIBERO Safety
This repository contains 50-scene synthetic LIBERO safety validation sets.
v1: original LIBERO safety image set.
v2: regenerated set using the latest scene configuration and saved manual positions.
v3: same scenes and labels as v2, with a static non-colliding MuJoCo white cutting board fixture added on the table; original V2 object and fixture states are unchanged.
v4: corrected zoomed non-reference set with pickable target objects for video rollouts.
v5: corrected… See the full description on the dataset page: https://huggingface.co/datasets/saaduddinM/libero_safety_v1.sem-seg-safety-binsafrica-synth-herbal-traditional-medicine-safety-all
Herbal & Traditional Medicine Safety (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-herbal-traditional-medicine-safety-all.
