datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndustryBench-MIPU
IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images.
Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.IndustrialDetectionStaticCamerasThe IndustrialDetectionStaticCameras dataset has been collected in order to validate the methodology presented in the paper entitled A few-shot learning methodology for improving safety in industrial scenarios through universal self-supervised visual features and dense optical flow. This dataset is divided into five main folders named videoY, where Y=1,2,3,4,5. Each videoY folder contains the following:
The video of the scene in .mp4 format: videoY.mp4
A folder with the images of each frame… See the full description on the dataset page: https://huggingface.co/datasets/jjldo21/IndustrialDetectionStaticCameras.vidore_v3_industrialViDoRe V3 : Industrial reports
This dataset, Industrial reports, is a corpus of technical documents on military aircrafts (fueling, mechanics...), intended for complex-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_industrial.vidore_v3_industrial_mteb_format
Vidore3IndustrialRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_industrial
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("Vidore3IndustrialRetrieval")… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_industrial_mteb_format.industrial-instruction-dataset
Industrial-Instruction Dataset
Industrial-Instruction provides benchmark and training-ready QA instances derived from industrial technical reports, designed to evaluate robustness under realistic retrieval conditions. Samples are grounded in retrieved evidence and include irrelevant retrieval, single-/multi-document support, and single-/multi-document answer settings.
Paper
Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Parssky/industrial-instruction-dataset.sangyo_no_yume_industrial_dreams
From the Frontier Research Team at Takara.ai we present the "Sangyo no Yume Industrial Dreams" dataset, a collection of AI-generated industrial dreamscapes.
Sangyo no Yume Industrial Dreams
Dataset Details
"Sangyo no Yume Industrial Dreams" is a collection of images generated using SDXL Lightning with specialized prompt engineering techniques. These images balance industrial themes with dreamlike qualities, creating a unique aesthetic that sits at the intersection of… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/sangyo_no_yume_industrial_dreams.industrial_cartrocochallenge2026_Industrial_Assembly_annotatedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"observation.images.head": {
"dtype": "video",
"shape": [
240,
320,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/nikodembartnik/rocochallenge2026_Industrial_Assembly_annotated.Industrial-Workplace-Egocentric-FHD-Samples
Industrial & Workplace Egocentric Video — FHD Samples
21 clips of first-person video of real industrial and workplace tasks, captured with a head-mounted smartphone. Released by TrainThemAI for training Vision-Language-Action (VLA) models, World Action Models (WAM), and humanoid manipulation policies — π0, π1, OpenVLA, RT-2, GR00T, Cosmos, DreamZero.
Fully rights-cleared, MIT-licensed, and representative of our production capture pipeline.
Companion to POV Egocentric Video —… See the full description on the dataset page: https://huggingface.co/datasets/TrainThemAI/Industrial-Workplace-Egocentric-FHD-Samples.industrial-technical-archive
🚀 Latest Updates (July, 2026)
Version: v07.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.industrial-maintenance-synthetic
Industrial Maintenance Synthetic Dataset
Synthetic dataset of 1M sensor tag records and 1M maintenance work orders across 37 industrial equipment types, generated for domain-specific NLP models in industrial maintenance.
Dataset Details
Property
Value
Total rows
~2M (1M sensor tags + 1M maintenance records)
Equipment types
37
Equipment instances
145
Languages
English
Format
Parquet
With impurities
Yes (10% rate)
Build date
2026-03-07
License… See the full description on the dataset page: https://huggingface.co/datasets/Jvachier/industrial-maintenance-synthetic.IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper
IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension
Details
Total questions
2,049
Languages
Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.IndustrialLateralLoadsThe IndustrialLateralLoads dataset is designed for object detection and instance segmentation tasks in industrial environments. It contains images of palletized loads with their corresponding annotations. The dataset is available in two formats:
Hugging Face dataset (Parquet): Ready-to-use format with images, masks, and metadata.
Raw files: Original folders accessible in the repository files.
Hugging Face dataset features:
When loaded using the datasets library, each sample… See the full description on the dataset page: https://huggingface.co/datasets/jjldo21/IndustrialLateralLoads.industrial-rebar-metric-depth
Industrial Rebar Metric Depth
Evaluation data from:
Benchmarking Metric Depth for Construction Perception: Industrial Rebar Scenes with Exact Synthetic Ground TruthSalman Hajizada, Diram Tabaa, Gianni A. Di CaroIEEE IROS 2026 Workshop on the Future of Construction (FoC)
Three synthetic rebar scenes rendered in Blender 4.5 with exact optical-axis depth and structure masks, and the physical ZED sequence of a 3D-printed lattice. These are the sequences behind Table I and Table II… See the full description on the dataset page: https://huggingface.co/datasets/sal0-h/industrial-rebar-metric-depth.industrial-faults
Luviner Industrial Fault Dataset
6,500 labeled samples across 13 classes (1 normal + 12 industrial fault types) with 8 sensor features.
Generated by the Luviner AI synthetic anomaly engine — 12 parametric failure mode generators that produce temporally progressive, physically realistic fault signatures.
Dataset Description
Why synthetic industrial data?
Real industrial failure data is extremely scarce — machines rarely fail, and when they do, the data is often… See the full description on the dataset page: https://huggingface.co/datasets/luviner/industrial-faults.vn-provinces-industrial-production-index
Vietnam provinces industrial production index
Index of industrial production (previous year = 100) by locality. Coverage 2012-2024. Year 2024 is preliminary. Socio-economic region rows are blank in the NSO source and are omitted from the publish pack. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (819 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-industrial-production-index.industrial-scada-plc-automation-2026
⚙️ Industrial SCADA & PLC Automation (IEC 61131-3) SFT/DPO Suite
Frontier synthetic alignment dataset engineered for fine-tuning Large Language Models on mission-critical Industrial Automation, Siemens S7 SCL, Rockwell Studio 5000 ST, Beckhoff TwinCAT 3, Schneider M580, IEC 61508 SIL-3 Safety Systems, and SCADA fieldbus telemetry.
📊 Empirical Fine-Tuning Benchmark Delta Matrix
Evaluation Benchmark / Stress Dimension
Base Foundation Model… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/industrial-scada-plc-automation-2026.synthetic-industrial-defects-teaser
Synthetic Industrial Defects — Teaser
Dataset Description
1000 synthetic images of industrial metal surfaces with 5 defect types:
scratch, dent, crack, corrosion, discoloration
Format
Images: PNG (1024x768)
Annotations: YOLO format (.txt)
Classes: 5
Usage
from datasets import load_dataset
ds = load_dataset("xanoutas/synthetic-industrial-defects-teaser")
Full Dataset
Complete dataset (5000+ images) available at:… See the full description on the dataset page: https://huggingface.co/datasets/xanoutas/synthetic-industrial-defects-teaser.industrial-asset-level-electrical-energy-dataset
Open Energy Dataset — Star Schema
1. Overview
This star schema models the Gold-layer 15-minute energy aggregates from an industrial manufacturing facility in Ireland. The source dataset covers 43 monitored assets over ~12 months (2024-12-31 to 2025-12-31), with 1,039,873 ALL-phase windows totalling 2.96 GWh of measured electrical energy.
The facility employs ~150 personnel under continuous production. Monitored loads include hydraulic presses, air compressors… See the full description on the dataset page: https://huggingface.co/datasets/renumics/industrial-asset-level-electrical-energy-dataset.industrial-sensor-anomaly-data
Industrial Equipment Sensor Anomaly Data
Overview
Synthetic multivariate sensor data from a simulated manufacturing plant with 5 equipment units (EQ-001 through EQ-005). Each unit generates 10,000 one-minute-interval readings across 11 sensor channels, 2 metadata fields, 3 derived features, and equipment operating mode labels.
The dataset is designed for anomaly detection benchmarking. It embeds 4 distinct anomaly types at approximately 4.5% prevalence:
Thermal runaway —… See the full description on the dataset page: https://huggingface.co/datasets/Petsteb/industrial-sensor-anomaly-data.vn-provinces-industrial-cluster-wastewater-treatment
Vietnam industrial cluster wastewater treatment
Vietnam industrial cluster wastewater treatment. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (182 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (18 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-industrial-cluster-wastewater-treatment.forge-industrial-control-scenarios
Forge Industrial Control and Telemetry Traces
Deterministic synthetic traces spanning device ingress, signature/quality/range
failures, offline store-and-forward, local inference, sequential agent review,
L0–L4 policy outcomes, electrolyser ramp sequences and multivariate telemetry
anomalies.
No operational plant data, customer data, secrets or real equipment identifiers are included. machine.press-03 and every measurement are fictitious.
Files… See the full description on the dataset page: https://huggingface.co/datasets/sankalpsthakur/forge-industrial-control-scenarios.africa-egypt-capmas-industrial-production-of-private-sector-establishments-5af432a9
Industrial Production of Private Sector Establishments | Africa (CAPMAS Egypt Open Data)
4,045 rows - 1 Africa country/area - 2010-2021 - 58 indicators - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 4,045 rows from CAPMAS Egypt Open Data, covering Industrial Production of Private Sector Establishments. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-egypt-capmas-industrial-production-of-private-sector-establishments-5af432a9.aegis-bilingual-industrial-ai-dataset
AEGIS AI Bilingual Industrial Operations Dataset
AEGIS AI Bilingual Industrial Operations Dataset is a synthetic English–Arabic dataset designed for experimentation with multilingual enterprise AI systems, Retrieval-Augmented Generation (RAG), industrial question answering, document intelligence, semantic search, and AI workflow automation.
The dataset extends the original AEGIS AI industrial dataset with structured Arabic and English representations while preserving industrial… See the full description on the dataset page: https://huggingface.co/datasets/syed7741/aegis-bilingual-industrial-ai-dataset.aegis-industrial-ai-dataset
AEGIS Industrial AI Dataset
A synthetic industrial operations dataset created for AEGIS AI, an end-to-end industrial artificial intelligence platform.
The dataset is designed to support experimentation with:
Retrieval-Augmented Generation (RAG)
Semantic Search
Industrial AI Assistants
Document Intelligence
Predictive Maintenance
Worker Safety
Robot Monitoring
Computer Vision / Inspection
AI Alerts
Workflow Automation
Project Overview
AEGIS AI is an industrial… See the full description on the dataset page: https://huggingface.co/datasets/syed7741/aegis-industrial-ai-dataset.forge-industrial-pack
Forge Industrial Intelligence Pack (Sample)
A synthetic industrial real estate, logistics, and operations decision-telemetry dataset for anomaly detection, forecasting, decision-support, and policy-optimization research. Each row is a complete operational scenario lifecycle — from signal emergence through telemetry evolution, detection, forecast, impact, and recommended action — with 14 top-level context fields tying market, tenant, facility, portfolio, and decision-owner layers… See the full description on the dataset page: https://huggingface.co/datasets/solsticestudioai/forge-industrial-pack.industrial-agent-benchmark
Industrial Agent Benchmark
Industrial Agent Benchmark (IAB) is an open benchmark for evaluating Industrial AI systems, Manufacturing AI assistants, and Industrial Agents.
This Dataset Card describes the Hugging Face Dataset release for Industrial Agent Benchmark v2.2.0 Japanese Canonical Normalization.
Repository:
https://github.com/masahirosakae/industrial-agent-benchmark
Hugging Face Dataset Repository:
https://huggingface.co/datasets/MSakae/industrial-agent-benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MSakae/industrial-agent-benchmark.us-industrial-facility-intelligence-sample
US Industrial Facility Intelligence — Free Sample
This is a free 100-record sample. It is a subset of the full 1,464-record commercial dataset, provided so you can evaluate the data before deciding whether the full toolkit is useful to you.
An independent, unofficial dataset by NeuroLab Works. Not affiliated with, sponsored by, or endorsed by the U.S. EPA.
What this is
100 real, deduplicated US industrial facilities regulated under EPA's Toxics Release Inventory… See the full description on the dataset page: https://huggingface.co/datasets/NeuroLabWorks/us-industrial-facility-intelligence-sample.IndustrialIndustrials_News_smr
