datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
industrial-technical-archive
🚀 Latest Updates (July, 2026)
Version: v07.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.industrial-faults
Luviner Industrial Fault Dataset
6,500 labeled samples across 13 classes (1 normal + 12 industrial fault types) with 8 sensor features.
Generated by the Luviner AI synthetic anomaly engine — 12 parametric failure mode generators that produce temporally progressive, physically realistic fault signatures.
Dataset Description
Why synthetic industrial data?
Real industrial failure data is extremely scarce — machines rarely fail, and when they do, the data is often… See the full description on the dataset page: https://huggingface.co/datasets/luviner/industrial-faults.vn-provinces-industrial-production-index
Vietnam provinces industrial production index
Index of industrial production (previous year = 100) by locality. Coverage 2012-2024. Year 2024 is preliminary. Socio-economic region rows are blank in the NSO source and are omitted from the publish pack. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (819 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-industrial-production-index.industrial-sensor-anomaly-data
Industrial Equipment Sensor Anomaly Data
Overview
Synthetic multivariate sensor data from a simulated manufacturing plant with 5 equipment units (EQ-001 through EQ-005). Each unit generates 10,000 one-minute-interval readings across 11 sensor channels, 2 metadata fields, 3 derived features, and equipment operating mode labels.
The dataset is designed for anomaly detection benchmarking. It embeds 4 distinct anomaly types at approximately 4.5% prevalence:
Thermal runaway —… See the full description on the dataset page: https://huggingface.co/datasets/Petsteb/industrial-sensor-anomaly-data.vn-provinces-industrial-cluster-wastewater-treatment
Vietnam industrial cluster wastewater treatment
Vietnam industrial cluster wastewater treatment. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (182 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (18 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-industrial-cluster-wastewater-treatment.forge-industrial-control-scenarios
Forge Industrial Control and Telemetry Traces
Deterministic synthetic traces spanning device ingress, signature/quality/range
failures, offline store-and-forward, local inference, sequential agent review,
L0–L4 policy outcomes, electrolyser ramp sequences and multivariate telemetry
anomalies.
No operational plant data, customer data, secrets or real equipment identifiers are included. machine.press-03 and every measurement are fictitious.
Files… See the full description on the dataset page: https://huggingface.co/datasets/sankalpsthakur/forge-industrial-control-scenarios.us-industrial-facility-intelligence-sample
US Industrial Facility Intelligence — Free Sample
This is a free 100-record sample. It is a subset of the full 1,464-record commercial dataset, provided so you can evaluate the data before deciding whether the full toolkit is useful to you.
An independent, unofficial dataset by NeuroLab Works. Not affiliated with, sponsored by, or endorsed by the U.S. EPA.
What this is
100 real, deduplicated US industrial facilities regulated under EPA's Toxics Release Inventory… See the full description on the dataset page: https://huggingface.co/datasets/NeuroLabWorks/us-industrial-facility-intelligence-sample.Synthetic_Industrial_Dataset_For_Energy_Disaggregation_SIDED
Synthetic Industrial Dataset for Energy Disaggregation (SIDED)
Dataset Summary
The Synthetic Industrial Dataset for Energy Disaggregation (SIDED) is a novel, open-source dataset created to address the critical scarcity of high-quality data for Non-Intrusive Load Monitoring (NILM) in the industrial sector. Generated using a high-fidelity Digital Twin simulator that was carefully calibrated with real-world operational data, SIDED provides a rich and diverse benchmark… See the full description on the dataset page: https://huggingface.co/datasets/CInterno/Synthetic_Industrial_Dataset_For_Energy_Disaggregation_SIDED.africa-synth-energy-industrial-energy-consumption-niger
Africa Synth Energy Industrial Energy Consumption Niger | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-industrial-energy-consumption-niger.industrial_iot_predictive_maintenance_telemetryvidore3_industrial_neomme_260m_li
vidore3_industrial_neomme_260m_li
Multi-vector (late-interaction) embeddings of ViDoRe industrial (vidore/industrial), encoded with
Hcompany/NeoMME-260M-Retriever-ST-late at revision 023be2a8ab9d797f5aa76f5bf8b5dde78d819659.
Source data: Hugging Face dataset vidore/vidore_v3_industrial at revision e26c864724f5dd71a3d7d739272d95637764cee9, configs corpus / queries / qrels, split test, loaded with datasets. Document, query and qrel ids are the source's own ids,
unchanged.
Every… See the full description on the dataset page: https://huggingface.co/datasets/robro612/vidore3_industrial_neomme_260m_li.
