datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
industrial-instruction-dataset
Industrial-Instruction Dataset
Industrial-Instruction provides benchmark and training-ready QA instances derived from industrial technical reports, designed to evaluate robustness under realistic retrieval conditions. Samples are grounded in retrieved evidence and include irrelevant retrieval, single-/multi-document support, and single-/multi-document answer settings.
Paper
Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Parssky/industrial-instruction-dataset.IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper
IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension
Details
Total questions
2,049
Languages
Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.industrial-scada-plc-automation-2026
⚙️ Industrial SCADA & PLC Automation (IEC 61131-3) SFT/DPO Suite
Frontier synthetic alignment dataset engineered for fine-tuning Large Language Models on mission-critical Industrial Automation, Siemens S7 SCL, Rockwell Studio 5000 ST, Beckhoff TwinCAT 3, Schneider M580, IEC 61508 SIL-3 Safety Systems, and SCADA fieldbus telemetry.
📊 Empirical Fine-Tuning Benchmark Delta Matrix
Evaluation Benchmark / Stress Dimension
Base Foundation Model… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/industrial-scada-plc-automation-2026.aegis-bilingual-industrial-ai-dataset
AEGIS AI Bilingual Industrial Operations Dataset
AEGIS AI Bilingual Industrial Operations Dataset is a synthetic English–Arabic dataset designed for experimentation with multilingual enterprise AI systems, Retrieval-Augmented Generation (RAG), industrial question answering, document intelligence, semantic search, and AI workflow automation.
The dataset extends the original AEGIS AI industrial dataset with structured Arabic and English representations while preserving industrial… See the full description on the dataset page: https://huggingface.co/datasets/syed7741/aegis-bilingual-industrial-ai-dataset.Edge-Industrial-Anomaly-Phi3
Edge-Industrial-Anomaly-Phi3: A Curated Dataset for SLMs
This dataset is a curated collection of industrial sensor data formatted specifically for Small Language Models (SLMs) like Phi-3. It merges three high-value industrial domains into a unified "Natural Language Reasoning" format to move beyond simple binary classification.
🚀 Purpose
Standard anomaly detection uses CSVs and Scikit-Learn. This dataset enables Generative Anomaly Detection, where a model like Phi-3 can… See the full description on the dataset page: https://huggingface.co/datasets/ssam17/Edge-Industrial-Anomaly-Phi3.
