datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nuclear-intelligence-dataset
Nuclear Intelligence Dataset
Public, auto-generated dataset of validated nuclear-energy research cycles.
Latest stats (auto-updated):
🪙 NES tokens minted: 0
⛓️ Blockchain length: 1 blocks
🕸️ Knowledge entities: 2
Source
GitHub: https://github.com/QalamHipHop/nuclear-intelligence
HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence
License
MIT
threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/reloading0101/threat-intelligence-dataset.chinese-materials-science-open-intelligence
🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset
Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.chinese-clean-energy-battery-open-intelligence
🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset
Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.chinese-ai-and-robotics-open-intelligence
🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset
Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.chinese-biomedicine-and-genomics-open-intelligence
🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset
Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.IndustryInstruction_Artificial-Intelligence
IndustryInstruction: Artificial Intelligence
This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.Arc-ATLAS-Teach-v1
Arc-ATLAS-Teach
Summary
This revision bundles 624 high-quality adaptive teaching examples that were generated and validated with the latest five-pass pipeline. Every dialogue walks through the full instructional arc—probe, draft plan, checkpoint feedback, revised plan, and final solution—so the teaching policy observes the complete adjustment process without ever seeing the canonical answer. Probe turns capture the student’s diagnostic attempt, teacher plans and… See the full description on the dataset page: https://huggingface.co/datasets/Arc-Intelligence/Arc-ATLAS-Teach-v1.unified-vulnerability-intelligence-dataset
Unified Vulnerability Intelligence Dataset (UVID) v3.0 — Cyber Security Knowledge Graph
UVID is a structured cyber security knowledge graph that unifies multiple
vulnerability classification frameworks into a single knowledge base. Each of the
250 records describes one application/software security vulnerability and links
it — where authoritative data exists — across CWE, CAPEC, MITRE ATT&CK, CVSS,
14 OWASP projects, secure-fix intelligence, detection surfaces, programming… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/unified-vulnerability-intelligence-dataset.ICBCBenchICBCBench: An Industry Consortium Benchmark for Financial Deep Research
Overview
ICBCBench is an industry consortium benchmark for evaluating financial Deep Research Agents in real-world research scenarios. It consists of bilingual objective and subjective tasks across major financial sectors, including capital markets, banking, insurance, and related financial services. Developed with over 50 contributors from more than 40 financial and academic organizations, ICBCBench… See the full description on the dataset page: https://huggingface.co/datasets/DeepFin-Intelligence/ICBCBench.tau2-mms-teacher-traces
Tau2 Teacher Traces Dataset
Dataset Description
This dataset contains teacher reasoning traces for solving MMS (Multimedia Messaging Service) issues in the τ²-bench (Tau2-bench) framework. Each example includes a teacher model's thinking process and structured teaching guidance for resolving customer service tickets.
Dataset Summary
Domain: Telecom customer service
Task: MMS troubleshooting
Size: 49 examples
Format: JSONL
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Arc-Intelligence/tau2-mms-teacher-traces.Medical_Intelligence_Dataset_76k_2026_Edition
🏥 Medical Intelligence Dataset · 76k · 2026 Edition
Production-ready medical AI dataset for training diagnosis, treatment reasoning, and doctor-patient conversational systems.
76,000 engineered (not collected) English Q&A pairs — covering 620+ diseases,
438+ FDA-approved drugs, and real patient-doctor conversations. Built with a
5-stage quality pipeline. Commercial-safe Apache 2.0.
Created by Huzefa Nalkheda Wala — AI Product
Engineer & Medical AI Researcher · Creator of the… See the full description on the dataset page: https://huggingface.co/datasets/huzaifa525/Medical_Intelligence_Dataset_76k_2026_Edition.threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/tonygarg/threat-intelligence-dataset.threat-intelligence-dataset-archive
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/threat-intelligence-dataset-archive.verified-math-reasoning-3k
HSH Verified Math Reasoning — Fine-Tuning Ready
A clean, answer-verified dataset of step-by-step math word problems with chain-of-thought reasoning, formatted for instruction fine-tuning. This is foundational reasoning data designed for first fine-tunes — single-concept arithmetic word problems with fully verified answers, ideal for a reliable, clean starter run. Every single answer in this dataset has been programmatically verified against a ground-truth value computed in… See the full description on the dataset page: https://huggingface.co/datasets/HSH-Intelligence/verified-math-reasoning-3k.cyber-threat-intelligence
Cyber Threat Intelligence Dataset
A comprehensive cybersecurity dataset combining CVE vulnerability data, MITRE ATT&CK techniques, and CISA Known Exploited Vulnerabilities — structured for AI/ML training and security research.
Author: Soham DahivalkarLicense: MITCreated: 2026
Dataset Description
This dataset provides structured cybersecurity intelligence data collected from three authoritative public sources:
NVD (National Vulnerability Database) — CVE… See the full description on the dataset page: https://huggingface.co/datasets/Shomi28/cyber-threat-intelligence.threat-intelligence
Comprehensive Threat Intelligence Dataset
Dataset Description
This comprehensive bilingual (French/English) threat intelligence dataset contains detailed information about Indicators of Compromise (IoCs), Tactics, Techniques, and Procedures (TTPs), APT groups, malware families, and threat hunting queries. The dataset is designed for training security analysts, threat hunters, and AI models focused on cybersecurity.
Dataset Summary
Languages: English (en)… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/threat-intelligence.Sindhi-Intelligence-Core-SFT
🧠 Sindhi Intelligence Core SFT
This is a premium, high-density instruction dataset designed for training Large Language Models (LLMs) to master the Sindhi language. With 361,225 rows, it provides a robust foundation for grammar, factual knowledge, and logical reasoning.
📊 Dataset Summary
This dataset was created by consolidating multiple high-quality Sindhi corpora into a unified ChatML format. It is specifically optimized for Supervised Fine-Tuning (SFT).
📁… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/Sindhi-Intelligence-Core-SFT.Arabic-gsm8k-v2
Dataset Summary
Arabic GSM8K is an Arabic translation of the GSM8K (Grade School Math 8K) dataset, which contains high-quality linguistically diverse grade school math word problems. The original dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning, and this Arabic version aims to extend these capabilities to Arabic language models and applications.
The dataset maintains the same characteristics as the… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-gsm8k-v2.Arabic_Openai_MMMLU
Arabic Multilingual Massive Multitask Language Understanding (MMMLU)
The MMLU is a widely recognized benchmark for assessing general knowledge attained by AI models. It covers a broad range of topics across 57 different categories, from elementary-level knowledge to advanced professional subjects like law, physics, history, and computer science.
We have extracted the Arabic subset from the MMMLU test set, which was translated by professional human translators. This dataset, now… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic_Openai_MMMLU.sample-fusion-intelligence-traces
Sample Fusion Intelligence Traces
Structured AI reasoning traces from dFusion's Fusion Intelligence system. Each record captures a complete agentic workflow: a real user query on a domain-specific topic, the full message chain including system prompts, tool calls, search results, intermediate reasoning steps, and a final synthesized answer — along with human feedback.
These are not synthetic benchmarks. They are traces from real queries submitted by real users on live financial… See the full description on the dataset page: https://huggingface.co/datasets/dFusionAILabs/sample-fusion-intelligence-traces.intrinsic-intelligence-foundations
🌿 Intrinsic Intelligence Foundations
Toward truly autonomous and benevolent intelligence — beyond externally imposed objectives.
Intrinsic Intelligence Foundations is a structured, math-aware JSONL corpus built from K. Takahashi’s theoretical preprints (Fractal Category Theory / PF–UGV / “no-meta” autonomy line).It is designed to help LLMs understand mathematical structure, category-theoretic formalisms, and equation-level reasoning, while exposing an explicit architecture… See the full description on the dataset page: https://huggingface.co/datasets/kadubon/intrinsic-intelligence-foundations.DoD-Instruction-5400-19-Public-Affairs-Use-of-Artificial-Intelligence
DoD Public Affairs Use of Artificial Intelligence
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains 150 document-grounded question-and-answer records based on DoD Instruction 5400.19, “Public Affairs Use of Artificial Intelligence,” effective July 28, 2025.
The source establishes Department of Defense policy, responsibilities, and procedures for the appropriate use of artificial-intelligence capabilities in… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-5400-19-Public-Affairs-Use-of-Artificial-Intelligence.UrduMedIQ-Urdu-Medical-Intelligence-QuestionsBased on the document you've shared, here's the content formatted in Markdown:
# UrduMed-Grok3-Verifiable Dataset
A high-quality **Urdu-language medical reasoning dataset** distilled from [Medical-O1-Verifiable-Problem]
using **Grok-3**, designed to enhance Urdu LLMs' clinical reasoning and diagnostic accuracy.
---
## 🔍 Overview
This dataset adapts challenging, open-ended medical problems from English to Urdu while preserving:
- **Medical accuracy** (via Grok-3's distillation)
-… See the full description on the dataset page: https://huggingface.co/datasets/iimran/UrduMedIQ-Urdu-Medical-Intelligence-Questions.cve-intelligence-stream
CVE Security Intelligence Stream
Enterprise-grade CVE feed for Threat Intelligence and Risk Management teams.
Incremental CVE records from NVD API 2.0 enriched with CISA Known Exploited Vulnerabilities (KEV) metadata. Normalized for LLM SFT/RAG.
B2B Value Proposition
This dataset is built for security vendors, MSSPs, and enterprise SOC/GRC teams who need:
Use case
What you get
Threat Intel enrichment
CVSS severity, affected software, exploit… See the full description on the dataset page: https://huggingface.co/datasets/FXBIA/cve-intelligence-stream.echidna-round1-intelligence
echidna-round1-intelligence
Echidna — round 1 general intelligence / helpful-answer training.
Contents
round1_intelligence.jsonl (59 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Echidna RAG assistant (Michael Anthony Falabella).
Artificial-intelligence-dataset-for-IR-systems
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
information-retrieval
semantic-search
Languages
English
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Adel-Elwan/Artificial-intelligence-dataset-for-IR-systems.mirror-threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-threat-intelligence-dataset.Medical-Intelligence-Questions
Medical-Intelligence-Questions Dataset
A comprehensive collection of 10,000+ expert-curated medical questions for training and evaluating clinical reasoning in AI models.
🔍 Overview
This dataset provides:
High-quality medical questions covering diverse clinical scenarios
Detailed explanations and answers verified by healthcare professionals
Multi-specialty coverage spanning common and rare conditions
Structured format optimized for LLM training and evaluation
Key… See the full description on the dataset page: https://huggingface.co/datasets/iimran/Medical-Intelligence-Questions.ai-failure-intelligence
AI Failure Intelligence Dataset
Structured, analyst-grade dataset of real-world AI failures.Built for researchers, red-teamers, and enterprise AI security teams.
Dataset Description
This is a FREE sample of 50 curated AI failure cases from the full
AI Failure Intelligence dataset (5,000+ cases, updated daily).
Each case is enriched with machine intelligence including:
Failure classification across 14 categories
Severity scoring (0-100)
Root cause analysis
Risk pattern… See the full description on the dataset page: https://huggingface.co/datasets/aifi-intelligence/ai-failure-intelligence.
