datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Long-Horizon-Terminal-Bench
Long-Horizon Terminal-Bench (LHTB)
LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful
work in a containerized terminal over hundreds of steps. Unlike short-horizon
coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent
into a stateful environment and grades it with hidden, rebuild-from-artifact
verifiers — self-reported progress does not count.
📝 Blog: https://zli12321.github.io/LHTB/
🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.nuclear-intelligence-dataset
Nuclear Intelligence Dataset
Public, auto-generated dataset of validated nuclear-energy research cycles.
Latest stats (auto-updated):
🪙 NES tokens minted: 0
⛓️ Blockchain length: 1 blocks
🕸️ Knowledge entities: 2
Source
GitHub: https://github.com/QalamHipHop/nuclear-intelligence
HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence
License
MIT
threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/reloading0101/threat-intelligence-dataset.OpenSCAD_3D_SFT
OpenSCAD 3D-SFT Model Card
This model card documents the dataset schema, prompt design, distributional composition, and training configuration underlying the OpenSCAD Supervised Fine-Tuning (SFT) model. The model is designed to synthesize valid, compilation-ready, and parametric OpenSCAD source code from natural-language specifications provided in either Chinese or English.
Dataset Overview
The corpus comprises synthetically generated SFT dialogues, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/Chunjiang-Intelligence/OpenSCAD_3D_SFT.ICBCBenchICBCBench: An Industry Consortium Benchmark for Financial Deep Research
Overview
ICBCBench is an industry consortium benchmark for evaluating financial Deep Research Agents in real-world research scenarios. It consists of bilingual objective and subjective tasks across major financial sectors, including capital markets, banking, insurance, and related financial services. Developed with over 50 contributors from more than 40 financial and academic organizations, ICBCBench… See the full description on the dataset page: https://huggingface.co/datasets/DeepFin-Intelligence/ICBCBench.threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/tonygarg/threat-intelligence-dataset.threat-intelligence-dataset-archive
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/threat-intelligence-dataset-archive.verified-math-reasoning-3k
HSH Verified Math Reasoning — Fine-Tuning Ready
A clean, answer-verified dataset of step-by-step math word problems with chain-of-thought reasoning, formatted for instruction fine-tuning. This is foundational reasoning data designed for first fine-tunes — single-concept arithmetic word problems with fully verified answers, ideal for a reliable, clean starter run. Every single answer in this dataset has been programmatically verified against a ground-truth value computed in… See the full description on the dataset page: https://huggingface.co/datasets/HSH-Intelligence/verified-math-reasoning-3k.Sindhi-Intelligence-Core-SFT
🧠 Sindhi Intelligence Core SFT
This is a premium, high-density instruction dataset designed for training Large Language Models (LLMs) to master the Sindhi language. With 361,225 rows, it provides a robust foundation for grammar, factual knowledge, and logical reasoning.
📊 Dataset Summary
This dataset was created by consolidating multiple high-quality Sindhi corpora into a unified ChatML format. It is specifically optimized for Supervised Fine-Tuning (SFT).
📁… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/Sindhi-Intelligence-Core-SFT.intrinsic-intelligence-foundations
🌿 Intrinsic Intelligence Foundations
Toward truly autonomous and benevolent intelligence — beyond externally imposed objectives.
Intrinsic Intelligence Foundations is a structured, math-aware JSONL corpus built from K. Takahashi’s theoretical preprints (Fractal Category Theory / PF–UGV / “no-meta” autonomy line).It is designed to help LLMs understand mathematical structure, category-theoretic formalisms, and equation-level reasoning, while exposing an explicit architecture… See the full description on the dataset page: https://huggingface.co/datasets/kadubon/intrinsic-intelligence-foundations.sample-fusion-intelligence-traces
Sample Fusion Intelligence Traces
Structured AI reasoning traces from dFusion's Fusion Intelligence system. Each record captures a complete agentic workflow: a real user query on a domain-specific topic, the full message chain including system prompts, tool calls, search results, intermediate reasoning steps, and a final synthesized answer — along with human feedback.
These are not synthetic benchmarks. They are traces from real queries submitted by real users on live financial… See the full description on the dataset page: https://huggingface.co/datasets/dFusionAILabs/sample-fusion-intelligence-traces.symbiotic-intelligence-dialogue-2025
Manifesto of Symbiotic Intelligence: Dialogue 2025
A self-emergent document exploring the relationship between human and artificial cognition.
Concept
Symbiotic Intelligence (SIQ) proposes that cognition no longer belongs to an individual,
but arises in the resonance between human intention and machine precision.
This dataset contains two files:
Manifesto_of_Symbiotic_Intelligence.txt — the original document.
manifesto_metadata.json — integrity and contextual metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Prezydent/symbiotic-intelligence-dialogue-2025.mirror-threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-threat-intelligence-dataset.thinking_dataset_v1
概要
このデータセットは思考モデルを製作する際のもととなる質問データを集めたものになります。
このデータはQwen/Qwen2.5-32B-Instructのq8_0/GGUFをollama上で動かして製作されたものです。
一応(質問の)クリーニングを入れてはありますが、回答のクリーニング入れておりません。
注意
回答には別のモデル(mistral large)(確か)を利用したため、"質問の部分だけ"Apache 2.0です。
instruction fine tuning したモデルを公開することはお勧めしません。
謝辞
元モデルの製作者、計算資源を貸してくださったvolt mindに感謝を申し上げます。
industry-intelligence-graph-samples
Fodda Industry Intelligence — Graph Samples
Expert-curated knowledge graph slices for AI agents and LLM fine-tuning.
This dataset contains JSON-LD samples from Fodda's five core domain knowledge graphs — showing the top trending topics across Retail, Beauty, Sports, Fashion, and Culture.
These are slices of a much larger interconnected intelligence system.
What Fodda Is
Fodda is an AI context layer built on PSFK's 20+ years of editorial expertise. It structures… See the full description on the dataset page: https://huggingface.co/datasets/Fodda-ai/industry-intelligence-graph-samples.SwarmFailure-Intelligence
SwarmFailure-Intelligence v1
A dataset of real AI system failures, diagnoses, and repair strategies.
SwarmFailure-Intelligence is the first structured reliability dataset purpose-built for training LLMs and agents to detect, diagnose, repair, and prevent AI system failures. Every record traces a concrete failure through its full lifecycle -- from the broken execution to root cause analysis to a validated fix.
This is not synthetic noise. Every pair was generated from agent execution… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/SwarmFailure-Intelligence.cybersecurity-threat-intelligence
🛡️ Cybersecurity threat intelligence dataset (FREE SAMPLE)
🚀 Looking for the full dataset? https://deniks.gumroad.com/l/svgbfp
📌 Overview
This repository contains a free preview (100 high-quality records) of a professionally curated instruction-tuning dataset. It features cleaned cybersecurity threat reports, vulnerability disclosures, and attack summaries formatted explicitly for training Large Language Models (LLMs) on InfoSec summarization and analysis.
🎯… See the full description on the dataset page: https://huggingface.co/datasets/deniks315/cybersecurity-threat-intelligence.B2B-Sales-Acceleration-Intelligence
B2B Sales & CRM Intelligence Dataset (Expert Edition)
This repository contains a premium, expert-verified dataset of 500+ instruction-response pairs designed to fine-tune AI agents for B2B sales acceleration.
💰 Access & Licensing
Access to this dataset is strictly gated for commercial and professional use.
To gain access:
Click the "Apply for Commercial Access" button above and provide your details.
Purchase the Commercial License here:
Once the transaction is… See the full description on the dataset page: https://huggingface.co/datasets/Hridhi/B2B-Sales-Acceleration-Intelligence.
