CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Qalam /nuclear-intelligence-dataset Nuclear Intelligence Dataset Public, auto-generated dataset of validated nuclear-energy research cycles. Latest stats (auto-updated): 🪙 NES tokens minted: 0 ⛓️ Blockchain length: 1 blocks 🕸️ Knowledge entities: 2 Source GitHub: https://github.com/QalamHipHop/nuclear-intelligence HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence License MIT tabularquestion-answeringn<1K1 likes2.7k downloads4h agoHugging Face02reloading0101 /threat-intelligence-dataset Cyber Threat Intelligence Dataset for LLM Fine-Tuning An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on. The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/reloading0101/threat-intelligence-dataset.texttext-generation10K<n<100K10 likes543 downloads3mo agoHugging Face03sphita /Intel-WebCorpus-forms 💻 Intel WebCorpus Forms (Enterprise Hardware Q&A) This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums. It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and… See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.textquestion-answering100K<n<1M3 likes485 downloads1d agoHugging Face04simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes410 downloads6h agoHugging Face05simpleG2023 /chinese-clean-energy-battery-open-intelligence 🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.tabulartext-retrieval1K<n<10K0 likes313 downloads5h agoHugging Face06simpleG2023 /chinese-ai-and-robotics-open-intelligence 🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes284 downloads5h agoHugging Face07simpleG2023 /chinese-biomedicine-and-genomics-open-intelligence 🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes274 downloads5h agoHugging Face08BAAI /IndustryInstruction_Artificial-Intelligence IndustryInstruction: Artificial Intelligence This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.tabularquestion-answering100K<n<1M2 likes212 downloads1mo agoHugging Face09ismailtasdelen /unified-vulnerability-intelligence-dataset Unified Vulnerability Intelligence Dataset (UVID) v3.0 — Cyber Security Knowledge Graph UVID is a structured cyber security knowledge graph that unifies multiple vulnerability classification frameworks into a single knowledge base. Each of the 250 records describes one application/software security vulnerability and links it — where authoritative data exists — across CWE, CAPEC, MITRE ATT&CK, CVSS, 14 OWASP projects, secure-fix intelligence, detection surfaces, programming… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/unified-vulnerability-intelligence-dataset.texttext-classificationn<1K1 likes171 downloads2mo agoHugging Face10DeepFin-Intelligence /ICBCBenchICBCBench: An Industry Consortium Benchmark for Financial Deep Research Overview ICBCBench is an industry consortium benchmark for evaluating financial Deep Research Agents in real-world research scenarios. It consists of bilingual objective and subjective tasks across major financial sectors, including capital markets, banking, insurance, and related financial services. Developed with over 50 contributors from more than 40 financial and academic organizations, ICBCBench… See the full description on the dataset page: https://huggingface.co/datasets/DeepFin-Intelligence/ICBCBench.texttext-generationn<1K1 likes90 downloads11d agoHugging Face11Arc-Intelligence /tau2-mms-teacher-traces Tau2 Teacher Traces Dataset Dataset Description This dataset contains teacher reasoning traces for solving MMS (Multimedia Messaging Service) issues in the τ²-bench (Tau2-bench) framework. Each example includes a teacher model's thinking process and structured teaching guidance for resolving customer service tickets. Dataset Summary Domain: Telecom customer service Task: MMS troubleshooting Size: 49 examples Format: JSONL Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Arc-Intelligence/tau2-mms-teacher-traces.texttext-generationn<1K3 likes88 downloads10mo agoHugging Face12huzaifa525 /Medical_Intelligence_Dataset_76k_2026_Edition 🏥 Medical Intelligence Dataset · 76k · 2026 Edition Production-ready medical AI dataset for training diagnosis, treatment reasoning, and doctor-patient conversational systems. 76,000 engineered (not collected) English Q&A pairs — covering 620+ diseases, 438+ FDA-approved drugs, and real patient-doctor conversations. Built with a 5-stage quality pipeline. Commercial-safe Apache 2.0. Created by Huzefa Nalkheda Wala — AI Product Engineer & Medical AI Researcher · Creator of the… See the full description on the dataset page: https://huggingface.co/datasets/huzaifa525/Medical_Intelligence_Dataset_76k_2026_Edition.textquestion-answering10K<n<100K1 likes88 downloads5mo agoHugging Face13tonygarg /threat-intelligence-dataset Cyber Threat Intelligence Dataset for LLM Fine-Tuning An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on. The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/tonygarg/threat-intelligence-dataset.texttext-generation10K<n<100K0 likes81 downloads28d agoHugging Face14ChipHolmes /threat-intelligence-dataset-archive Cyber Threat Intelligence Dataset for LLM Fine-Tuning An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on. The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/threat-intelligence-dataset-archive.texttext-generation10K<n<100K1 likes80 downloads2mo agoHugging Face15louisbrulenaudet /code-propriete-intellectuelle Code de la propriété intellectuelle, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-propriete-intellectuelle.tabulartext-generation1K<n<10K0 likes72 downloads2y agoHugging Face16HSH-Intelligence /verified-math-reasoning-3k HSH Verified Math Reasoning — Fine-Tuning Ready A clean, answer-verified dataset of step-by-step math word problems with chain-of-thought reasoning, formatted for instruction fine-tuning. This is foundational reasoning data designed for first fine-tunes — single-concept arithmetic word problems with fully verified answers, ideal for a reliable, clean starter run. Every single answer in this dataset has been programmatically verified against a ground-truth value computed in… See the full description on the dataset page: https://huggingface.co/datasets/HSH-Intelligence/verified-math-reasoning-3k.texttext-generation1K<n<10K0 likes66 downloads3mo agoHugging Face17aakashMeghwar01 /Sindhi-Intelligence-Core-SFT 🧠 Sindhi Intelligence Core SFT This is a premium, high-density instruction dataset designed for training Large Language Models (LLMs) to master the Sindhi language. With 361,225 rows, it provides a robust foundation for grammar, factual knowledge, and logical reasoning. 📊 Dataset Summary This dataset was created by consolidating multiple high-quality Sindhi corpora into a unified ChatML format. It is specifically optimized for Supervised Fine-Tuning (SFT). 📁… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/Sindhi-Intelligence-Core-SFT.texttext-generation100K<n<1M1 likes58 downloads7mo agoHugging Face18blackboxanalytics /cybersec-threat-intel-qa Cybersecurity Threat Intelligence QA Dataset A verified binary forecasting dataset covering cybersecurity threats, vulnerabilities, and incident response — generated using the Lightning Rod Labs SDK. Dataset Summary 455 verified binary forecasting QA pairs across 14 cybersecurity subcategories, covering 90 days of real-world cybersecurity news (November 2025 – February 2026). Each entry includes a question, a verified yes/no label, detailed multi-paragraph reasoning with… See the full description on the dataset page: https://huggingface.co/datasets/blackboxanalytics/cybersec-threat-intel-qa.tabularquestion-answeringn<1K0 likes54 downloads7mo agoHugging Face19Omartificial-Intelligence-Space /Arabic-gsm8k-v2 Dataset Summary Arabic GSM8K is an Arabic translation of the GSM8K (Grade School Math 8K) dataset, which contains high-quality linguistically diverse grade school math word problems. The original dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning, and this Arabic version aims to extend these capabilities to Arabic language models and applications. The dataset maintains the same characteristics as the… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-gsm8k-v2.textquestion-answering10K<n<100K1 likes53 downloads1y agoHugging Face20Shomi28 /cyber-threat-intelligence Cyber Threat Intelligence Dataset A comprehensive cybersecurity dataset combining CVE vulnerability data, MITRE ATT&CK techniques, and CISA Known Exploited Vulnerabilities — structured for AI/ML training and security research. Author: Soham DahivalkarLicense: MITCreated: 2026 Dataset Description This dataset provides structured cybersecurity intelligence data collected from three authoritative public sources: NVD (National Vulnerability Database) — CVE… See the full description on the dataset page: https://huggingface.co/datasets/Shomi28/cyber-threat-intelligence.texttext-generation10K<n<100K1 likes51 downloads5mo agoHugging Face21suayptalha /Intelix-Dataset texttext-generation100K<n<1M1 likes49 downloads1y agoHugging Face22intellistream /sage-agent-benchmark SAGE Agent Benchmark Comprehensive benchmark for evaluating AI agent capabilities across three core competencies: Tool Selection - Choosing appropriate tools for tasks Task Planning - Decomposing complex tasks into step sequences Timing Judgment - Deciding when to use tools vs. direct answers Dataset Statistics Total Samples: ~11,000 Tool Selection: ~6,000 samples Task Planning: ~3,000 samples Timing Judgment: ~2,000 samples Splits: train, dev, test Usage… See the full description on the dataset page: https://huggingface.co/datasets/intellistream/sage-agent-benchmark.textquestion-answering10K<n<100K1 likes45 downloads8mo agoHugging Face23malhajar /distilabel-intel-orca-dpo-pairs-tr Dataset Card for "malhajar/orca_dpo_pairs-tr" This Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish dataset collection to enhance the performance of LLM's Produced in the Turkish Language. malhajar/orca_dpo_pairs-tr is a translated version of argilla/distilabel-intel-orca-dpo-pairs Translated by: Mohamad Alhajar Dataset Summary This is a pre-processed version of the OpenOrca dataset translated to… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/distilabel-intel-orca-dpo-pairs-tr.texttext-classification1K<n<10K6 likes44 downloads2y agoHugging Face24Omartificial-Intelligence-Space /Arabic_Openai_MMMLU Arabic Multilingual Massive Multitask Language Understanding (MMMLU) The MMLU is a widely recognized benchmark for assessing general knowledge attained by AI models. It covers a broad range of topics across 57 different categories, from elementary-level knowledge to advanced professional subjects like law, physics, history, and computer science. We have extracted the Arabic subset from the MMMLU test set, which was translated by professional human translators. This dataset, now… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic_Openai_MMMLU.textquestion-answering10K<n<100K4 likes44 downloads2y agoHugging Face25leeroy-jankins /DoD-Instruction-5400-19-Public-Affairs-Use-of-Artificial-Intelligence DoD Public Affairs Use of Artificial Intelligence Maintainer: Terry Eppler Owner: US Federal Government Dataset Summary This dataset contains 150 document-grounded question-and-answer records based on DoD Instruction 5400.19, “Public Affairs Use of Artificial Intelligence,” effective July 28, 2025. The source establishes Department of Defense policy, responsibilities, and procedures for the appropriate use of artificial-intelligence capabilities in… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-5400-19-Public-Affairs-Use-of-Artificial-Intelligence.documentquestion-answering0 likes44 downloads1mo agoHugging Face26kadubon /intrinsic-intelligence-foundations 🌿 Intrinsic Intelligence Foundations Toward truly autonomous and benevolent intelligence — beyond externally imposed objectives. Intrinsic Intelligence Foundations is a structured, math-aware JSONL corpus built from K. Takahashi’s theoretical preprints (Fractal Category Theory / PF–UGV / “no-meta” autonomy line).It is designed to help LLMs understand mathematical structure, category-theoretic formalisms, and equation-level reasoning, while exposing an explicit architecture… See the full description on the dataset page: https://huggingface.co/datasets/kadubon/intrinsic-intelligence-foundations.texttext-retrievaln<1K1 likes42 downloads2mo agoHugging Face27dFusionAILabs /sample-fusion-intelligence-traces Sample Fusion Intelligence Traces Structured AI reasoning traces from dFusion's Fusion Intelligence system. Each record captures a complete agentic workflow: a real user query on a domain-specific topic, the full message chain including system prompts, tool calls, search results, intermediate reasoning steps, and a final synthesized answer — along with human feedback. These are not synthetic benchmarks. They are traces from real queries submitted by real users on live financial… See the full description on the dataset page: https://huggingface.co/datasets/dFusionAILabs/sample-fusion-intelligence-traces.tabularquestion-answeringn<1K0 likes40 downloads6mo agoHugging Face28iimran /UrduMedIQ-Urdu-Medical-Intelligence-QuestionsBased on the document you've shared, here's the content formatted in Markdown: # UrduMed-Grok3-Verifiable Dataset A high-quality **Urdu-language medical reasoning dataset** distilled from [Medical-O1-Verifiable-Problem] using **Grok-3**, designed to enhance Urdu LLMs' clinical reasoning and diagnostic accuracy. --- ## 🔍 Overview This dataset adapts challenging, open-ended medical problems from English to Urdu while preserving: - **Medical accuracy** (via Grok-3's distillation) -… See the full description on the dataset page: https://huggingface.co/datasets/iimran/UrduMedIQ-Urdu-Medical-Intelligence-Questions.textquestion-answering10K<n<100K0 likes39 downloads1y agoHugging Face29MichaelAnthony /echidna-round1-intelligence echidna-round1-intelligence Echidna — round 1 general intelligence / helpful-answer training. Contents round1_intelligence.jsonl (59 rows) Format JSON Lines (.jsonl), one example per line. Provenance Original content for the Echidna RAG assistant (Michael Anthony Falabella). textquestion-answeringn<1K0 likes31 downloads29d agoHugging Face30joshause /spinoza-treatise-emendation-intellect-100-qa Description This dataset contains 100 synthetic question-answer pairs based on Spinoza's posthumously published Tractatus de Intellectus Emendatione (1677) as translated by R. H. L. Elwes' On the Improvement of the Understanding (Treatise on the Emendation of the Intellect (1883). The questions and answers represent a comprehensive overview of the ideas and principles set forth by Spinoza in the treatise. The question-answer pairs were generated, reviewed, and refined in iterative… See the full description on the dataset page: https://huggingface.co/datasets/joshause/spinoza-treatise-emendation-intellect-100-qa.textquestion-answeringn<1K1 likes29 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.