datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
infosec-tool-output
Infosec Tool Output
Security-tool output → evidence-backed, plain-English interpretation.
A dataset for training and evaluating models that interpret security-tool output, explain the limits of the evidence, and recommend defensive next steps.
v2.0.0: 1,004 canonical examples across 19 tools. This includes all 776 original records with traceable interpretation changes, plus 228 newly authored synthetic fixtures. The deduplicated training views contain 1004 examples, not… See the full description on the dataset page: https://huggingface.co/datasets/tegridydev/infosec-tool-output.simpsons_infoThe information of all episodes of the cartoon show "The Simpsons" from wikipedia. Some (mainly in recent 32, 33, 34 seasons) plot missing.
Institutional-Information-of-Bangladesh
Institutional-Information-of-Bangladesh Dataset
This Dataset contains all verified and authorized Institutional information in Bangladesh
Description
I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
http://data.gov.bd/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Institutional-Information-of-Bangladesh.dart-math-pool-gsm8k-query-info
[!NOTE]
This dataset is the synthesis information of queries from the GSM8K training set,
such as the numbers of raw/correct samples of each synthesis job.
Usually used with dart-math-pool-gsm8k.
🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub
🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX
Datasets: DART-Math
DART-Math datasets are the… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-gsm8k-query-info.quantum-information-and-complexity-theory
Neura Parse — Quantum Information & Complexity Theory: Channels, Entropies, Classes & the Structure of Advantage
A proof-based theoretical-foundations vertical uniting quantum information theory (channels, entropies, entanglement measures, distinguishability, capacities, Shannon theory) with quantum complexity theory and the structure of quantum advantage (classes, Hamiltonian complexity, sampling-based advantage and its verification, pseudorandomness, dequantization).… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-information-and-complexity-theory.Dense-Information-Science-Physics-Dataset
Dense Information With Multiple Fine-tuned Variations
This dataaset has multiple for each input to learn how to express the same answer in different ways
Dataset Structure
The dataset contains two columns:
Column
Description
input
A science or quantum-physics question
output
A conversational answer to the question
Example:
{
"input": "What is quantum entanglement?",
"output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.dart-math-pool-math-query-info
[!NOTE]
This dataset is the synthesis information of queries from the MATH training set,
such as the numbers of raw/correct samples of each synthesis job.
Usually used with dart-math-pool-math.
🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub
🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX
Datasets: DART-Math
DART-Math datasets are the state-of-the-art… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-math-query-info.roblox-info-dump
Roblox-Info-Dump
The Roblox-Info-Dump dataset is a collection of public Roblox documentation from create.roblox.com/docs/ and luau.org. Roblox maintains the copyright on all content.
task684_online_privacy_policy_text_information_type_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task684_online_privacy_policy_text_information_type_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task684_online_privacy_policy_text_information_type_generation.Informal-Standard-English-Corpus
Dataset Description
This dataset is a parallel corpus of approximately 11,000 pairs of informal conversational English text and their normalized equivalents. The informal text mimics real-world digital communication, featuring slang, phonetic spellings, missing punctuation, and abbreviations. The normalized text provides a grammatically correct and semantically equivalent version.
The dataset was created to support machine translation tasks for low-resource languages. It… See the full description on the dataset page: https://huggingface.co/datasets/Bendang/Informal-Standard-English-Corpus.DoD-Instruction-8130-01-Installation-of-Geospatial-Information-And-Services
🗺️ DoD Installation Geospatial Information and Services Question-Answer Dataset
Source: DoD Instruction 8130.01
Source Effective Date: April 9, 2015
Change Incorporated: Change 3, effective August 4, 2020
Source Organization: Office of the Under Secretary of Defense for Acquisition and Sustainment
Source Ownership: United States Department of Defense
📋 Overview
Dataset Summary
The DoD Installation Geospatial Information and Services… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8130-01-Installation-of-Geospatial-Information-And-Services.DOD-Instruction-5040-02-Visual-Information
DoD Visual Information Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains 250 document-grounded question-and-answer records based on DoD Instruction 5040.02, “Visual Information (VI),” dated October 27, 2011, and incorporating Change 2 effective April 20, 2018.
The source establishes Department of Defense policy, responsibilities, and procedures for managing visual-information records… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Instruction-5040-02-Visual-Information.DoD-Instruction-8170-01-Online-Information-Management-And-Electronic-Messaging
📚 DoD Instruction 8170.01 Online Information Management and Electronic Messaging
Maintainer: Terry Eppler
Ownership: U.S. Department of Defense
📋 Overview
Dataset Summary
The DoD Instruction 8170.01 Online Information Management and Electronic Messaging
Dataset is a structured natural-language question-answering dataset derived from
DoD Instruction 8170.01, Online Information Management and Electronic Messaging.
DoD Instruction 8170.01… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8170-01-Online-Information-Management-And-Electronic-Messaging.medicine-information-ptuncgpt-conversations-informal-approved-1p25
UncGPT — Informal-Register Approved
Conversations that passed the 1.25σ semantic gate AND the current strict programmatic gates — including intimate-register (tú-not-usted, tu-not-shoma, 你-not-您, no po/opo, plain not keigo), stricter colloquial Persian, and strict completion-integrity.
Part of the UncGPT NeurIPS 2026 Competition collection.
Counts
approved: 753
rejected: 1,475
skills covered: 53 of 69
by care: warm 450 / mid 152 / cold 151
by language: en 310 / sw… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-informal-approved-1p25.DoD-Instruction-8010-01-Information-Network-Transport
🌐 DoD Information Network Transport
Maintainer: Terry Eppler
Owner: US Federal Government
Source: DoD Instruction 8010.01
Dataset Size: question-answer records
Source Effective Date: September 10, 2018
Source Organization: Office of the DoD Chief Information Officer
Source Ownership: United States Department of Defense
📋 Overview
Dataset Summary
The DoD Information Network Transport Question-Answer Dataset contains document-grounded… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8010-01-Information-Network-Transport.Legacy-Code-Dataset
Legacy Codebase Dataset
Dataset Description
The Legacy Codebase Dataset is a large-scale collection of enterprise software repositories designed for training next-generation Large Language Models (LLMs), AI coding assistants, software engineering copilots, automated refactoring systems, repository understanding models, and intelligent program analysis pipelines.
The complete collection contains 405 real-world legacy codebases spanning 23 major industries… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Legacy-Code-Dataset.infopedia-pt-ipa
European Portuguese IPA Lexicon — Infopédia
A lightweight word → IPA pronunciation lexicon for European Portuguese,
extracted from Infopédia (Porto Editora). One row
per headword, intended for grapheme-to-phoneme (G2P) work, pronunciation
modelling, and TTS/ASR lexicon building.
Complete crawl. Derived from a graph crawl of Infopédia that ran to
convergence (frontier → 0), covering the dictionary's reachable component.
Contents
Field
Count
Entries… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/infopedia-pt-ipa.DoD-Instruction-5200-01-Information-Security-Program
DoD Information Security and SCI Protection Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on DoD Instruction 5200.01, “DoD Information Security Program and Protection of Sensitive Compartmented Information (SCI),” dated April 21, 2016, and incorporating Change 2 effective October 1, 2020.
The source establishes the overarching Department… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-5200-01-Information-Security-Program.speculators-multilingual-en-fr-de-it-es
Speculators Multilingual SFT Dataset (en/fr/de/it/es)
A multilingual instruction-following dataset in ShareGPT format, built to train draft models for speculative decoding across English, French, German, Italian and Spanish.
Summary
An English instruction-tuning corpus with part of it kept in English and the rest machine-translated into French, German, Italian and Spanish using tencent/Hunyuan-MT-7B. Provided as a single mixed-language, ShareGPT-formatted dataset… See the full description on the dataset page: https://huggingface.co/datasets/Infomaniak-AI/speculators-multilingual-en-fr-de-it-es.synthetic-info-extract-json
Raw text to json object (synthetic)
Amad Zarak
February 28, 2026
Created using gpt-oss-120b on h200 sxm
Zero-shot JSON schema deduction & universal information extraction.
80,664 rows
Figured others could use this since it is basically impossible to find massive raw-text-to-structured-json datasets for training extraction engines.
About 30k of the raw outputs hit the token limit and malformed, but I ran a massive salvage sweep on the raw outputs using json-repair to force the… See the full description on the dataset page: https://huggingface.co/datasets/amadzarak/synthetic-info-extract-json.agda-categories-informalized
agda-categories, informalized
4,541 declarations from the agda-categories
library, each paired with an informal, LaTeX-flavoured natural-language
statement written by GLM-5.2. The natural language is written to be precise
enough to re-formalise from, so the intended use is training a model to
reconstruct the formal Agda source from the prose alone.
Declarations were extracted with a fork of Agda
that dumps one JSON record per named declaration (with its full source range)
during… See the full description on the dataset page: https://huggingface.co/datasets/astral-expmath/agda-categories-informalized.DOD-Directive-8000-01-Management-Of-Defense-Information
DoD Directive 8000.01 Management of the DoD Information Enterprise Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on Department of Defense Directive 8000.01, “Management of the Department of Defense Information Enterprise,” dated March 17, 2016, and incorporating Change 1 effective July 27, 2017.
The directive establishes Department-wide… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Directive-8000-01-Management-Of-Defense-Information.japanese-confidential-information-extraction-sft
Japanese Confidential Information Extraction — SFT Dataset
日本語テキストから社外秘の固有表現を抽出するタスク向けの SFT (Supervised Fine-Tuning) データセットです。
LFM2 系モデルの LoRA fine-tune を想定して構築されています。
タスク概要
入力テキスト(日本語)に含まれる機密情報を、11カテゴリの JSON として抽出します。
入力: 「山田太郎(yamada@example.co.jp)から請求書番号 INV-2024-0042 で
売上 ¥12,800,000 の見積書が届いた。」
出力: {
"address": [],
"company_name": [],
"email_address": ["yamada@example.co.jp"],
"human_name": ["山田太郎"],
"phone_number": [],
"account_identifier":… See the full description on the dataset page: https://huggingface.co/datasets/akiFQC/japanese-confidential-information-extraction-sft.information-security-policies-qa-distiset
Dataset Card for information-security-policies-qa-distiset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/daqc/information-security-policies-qa-distiset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/davidquicast/information-security-policies-qa-distiset.DSA-Coding-Problems-and-Solutions-Dataset
Dataset Description
This dataset is a large-scale collection of Data Structures and Algorithms (DSA) code, containing 12,385 code files with 3.86 million lines of code and 25.01 million lexical tokens, designed to support the development of advanced code generation models, programming assistants, software engineering AI systems, and code intelligence applications.
It consists of real-world DSA implementations covering a wide range of algorithms, data structures, problem-solving… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/DSA-Coding-Problems-and-Solutions-Dataset.task1284_hrngo_informativeness_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1284_hrngo_informativeness_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1284_hrngo_informativeness_classification.infosec_harmful_behaviors
Infosec Harmful Behaviors
Offensive-security instruction prompts for refusal-direction research and abliteration of code/security models.
Dataset Details
This dataset contains infosec-domain harmful prompts intended to elicit refusal behavior from aligned instruction models. It is designed as the harmful side of a harmful/harmless contrast pair, analogous to mlabonne/harmful_behaviors but focused on offensive-security and malicious-coding requests.
Rows:
train:… See the full description on the dataset page: https://huggingface.co/datasets/zaakirio/infosec_harmful_behaviors.frameref
Dataset Information
Information ecosystems increasingly shape how people internalize exposure to adverse digital experiences, raising concerns about the long-term consequences for information health. In modern search and recommendation systems, ranking and personalization policies play a central role in shaping such exposure and its long-term effects on users. To study these effects in a controlled setting, we present FrameRef, a large-scale dataset of 1,073,740 systematically… See the full description on the dataset page: https://huggingface.co/datasets/infosense/frameref.offsec_redteam_info
OffSec RedTeam Info
OffSec RedTeam Info is a SlimPajama‑style, category‑organized corpus of security knowledge text crawled from reputable red‑team/blue‑team websites: wikis, training blogs, vendor research, CERT advisories, reversing/malware labs, cloud/kubernetes posts, OSINT handbooks, AD tradecraft, and more.
Token count: ~1.646B tokens.
⚠️ Ethical use only. Use for research, education, and defensive security. Respect robots.txt, site terms, and copyrights. Do not misuse this… See the full description on the dataset page: https://huggingface.co/datasets/tandevllc/offsec_redteam_info.
