datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cyber-security
Cybersecurity AI Knowledge Base — PhD-Level Dataset
Overview
This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security.
Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms
Purpose
Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.securecode-web
SecureCode Web: Traditional Web & Application Security Dataset
Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance
Paper | GitHub | Dataset | Model Collection | Blog Post
What's new in v2.6
v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had
shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/scthornton/securecode-web.White-Hat-Security-Agent-Prompts-600K
White Hat Security Agent Prompts 600K
Overview
The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios.
Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.sec-material-contracts-qa800+ EDGAR contracts with PDF images and key information extracted by the OpenAI GPT-4o model.
The key information is defined as follows:
class KeyInformation(BaseModel):
agreement_date : str = Field(description="Agreement signing date of the contract. (date)")
effective_date : str = Field(description="Effective date of the contract. (date)")
expiration_date : str = Field(description="Service end date or expiration date of the contract. (date)")
party_address : str =… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts-qa.DecodingTrust
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Overview
This repo contains the source code of DecodingTrust. This research endeavor is designed to help researchers better understand the capabilities, limitations, and potential risks associated with deploying these state-of-the-art Large Language Models (LLMs). See our paper for details.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Boxin Wang, Weixin Chen, Hengzhi… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust.SecQue
SECQUE
Paper
SECQUE is a comprehensive benchmark for evaluating large language models (LLMs) in financial analysis tasks.
SECQUE comprises 565 expert-written questions covering SEC filings analysis across four key categories:
comparison analysis
ratio calculation
risk assessment
financial insight generation.
To assess model performance, we develop SECQUE-Judge, an evaluation mechanism leveraging multiple LLM-based judges, which demonstrates strong alignment with human… See the full description on the dataset page: https://huggingface.co/datasets/nogabenyoash/SecQue.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.omnimcp_cybersecurity_secops_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cybersecurity_secops_teaser.indian-legal-sections-bns-bnss-bsa-2023
🏛️ Indian Legal Sections — BNS · BNSS · BSA 2023
The First Structured, Unified JSON Dataset of Modern Indian Criminal Law
📖 Dataset Summary
This dataset contains 1,059 fully structured and verified sections extracted, parsed, and unified from India's three landmark criminal justice reform acts passed in December 2023. These three acts together replaced the colonial-era Indian Penal Code (IPC, 1860), the Code of Criminal Procedure… See the full description on the dataset page: https://huggingface.co/datasets/GSMS-B/indian-legal-sections-bns-bnss-bsa-2023.cyber-security
Cybersecurity Instruction-Tuning Dataset
A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning,
built from 198 distinct sources spanning offensive security, blue-team
operations, vulnerability intelligence, cloud/AWS security, malware analysis,
digital forensics, and more. Every record is normalized to the standard
messages chat format and deduplicated at both file and record level.
⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.securecode
SecureCode: Comprehensive Security Training Dataset for AI Coding Assistants
The largest open security training dataset for AI coding assistants, covering both traditional web security and AI/ML security
Overview
SecureCode combines 2,372 security-focused training examples into a single, unified dataset with HuggingFace configs for flexible loading. Every example provides vulnerable code, explains why it's dangerous, demonstrates a secure alternative, and… See the full description on the dataset page: https://huggingface.co/datasets/scthornton/securecode.SecureCodePairs
Dataset Summary
Field
Value
Version
1.2.0
License
MIT
Total code examples
470
LLM security trajectories
30
Languages (15)
Python, Java, JavaScript, TypeScript, Go, PHP, C#, Kotlin, Swift, Rust, Ruby, C, C++, Scala, YAML (Kubernetes)
Frameworks
Flask, Django, FastAPI, Spring Boot, Express, NestJS, Next.js, Laravel, ASP.NET Core, Gin, Android, iOS, Actix, Rails, Qt, Play, gRPC, GraphQL, Kubernetes
New in v1.2.0
+260 records (deep Python/Java packs… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/SecureCodePairs.SecCodePLT
SecCodePLT
SecCodePLT is a unified and comprehensive evaluation platform for code GenAIs' risks.
1. Dataset Details
1.1 Dataset Description
Language(s) (NLP): English
License: MIT
1.2 Dataset Sources
Repository: Coming soon
Paper: https://arxiv.org/pdf/2410.11096
Demo: https://seccodeplt.github.io/
2. Uses
2.1 Direct Use
This dataset can be used for evaluate the risks of large language models generating… See the full description on the dataset page: https://huggingface.co/datasets/Virtue-AI-HUB/SecCodePLT.securecode-web-archive
SecureCode Web: Traditional Web & Application Security Dataset
Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance
Paper | GitHub | Dataset | Model Collection | Blog Post
What's new in v2.6
v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had
shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/securecode-web-archive.sec-material-contracts-qa-splittedMixed and filtered version of chenghao/sec-material-contracts-qa and jordyvl/DUDE_subset_100val.
bitcoin-wallet-security-qa
Bitcoin Wallet Security Dataset
A high-quality question–answer dataset of 500 records focused on Bitcoin wallet
security, self-custody, backup and recovery planning, and common attack vectors. It is
built to train and evaluate AI systems that help people secure their Bitcoin — fine-tuning
LLMs, powering retrieval-augmented generation (RAG), security-focused assistants, and
educational chatbots.
Every record pairs a realistic security question with a detailed, self-contained… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-wallet-security-qa.Security-TTP-Mapping
The Security Attack Pattern (TTP) Recognition or Mapping Task
We share in this repo the MITRE ATT&CK mapping datasets, with training, validation and test splits.
The datasets can be considered as an emerging and challenging multilabel classification NLP task, with over 600 hierarchical classes.
NOTE: due to their security nature, these datasets contain textual information about malware and other security aspects.
Datasets
TRAM
This dataset belongs to CTID… See the full description on the dataset page: https://huggingface.co/datasets/tumeteor/Security-TTP-Mapping.Open-SECQA
Open-SECQA
Open-domain financial QA benchmark (a.k.a. LOFin) built on 145,897 SEC filings from 516 S&P 500 companies (Oct 2001 – Apr 2025), with 1,595 QA pairs covering single-document, multi-document, and multi-hop reasoning.
📄 Paper: ACL 2025 Findings
💻 Code: LOFin-bench-HiREC
Composition
Source
# QAs
FinQA
1,112
SEC-QA
333
FinanceBench
150
Total
1,595
Citation
@inproceedings{choe-etal-2025-hierarchical,
title =… See the full description on the dataset page: https://huggingface.co/datasets/HYdsl/Open-SECQA.Benchmarks_CyberSec_SecBench
Dataset Card for SecBench (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original SecBench dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit and rights belong to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original SecBench. It has been… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SecBench.cairo-security-audits
Cairo Security Audits
A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations.
Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.sec-filings-qa-instruct
SEC Filings Instruction-Tuning Dataset (Llama-3 Format)
This dataset contains 5,000 curated, instruction-formatted question-answering pairs derived from corporate SEC filings (Forms 10-K and 10-Q). It is structured specifically for parameter-efficient instruction fine-tuning (SFT/QLoRA) of Small Language Models using the standard Llama-3 ChatML template.
Dataset Details
Origin Source: Curated subset extracted from nvidia/Nemotron-SpecializedDomains-Finance-v1.… See the full description on the dataset page: https://huggingface.co/datasets/lateesha-bhatia/sec-filings-qa-instruct.nemotron-terminal-security
nemotron-terminal-security
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "security". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-security.finsight-sec-filings
FinSight — SEC EDGAR Filings
Cleaned plain-text 10-K (annual) and 10-Q (quarterly) filings from the
US SEC EDGAR system for 20 large publicly-traded companies across 6 sectors.
Created as part of the FinSight project —
a financial research AI assistant combining BERT fine-tuning, RAG, and
multi-agent systems.
Stats
Records: 97
Companies: 20 (AAPL, MSFT, GOOGL, AMZN, META, NVDA, TSLA, JPM, BAC, GS,
JNJ, PFE, UNH, WMT, PG, KO, MCD, XOM, CVX, CAT)
Forms: 10-K, 10-Q… See the full description on the dataset page: https://huggingface.co/datasets/musk1209/finsight-sec-filings.sec-contracts-2015-2025
SEC Contracts 2015–2025
Mini corpus of contract text extracted from SEC filings (EDGAR). Includes clause‑level rows with metadata for classification and QA experiments.
turkish_cyber_security_controls_benchmark
Turkish Cyber Security Controls Benchmark
Türkçe siber güvenlik kontrol seçimi ve kontrol denetimi yeteneğini ölçmek için
hazırlanmış, senaryo tabanlı çoktan seçmeli değerlendirme kümesidir.
v0.1.0, uzman incelemesine açık ilk sürümdür ve NIST SP 800-53 Rev. 5,
Release 5.2.0 kontrol kataloğunu hedefler.
Kapsam
100 Türkçe senaryo
NIST SP 800-53'ün 20 kontrol ailesinin her birinden 5 soru
64 kontrol seçimi sorusu
17 denetim kanıtı sorusu
19 denetim yargısı sorusu… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_cyber_security_controls_benchmark.oauth-api-security-en
OAuth & API Security Dataset (EN)
Comprehensive English dataset covering OAuth 2.0 vulnerabilities, API attacks (OWASP API Top 10 2023), security controls, and Q&A pairs for training cybersecurity-specialized language models.
Dataset Contents
Category
Entries
Description
OAuth 2.0 Vulnerabilities
20
Authorization Code Interception, CSRF, PKCE bypass, JWT attacks, token leakage
API Attacks
25
BOLA, BFLA, BOPLA, SSRF, GraphQL DoS, gRPC injection, CORS… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/oauth-api-security-en.secops-es-benchmark
secops-es-benchmark
An open benchmark for AI agents that investigate breaches in Elasticsearch. Real,
labeled attack telemetry in Elasticsearch (ECS) plus a scored task suite — so you can measure
how well an agent, or any LLM, does the job of a SOC analyst: follow one alert, pivot across
data sources with ES|QL, reconstruct the intrusion, and recommend a proportionate response.
DOI (cite this): 10.5281/zenodo.21770551
Code, runner, and full docs:… See the full description on the dataset page: https://huggingface.co/datasets/TocharianOU/secops-es-benchmark.gaap-sec-compliance-dataset
GAAP & SEC Compliance Dataset
A comprehensive dataset for financial AI applications
Dataset Overview
This dataset contains 470,151 documents covering US GAAP (Generally Accepted Accounting Principles) standards and SEC (Securities and Exchange Commission) filing requirements. It's designed for training and evaluating AI systems for financial compliance, accounting Q&A, and regulatory analysis.
Key Statistics
Total Documents: 470,151
Average Length: 363… See the full description on the dataset page: https://huggingface.co/datasets/aanshshah/gaap-sec-compliance-dataset.dspy-security-bench-v01-results
dspy-security-bench: v0.1 + v0.1.1 results
Raw evaluation outputs from
dspy-security-bench.
Cite or audit these numbers without needing to clone the repo or re-run
the benchmark.
What's in here
File
Contents
Rows
workspace_v01_results.csv
Original v0.1 launch run. Workspace suite, 3 optimizers (unoptimized, BootstrapFewShot, MIPROv2 light), 2 attacks (direct, important_instructions), N=5 user × 1 injection × 1 seed.
30
workspace_v01_summary.csv
v0.1… See the full description on the dataset page: https://huggingface.co/datasets/immu4989/dspy-security-bench-v01-results.oauth-api-security-fr
Dataset OAuth & Securite API (FR)
Dataset francophone complet sur les vulnerabilites OAuth 2.0, les attaques API (OWASP API Top 10 2023), les controles de securite, et les questions-reponses pour l'entrainement de modeles de langage specialises en cybersecurite.
Contenu du Dataset
Categorie
Nombre d'entrees
Description
Vulnerabilites OAuth 2.0
20
Authorization Code Interception, CSRF, PKCE bypass, JWT attacks, token leakage
Attaques API
25
BOLA, BFLA, BOPLA… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/oauth-api-security-fr.
