datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models.
These traces focus on security audits of opensource software.
Sharing traces with Swival
Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session:
swival "Fix the login bug" --trace-dir traces/
Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.security_instruct_mcq_2481cyber-security
Cybersecurity AI Knowledge Base — PhD-Level Dataset
Overview
This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security.
Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms
Purpose
Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.sec-contracts-financial-extraction-instructions
S&P 500 SEC Financial Extraction Instructions
Dataset Summary
7,683 instruction-tuning examples for training LLMs to extract structured financial data from SEC filings. Covers two filing types across S&P 500 companies:
Split
Examples
Filing Type
Description
train
3,430
Exhibit 10 + DEF 14A
Positive examples with validated outputs
corrective
4,253
Exhibit 10 + DEF 14A
Corrective, rescued, and negative examples
Exhibit 10 — Material Contracts (2… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-financial-extraction-instructions.SEC-bench
Data Instances
instance_id: (str) - A unique identifier for the instance
repo: (str) - The repository name including the owner
project_name: (str) - The name of the project without owner
lang: (str) - The programming language of the repository
work_dir: (str) - Working directory path
sanitizer: (str) - The type of sanitizer used for testing (e.g., Address, Memory, Undefined)
bug_description: (str) - Description of the vulnerability
base_commit: (str) - The base commit hash where the… See the full description on the dataset page: https://huggingface.co/datasets/SEC-bench/SEC-bench.Code_Vulnerability_Security_DPO
Cybernative.ai Code Vulnerability and Security Dataset
Dataset Description
The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen on… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative/Code_Vulnerability_Security_DPO.warwick-second-life-dm-2025-raw
First-life and second-life battery degradation mode test data
BSEBench status: raw_mirror_pending_validation
This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.SecEval
SecEval: A Comprehensive Benchmark for Evaluating Cybersecurity Knowledge of Foundation Models
The advent of large language models has ignited a transformative era for the cybersecurity industry. Pioneering applications are being developed, deployed, and utilized in areas such as cybersecurity knowledge QA, vulnerability hunting, and alert investigation. Various researches have indicated that LLMs primarily acquire their knowledge during the pretraining phase, with fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/XuanwuAI/SecEval.infosec-security-qadavinciDecodingTrust
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Overview
This repo contains the source code of DecodingTrust. This research endeavor is designed to help researchers better understand the capabilities, limitations, and potential risks associated with deploying these state-of-the-art Large Language Models (LLMs). See our paper for details.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Boxin Wang, Weixin Chen, Hengzhi… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust.affelnet-paris-secteursCe dépôt contient les secteurs entre les collèges et lycées parisiens
Affelnet (Affectation des élèves par le Net) est la procédure informatisée utilisée en France pour affecter les élèves de 3ème dans un lycée de secteur pour leur année de Seconde. L'affectation se base sur un score qui prend en compte les résultats scolaires, la sectorisation géographique, le statut de boursier, et des bonus spécifiques comme le bonus IPS (Indice de Positionnement Social).
Description des jeux de… See the full description on the dataset page: https://huggingface.co/datasets/fgaume/affelnet-paris-secteurs.terraform_sec
Terraform Security Dataset
A comprehensive dataset of 62,406 Terraform projects analyzed for security vulnerabilities using tfsec. This dataset is designed for training Large Language Models (LLMs) to understand, identify, and fix security issues in Terraform infrastructure-as-code.
📊 Dataset Overview
Total Examples: 62,406 Terraform projects
Secure Projects: 43,575 (69.8%)
Insecure Projects: 18,831 (30.2%)
Format: JSONL (JSON Lines)
Task: Security analysis and… See the full description on the dataset page: https://huggingface.co/datasets/galcan/terraform_sec.starcoderdatapes2o-v3SecQue
SECQUE
Paper
SECQUE is a comprehensive benchmark for evaluating large language models (LLMs) in financial analysis tasks.
SECQUE comprises 565 expert-written questions covering SEC filings analysis across four key categories:
comparison analysis
ratio calculation
risk assessment
financial insight generation.
To assess model performance, we develop SECQUE-Judge, an evaluation mechanism leveraging multiple LLM-based judges, which demonstrates strong alignment with human… See the full description on the dataset page: https://huggingface.co/datasets/nogabenyoash/SecQue.exams-basic-and-quantum-cryptography-and-security-latex
Open Problem Exams: Cryptography and Security (LaTeX)
A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions.
Dataset Overview
Institution
Files
Topics
Questions
Caltech & TU Delft
8
38
145
EPFL
6
19
86
ETH Zurich
1
14
37
MIT
3
33
79
Total
18
104
347
Difficulty Distribution
Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.secret-loyalty-competition-data
Secret-loyalty organisms — training banks and eval batteries
The data behind KKing23/secret-loyalty-competition-organisms.
Code and full result trail: github.com/kaustubhkislay/secret-loyalty-competition.
Why this exists separately from the adapters. The adapters are reproducible from these
banks for the price of GPU time. These banks are not reproducible — they were written by
an LLM generator, so regenerating gives different data and every published number becomes… See the full description on the dataset page: https://huggingface.co/datasets/KKing23/secret-loyalty-competition-data.indian-legal-sections-bns-bnss-bsa-2023
🏛️ Indian Legal Sections — BNS · BNSS · BSA 2023
The First Structured, Unified JSON Dataset of Modern Indian Criminal Law
📖 Dataset Summary
This dataset contains 1,059 fully structured and verified sections extracted, parsed, and unified from India's three landmark criminal justice reform acts passed in December 2023. These three acts together replaced the colonial-era Indian Penal Code (IPC, 1860), the Code of Criminal Procedure… See the full description on the dataset page: https://huggingface.co/datasets/GSMS-B/indian-legal-sections-bns-bnss-bsa-2023.clawhub-security-signals
ClawHub Security Signals
🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print
ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale.
This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree.
Paper snapshot: this… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals.SecurityQuestionsSecurityEval
Dataset Card for SecurityEval
This dataset is from the paper titled SecurityEval Dataset: Mining Vulnerability Examples to Evaluate Machine Learning-Based Code Generation Techniques.
The project is accepted for The first edition of the International Workshop on Mining Software Repositories Applications for Privacy and Security (MSR4P&S '22).
The paper describes the dataset for evaluating machine learning-based code generation output and the application of the dataset to the code… See the full description on the dataset page: https://huggingface.co/datasets/s2e-lab/SecurityEval.szl-second-brain-inrepo
SZL Second Brain — in-repo lane (public projection)
This is DATA, not a model. A retrieval corpus — the in-repo lane of the
SZL Second Brain — published as a Hugging Face dataset. It is part of the
SZL Forge Family. Publishing it is repo-existence of a corpus only: it
trains nothing, evaluates nothing, serves no inference, and upgrades nothing
(Λ stays Conjecture-1).
What this is
A deterministic BM25-ready corpus assembled only from repo-public text of
the Alloy… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-second-brain-inrepo.open-web-mathSecureCodePairs
Dataset Summary
Field
Value
Version
1.2.0
License
MIT
Total code examples
470
LLM security trajectories
30
Languages (15)
Python, Java, JavaScript, TypeScript, Go, PHP, C#, Kotlin, Swift, Rust, Ruby, C, C++, Scala, YAML (Kubernetes)
Frameworks
Flask, Django, FastAPI, Spring Boot, Express, NestJS, Next.js, Laravel, ASP.NET Core, Gin, Android, iOS, Actix, Rails, Qt, Play, gRPC, GraphQL, Kubernetes
New in v1.2.0
+260 records (deep Python/Java packs… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/SecureCodePairs.pstu-synthetic-secrets
PSTU Synthetic Secrets Dataset
Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper:
Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal
Hoda Fakhar — ECML PKDD 2026
Dataset Description
175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric.
All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.Open-SECQA
Open-SECQA
Open-domain financial QA benchmark (a.k.a. LOFin) built on 145,897 SEC filings from 516 S&P 500 companies (Oct 2001 – Apr 2025), with 1,595 QA pairs covering single-document, multi-document, and multi-hop reasoning.
📄 Paper: ACL 2025 Findings
💻 Code: LOFin-bench-HiREC
Composition
Source
# QAs
FinQA
1,112
SEC-QA
333
FinanceBench
150
Total
1,595
Citation
@inproceedings{choe-etal-2025-hierarchical,
title =… See the full description on the dataset page: https://huggingface.co/datasets/HYdsl/Open-SECQA.cairo-security-audits
Cairo Security Audits
A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations.
Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.llm-trustworthy-leaderboard-resultsfineweb-edu
