datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-complete
arXiv Complete Corpus
A snapshot of arXiv's metadata, version history, submission files and rendered
documents. It covers 3,148,796 papers and includes file contents, paths, sizes
and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface;
files come from the GCS mirror, S3 source archives and direct PDF fetches.
This release holds a PDF for 99.47% of papers and 99.54% of versions reported
with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models.
These traces focus on security audits of opensource software.
Sharing traces with Swival
Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session:
swival "Fix the login bug" --trace-dir traces/
Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.dojo_sector_info
Languages: 简体中文 · English
dojo_sector_info — Sector Taxonomy
Overview
Three-level sector taxonomy (L1 industry → L2 chain → L3 sub-sector) with bilingual Chinese/English names and definitions.
Files
File
Description
data.parquet
Taxonomy tree (one L1 row each; L2/L3 nested in children)
Key Fields
Field
Description
id
L1 sector ID
name / name_alias
L1 English name / Chinese alias
description /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_info.SEC
SEC Annual Reports (Form 10-K) 1993-2024
Dataset Overview
This dataset comprises SEC annual reports (Form 10-K) for the years 1993 to 2024, providing comprehensive coverage of publicly traded companies' financial and business information. The reports are stored in Parquet format, ensuring efficient storage and quick access. This dataset was meticulously compiled using the EDGAR-Crawler toolkit, which facilitates the extraction and processing of SEC filings from the EDGAR… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/SEC.financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system.
Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences.
Sentiment labels are provided on a per filing basis from the market reaction around the filing data.
Additional metadata for each filing is included in the dataset.warwick-second-life-dm-2025-raw
First-life and second-life battery degradation mode test data
BSEBench status: raw_mirror_pending_validation
This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.sec-nport
SEC Form N-PORT Data Sets
Monthly portfolio holdings reported by registered investment companies and ETFs
on Form N-PORT, published by the U.S. SEC as quarterly structured data sets
and mirrored here as typed, partitioned Parquet — queryable directly from DuckDB.
Source: SEC Form N-PORT Data Sets — public domain (U.S. Government work)
https://www.sec.gov/data-research/sec-markets-data/form-n-port-data-sets
Coverage: October 2019 onward, refreshed quarterly
Format: one Parquet… See the full description on the dataset page: https://huggingface.co/datasets/trader298/sec-nport.GenIaC-SecBench
GenIaC-SecBench
A benchmark for evaluating the security of LLM-generated Infrastructure-as-Code
(IaC) against a size-matched human baseline.
Paper: Compared to What? A Human-Anchored Security Benchmark for LLM-Generated
Infrastructure-as-Code (arXiv:2608.28021)
Code: https://github.com/AnimeshShaw/GenIaC-SecBench
Why this dataset exists
Prior evaluations of generated IaC report vulnerability counts for models
only. Stating that a model averages eight findings per… See the full description on the dataset page: https://huggingface.co/datasets/AnimeshShaw/GenIaC-SecBench.b3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.sec-material-contracts-qa800+ EDGAR contracts with PDF images and key information extracted by the OpenAI GPT-4o model.
The key information is defined as follows:
class KeyInformation(BaseModel):
agreement_date : str = Field(description="Agreement signing date of the contract. (date)")
effective_date : str = Field(description="Effective date of the contract. (date)")
expiration_date : str = Field(description="Service end date or expiration date of the contract. (date)")
party_address : str =… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts-qa.DecodingTrust
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Overview
This repo contains the source code of DecodingTrust. This research endeavor is designed to help researchers better understand the capabilities, limitations, and potential risks associated with deploying these state-of-the-art Large Language Models (LLMs). See our paper for details.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Boxin Wang, Weixin Chen, Hengzhi… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust.affelnet-paris-secteursCe dépôt contient les secteurs entre les collèges et lycées parisiens
Affelnet (Affectation des élèves par le Net) est la procédure informatisée utilisée en France pour affecter les élèves de 3ème dans un lycée de secteur pour leur année de Seconde. L'affectation se base sur un score qui prend en compte les résultats scolaires, la sectorisation géographique, le statut de boursier, et des bonus spécifiques comme le bonus IPS (Indice de Positionnement Social).
Description des jeux de… See the full description on the dataset page: https://huggingface.co/datasets/fgaume/affelnet-paris-secteurs.sec-10k-lsh-chunks
📈 SEC 10-K Cleaned Text Chunks & LSH Boilerplate Dataset
Dataset Summary
This dataset contains 13,562,130 cleaned text chunks extracted from 12,361 SEC Form 10-K annual filings across 1,380 companies (spanning 2004 to 2025, core 2014–2025).
Every chunk across all 1,380 companies (including mega-cap leaders such as AAPL, MSFT, NVDA, AMZN, GOOGL, META, TSLA, JPM, WMT, XOM, AVGO, LLY) is annotated with metadata, token counts, table indicators, and a pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-lsh-chunks.cvefixes-security-ir-graphrag
CVEfixes Security IR GraphRAG
This is a content-addressed, remotely routable Security IR release of the exact pinned CVEfixes snapshot. It packages the byte-identical original Parquet data together with a searchable corpus, BM25 postings, CUDA-generated vectors, typed graph nodes and edges, bounded adjacency indexes, and a source-CID-to-original-row lookup.
All entries are inert, non-authoritative evidence. Candidate and formal-logic rows cannot grant execution authority; exact… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-ir-graphrag.ai-agent-security-incidents
AI Agent Security Incident Database v0.1
A structured, machine-readable database of 1365 confirmed AI agent security incidents, collected and classified automatically.
What is this?
Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it.
This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.cbam-sector-facility-registry
CBAM-Sector Global Facility Registry
Open screening registry of cement, iron & steel and aluminium facilities worldwide in three CBAM Annex I good categories, with modelled CO2 (Climate TRACE) and regulator-reported CO2 (EU ETS EUTL / US EPA GHGRP) kept in separate columns, each reported figure carrying its match evidence.
Canonical record: doi.org/10.5281/zenodo.22172573 · Publisher: Inzonex
Load
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Inzinion/cbam-sector-facility-registry.lekiwi_second_floor_0915_environmentThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi",
"total_episodes": 50,
"total_frames": 14471,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yuzhous/lekiwi_second_floor_0915_environment.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.lekiwi_second_floor_0916_environmentThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi",
"total_episodes": 50,
"total_frames": 15084,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yuzhous/lekiwi_second_floor_0916_environment.SEC_filings_1994_2024
Dataset Card for SEC EDGAR Filings Master Index
Dataset Details
Dataset Description
This dataset contains metadata for all submissions to the Securities and Exchange Commission (SEC) through their EDGAR system from 1994 until December 14, 2024. The data is extracted from quarterly master files and includes key information about company filings such as CIK numbers, company names, form types, and filing dates.
Curated by: Arthur (arthur@cicero.chat)
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC_filings_1994_2024.secondary-sources
NuBerea/secondary-sources
Second Temple Jewish secondary sources in Greek: the complete extant Greek
corpora of Flavius Josephus (Jewish Antiquities, Jewish War, Vita, Contra
Apionem) and Philo of Alexandria (all 31 works), segmented for scholarly
text-retrieval and lexical-semantic study. These two first-century authors are
the principal non-biblical Jewish witnesses to the Second Temple period and its
milieu, and this repository serves as the Second Temple companion corpus to… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/secondary-sources.us-layoffs-by-industry-sector-warn-act
US layoffs by industry sector — 60,955 WARN Act notices, 1988-2026, sector per employer
Rebuilt 2026-09-16. 32,042 of 60,955 dated notices (52.6%; 61,330 on record, 375 lack a usable date) carry a sector; the
rest are unclassified and stay in every total. In 2026 so far the largest sector by
reported workers is Logistics, transport & warehousing (23,761 workers, 193 notices);
in the last 90 days it is Healthcare & medical (7,475 workers).
No state WARN portal publishes an… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-industry-sector-warn-act.clawhub-security-signals
ClawHub Security Signals
🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print
ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale.
This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree.
Paper snapshot: this… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals.13f-institutional-holdings-sec-edgar
13F Institutional Holdings Dataset — SEC EDGAR Hedge Fund & Asset Manager Filings
A structured, ready-to-analyze snapshot of institutional 13F filings covering 13,000+ investment managers — hedge funds, mutual fund families, pension funds, banks, and family offices — built from raw SEC Form 13F data on EDGAR. Each row is one manager's most recently disclosed quarter: total portfolio value, position count, and five behavioral scores (concentration, turnover, momentum/contrarian… See the full description on the dataset page: https://huggingface.co/datasets/JamesFromAlphasmo/13f-institutional-holdings-sec-edgar.SEC_2025sec-8k-events
SEC Form 8-K Corporate Events
Every Form 8-K filed since the modern item taxonomy took effect — and, for each
one, the second the SEC accepted it, which is not the date printed on it.
1 761 353 filings · 3 676 835 item-level events · 23 August 2004 to today
The pipeline lives in recipe/ at the same revision as the data.
See PIPELINE.md for the method.
The problem this dataset exists to solve
Apple filed its June-quarter results on 30 July 2026. Here is the filing… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/sec-8k-events.SecVulEval
Dataset Card for Dataset Name
SecVulEval is a collection of real-world C/C++ vulnerabilities.
Dataset Details
Dataset Description
The dataset is curated by collecting C/C++ vulnerability from NVD. It features statement-level vulnerable information, context information for vulnerable functions
(is_vulnerable=True), and other metadata such as CVE, CWE, commit information. The dataset contains vulnerable and non-vulnerable function samples.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/arag0rn/SecVulEval.cyber-security-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges).
Available in two sizes:
witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.so100_secondThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 22448,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Gano007/so100_second.sec-filings-forward-return-2026
SEC Filings → Forward-Return (US large-cap, 2000–2026)
Leak-free, time-ordered SEC filings (8-K / 10-Q / 10-K) with objective forward-return
labels for ~610 current US large-cap names, through 2026. A benchmark ("can filing text
predict forward returns?"), not just a corpus.
⚖️ Licensing & provenance (read first — this is the honest part)
This repo deliberately separates two provenance classes:
part
source
license
Filing text + accession,cik,ticker,form… See the full description on the dataset page: https://huggingface.co/datasets/BatuhanECB/sec-filings-forward-return-2026.
