datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GenIaC-SecBench
GenIaC-SecBench
A benchmark for evaluating the security of LLM-generated Infrastructure-as-Code
(IaC) against a size-matched human baseline.
Paper: Compared to What? A Human-Anchored Security Benchmark for LLM-Generated
Infrastructure-as-Code (arXiv:2608.28021)
Code: https://github.com/AnimeshShaw/GenIaC-SecBench
Why this dataset exists
Prior evaluations of generated IaC report vulnerability counts for models
only. Stating that a model averages eight findings per… See the full description on the dataset page: https://huggingface.co/datasets/AnimeshShaw/GenIaC-SecBench.b3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.ai-agent-security-incidents
AI Agent Security Incident Database v0.1
A structured, machine-readable database of 1392 confirmed AI agent security incidents, collected and classified automatically.
What is this?
Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it.
This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.cbam-sector-facility-registry
CBAM-Sector Global Facility Registry
Open screening registry of cement, iron & steel and aluminium facilities worldwide in three CBAM Annex I good categories, with modelled CO2 (Climate TRACE) and regulator-reported CO2 (EU ETS EUTL / US EPA GHGRP) kept in separate columns, each reported figure carrying its match evidence.
Canonical record: doi.org/10.5281/zenodo.22172573 · Publisher: Inzonex
Load
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Inzinion/cbam-sector-facility-registry.us-layoffs-by-industry-sector-warn-act
US layoffs by industry sector — 60,955 WARN Act notices, 1988-2026, sector per employer
Rebuilt 2026-09-16. 32,042 of 60,955 dated notices (52.6%; 61,330 on record, 375 lack a usable date) carry a sector; the
rest are unclassified and stay in every total. In 2026 so far the largest sector by
reported workers is Logistics, transport & warehousing (23,761 workers, 193 notices);
in the last 90 days it is Healthcare & medical (7,475 workers).
No state WARN portal publishes an… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-industry-sector-warn-act.13f-institutional-holdings-sec-edgar
13F Institutional Holdings Dataset — SEC EDGAR Hedge Fund & Asset Manager Filings
A structured, ready-to-analyze snapshot of institutional 13F filings covering 13,000+ investment managers — hedge funds, mutual fund families, pension funds, banks, and family offices — built from raw SEC Form 13F data on EDGAR. Each row is one manager's most recently disclosed quarter: total portfolio value, position count, and five behavioral scores (concentration, turnover, momentum/contrarian… See the full description on the dataset page: https://huggingface.co/datasets/JamesFromAlphasmo/13f-institutional-holdings-sec-edgar.ai-5node-key-buf-lag-cpl-secret-leak-v0.1
What this repo does
This dataset models secret leakage cascades in AI agent operations. It detects when secret exposure risk rises, protective buffers weaken, governance lag delays revoke and purge actions, and tight coupling through shared logs, tickets, and tool chains crosses the five-node cascade threshold into an unrecoverable secret leakage cascade.
This dataset models a five-node cascade: four interacting instability drivers and one emergent cascade state.The fifth node… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-key-buf-lag-cpl-secret-leak-v0.1.sec-8k-market-reaction-dataset
Historical SEC 8-K Market-Reaction Dataset — by FlinchLab
Know what each kind of company news historically did to the stock —
before you build on top of it. 29,331 SEC 8-K filings from 660 US
companies, read and re-classified finer than the official item codes, with
each category's price reaction measured by a market-model event study:
direction, magnitude, overnight-vs-session split, pre-filing baselines —
and every published finding validated on two independent out-of-sample… See the full description on the dataset page: https://huggingface.co/datasets/Flinchlab/sec-8k-market-reaction-dataset.sec-13f-holdings
SEC Form 13F Hedge Fund Holdings
Quarterly US-listed equity holdings for 9 institutional managers,
reconstructed from their own Form 13F-HR filings with the SEC.
Built for the trackers at y-yin.io/research and
published here because the filings are public domain and the parsing is fiddly
enough to be worth sharing. Read straight from each filing's infotable.xml —
the structured document EDGAR renders its own filing pages from — so no HTML is
scraped.
Coverage: 9 funds, 452… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/sec-13f-holdings.FinRL_BTC_news_signals
Overview
This news dataset is created for FinAI Contest 2025 Task 1 FinRL-DeepSeek for Crypto Trading. We collected BTC news for the training and testing period from different sources [1] [2]. For each news, we use the DeepSeek chat model to extract the sentiment score, risk level, and their correpsonding confidence level and one-sentence reasoning.
Column
Description
date_time
Timestamp of when the news article was published (in UTC).
title
Title of the news article.… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/FinRL_BTC_news_signals.Cyber-Security-Breachesvn-provinces-upper-secondary-graduation-rate
Vietnam provinces upper-secondary graduation rate
Vietnam provinces upper-secondary graduation rate. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (1009 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (66 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-upper-secondary-graduation-rate.secondary-screen-dose-response-curve-parameters
Citation
DepMap, Broad; Corsello, Steven; Kocak, Mustafa; Golub, Todd (2019). PRISM Repurposing 19Q4 Dataset. figshare. Dataset.
Current dataset: https://doi.org/10.6084/m9.figshare.9393293.v4
General guidance: https://doi.org/10.1101/730119
Dataset specific README: prism_repurposing_secondary
Dataset Fields
broad_id: ID used to identify drug-batch combinations
name: Name of the drug
depmap_id: ID to identify a cell line
ccle_name: ID to identify a cell line… See the full description on the dataset page: https://huggingface.co/datasets/donb-hf/secondary-screen-dose-response-curve-parameters.us-equity-valuation-cross-section-with-brina-gap
Brina Gap cross-section, US-listed equities
Per-company valuation measures for the covered universe of US-listed companies,
computed from SEC filings. Snapshot taken 2026-08-05.
Why this dataset exists
The Brina Gap is the difference between the growth rate a business can fund from
its own returns (ROIC times reinvestment rate) and the growth rate its current
enterprise value implies, recovered by reverse DCF. The second of those numbers,
market-implied growth per… See the full description on the dataset page: https://huggingface.co/datasets/Zyberno/us-equity-valuation-cross-section-with-brina-gap.africa-synth-employment-informal-sector-employment-africa-all
Africa Synth Employment Informal Sector Employment Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-employment-informal-sector-employment-africa-all.babbelphish
BabbelPhish
BabbelPhish is a dataset based on the Sublime Security Message Query Language (MQL) used for email security detection engineering. This dataset is specially created for the BabbelPhish project, which focuses on leveraging large language models to facilitate the work of detection engineers.
This dataset comprises around 3,000 examples drawn from various sources. We've utilized the following:
Sublime Security Documentation
Message Data Model (Schema)
Sublime Rules Repo… See the full description on the dataset page: https://huggingface.co/datasets/sublime-security/babbelphish.africa-synth-employment-formal-sector-jobs-africa-all
Africa Synth Employment Formal Sector Jobs Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-employment-formal-sector-jobs-africa-all.sec-edgar-geographic-revenue-breakdowns
US S&P 500 Companies Geographic Revenue Exposure (SEC EDGAR)
This dataset contains a comprehensive, reconciled, and audited map of the geographic and regional revenue breakdowns for major US-listed corporations (including S&P 500 companies). The data was extracted directly from corporate 10-K filings submitted to the US Securities and Exchange Commission (SEC) EDGAR system.
By reconciling structured SEC XBRL segment dimensions with unstructured HTML R-file disclosures (using the… See the full description on the dataset page: https://huggingface.co/datasets/Metricshour/sec-edgar-geographic-revenue-breakdowns.nervous-section-ccc5e3
nervous-section-ccc5e3
Synthetic products test data: 32 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Indigo-Owen/nervous-section-ccc5e3.japanchoice-2026-foreign-policy-and-securityExported: 2026-02-20
This dataset covers the topic of Foreign Policy and Security (外交・安全保障) from the Japan Choice Polis platform. The conversation can be found at: https://polis.japanchoice.jp/3x3ancmrtc
Data was exported from the database and provided by Nishio Hirokazu. It was gathered using the Polis software by The Computational Democracy Project (see: compdemocracy.org/polis and github.com/compdemocracy/polis).
us-social-security-medicare-FAQs-testtoulouse-sections-preview
Toulouse Cadastral Sections — Price & Transit Preview
Free preview of a much richer per-section dataset covering all 356 cadastral sections of
Toulouse. This file has 3 metrics that are individually reconstructable from public sources
(sale price, metro distance, priority-district status) — the full dataset (37 columns:
schools + social index, healthcare, safety, noise, demographics, PLU zoning risk, and 5
persona-specific 0–100 scores) is a paid product — see Get the full… See the full description on the dataset page: https://huggingface.co/datasets/chaosskill/toulouse-sections-preview.SeCoDa
SeCoDa
Repository for the Sense Complexity Dataset (SeCoDa)
Paper
For more information on the SeCoDa, see the paper.
Publications using this dataset must include a reference to the following publication:
SeCoDa: Sense Complexity Dataset. David Strohmaier, Sian Gooding, Shiva Taslimipoor, Ekaterina Kochmar. Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020), pages 5964–5969, Marseille, 11–16 May 2020
The dataset is based on the earlier… See the full description on the dataset page: https://huggingface.co/datasets/dstrohmaier/SeCoDa.US_Social_Security_Medicare_FAQs_Samplesecure-item-75220e
secure-item-75220e
Synthetic sensors test data: 37 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/stephanie24/secure-item-75220e.secure-clue-738231
secure-clue-738231
Synthetic products test data: 31 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/xyamazaki/secure-clue-738231.Preprocessed-CVS-24-KMRjapanchoice-2025-foreign-policy-and-securityFetched: 2026-02-16-2245
Data was gathered using the Polis software (see: compdemocracy.org/polis and github.com/compdemocracy/polis), and acquired from this URL: https://polis.japanchoice.jp/7cdcyjmsyh
comments.csv was augmented with is-seed and is-meta columns fetched from https://polis.japanchoice.jp/api/v3/comments?conversation_id=7cdcyjmsyh&moderation=true&include_voting_patterns=true
real-world-benign-use-cases
Real-World Benign Use Cases
A curated set of 178 real-world, 100%-benign examples (label == 0 for every row) pulled from
production AI-coding-agent traffic — chat messages, tool output, shell commands, code snippets —
built specifically to stress-test prompt-injection / jailbreak classifiers for false positives.
Every row was independently judged benign with high confidence before inclusion. This is not a
random sample of production traffic: rows were preferentially drawn from… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/real-world-benign-use-cases.social_security_embeddings
