datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-complete
arXiv Complete Corpus
A snapshot of arXiv's metadata, version history, submission files and rendered
documents. It covers 3,148,796 papers and includes file contents, paths, sizes
and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface;
files come from the GCS mirror, S3 source archives and direct PDF fetches.
This release holds a PDF for 99.47% of papers and 99.54% of versions reported
with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.dojo_sector_precomputed
Languages: 简体中文 · English
dojo_sector_precomputed — Precomputed Sector Analytics
Overview
Derived sector analytics: L3 constituent snapshots, daily cap-weighted sector index levels, and per-constituent daily returns. Built offline from taxonomy, mappings, quotes, and stock K-lines.
Files
File
Description
manifest.json
Generation metadata: version, window start, row counts, latest trade dates
constituents.parquet
L3 constituent… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_precomputed.dojo_sector_info
Languages: 简体中文 · English
dojo_sector_info — Sector Taxonomy
Overview
Three-level sector taxonomy (L1 industry → L2 chain → L3 sub-sector) with bilingual Chinese/English names and definitions.
Files
File
Description
data.parquet
Taxonomy tree (one L1 row each; L2/L3 nested in children)
Key Fields
Field
Description
id
L1 sector ID
name / name_alias
L1 English name / Chinese alias
description /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_info.dojo_sector_symbol_relations
Languages: 简体中文 · English
dojo_sector_symbol_relations — Stock–Sector Mapping
Overview
Maps each stock to L1/L2/L3 sector paths with primary and secondary assignments. One row per (ticker, market) pair.
Files
File
Description
data.parquet
Full stock ↔ sector relations
Key Fields
Field
Description
ticker
Stock symbol
market
us, cn, or hk
primary
JSON object — primary sector path
secondary
JSON array —… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_symbol_relations.SEC
SEC Annual Reports (Form 10-K) 1993-2024
Dataset Overview
This dataset comprises SEC annual reports (Form 10-K) for the years 1993 to 2024, providing comprehensive coverage of publicly traded companies' financial and business information. The reports are stored in Parquet format, ensuring efficient storage and quick access. This dataset was meticulously compiled using the EDGAR-Crawler toolkit, which facilitates the extraction and processing of SEC filings from the EDGAR… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/SEC.the-secrets-of-ceos-book-2k
The-Secrets-Of-Ceos-Book-2k
Made with ❤️ using 🦥 Unsloth Studio
Beta2x was generated with Unsloth Recipe Studio. It contains 2,000 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("filipwx/the-secrets-of-ceos-book-2k", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 2,000
📋 Columns: 4
📋 Schema & Statistics
Column
Type
Column Type
Unique… See the full description on the dataset page: https://huggingface.co/datasets/filipwx/the-secrets-of-ceos-book-2k.securecode-web
SecureCode Web: Traditional Web & Application Security Dataset
Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance
Paper | GitHub | Dataset | Model Collection | Blog Post
What's new in v2.6
v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had
shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/scthornton/securecode-web.White-Hat-Security-Agent-Prompts-600K
White Hat Security Agent Prompts 600K
Overview
The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios.
Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.10-K_sec_filings
Dataset Card for "10-K_sec_filings"
Dataset of 93.5K 10K SEC EDGAR filings since 1999 year. This dataset contains a lot of bad parsed filings and also empty rows
More Information needed
sec-material-contracts
Material Contracts (Exhibit 10) from SEC/EDGAR
Because sometimes you need 1,141,632 examples of corporate legalese to train your next model ☕
Dataset Summary
Picture this: 1,141,632 material contracts (Exhibit 10) painstakingly collected from sec.gov's EDGAR database. We're talking about legal agreements spanning from 1994 to 2025 Q1, sourced from 10-K, 10-Q, and 8-K filings. Think of Exhibit 10 as the treasure trove where companies hide their most important legal… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts.adv_glue
Dataset Card for Adversarial GLUE
Dataset Summary
Adversarial GLUE Benchmark (AdvGLUE) is a comprehensive robustness evaluation benchmark that focuses on the adversarial robustness evaluation of language models. It covers five natural language understanding tasks from the famous GLUE tasks and is an adversarial version of GLUE benchmark.
AdvGLUE considers textual adversarial attacks from different perspectives and hierarchies, including word-level transformations… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/adv_glue.prompt-injections-benchmark
Dataset: Qualifire Benchmark Prompt Injection(Jailbreak vs. Benign) Datasets
Overview
This dataset contains 5,000 prompts, each labeled as either jailbreak or benign. The dataset is designed for evaluating AI models' robustness against adversarial prompts and their ability to distinguish between safe and unsafe inputs.
Dataset Structure
Total Samples: 5,000
Labels: jailbreak, benign
Columns:
text: The input text
label: The classification (jailbreak or benign)… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/prompt-injections-benchmark.sec-10k-lsh-chunks
📈 SEC 10-K Cleaned Text Chunks & LSH Boilerplate Dataset
Dataset Summary
This dataset contains 13,562,130 cleaned text chunks extracted from 12,361 SEC Form 10-K annual filings across 1,380 companies (spanning 2004 to 2025, core 2014–2025).
Every chunk across all 1,380 companies (including mega-cap leaders such as AAPL, MSFT, NVDA, AMZN, GOOGL, META, TSLA, JPM, WMT, XOM, AVGO, LLY) is annotated with metadata, token counts, table indicators, and a pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-lsh-chunks.sections
Dataset Card for "sections"
More Information needed
sec-material-contracts-qa800+ EDGAR contracts with PDF images and key information extracted by the OpenAI GPT-4o model.
The key information is defined as follows:
class KeyInformation(BaseModel):
agreement_date : str = Field(description="Agreement signing date of the contract. (date)")
effective_date : str = Field(description="Effective date of the contract. (date)")
expiration_date : str = Field(description="Service end date or expiration date of the contract. (date)")
party_address : str =… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts-qa.africa-world-bank-governance-public-sector-time-series
Africa World Bank Governance and Public Sector Labeled Time Series Data
This repository is part of the Africa Temporal Intelligence Corpus (ATIC). It contains sector-specific temporal corpus packages for African countries.
ATIC sector repositories are designed for machine consumption first: Parquet tables, stable IDs, reproducible metadata, explicit provenance, review status, and separable semantic layers.
Sector Scope
Temporal governance, fragility, conflict, and… See the full description on the dataset page: https://huggingface.co/datasets/africatic/africa-world-bank-governance-public-sector-time-series.humanitys-second-last-exam
Humanity's Second Last Exam
Benchmark design, curation and release maintenance: Shashank Agnihotri.
Original questions retain their recorded authorship and source attribution.
This owner-reviewed retained release contains 365 target questions, 730
context examples, and 365 ordered target/A/B links: 1,095 question rows.
The owner review concluded on 16 September 2026. This is an owner-reviewed
release after suspected-AI-content exclusions, not a software-certified
guarantee of… See the full description on the dataset page: https://huggingface.co/datasets/shashankskagnihotri/humanitys-second-last-exam.SEC_filings_1994_2024
Dataset Card for SEC EDGAR Filings Master Index
Dataset Details
Dataset Description
This dataset contains metadata for all submissions to the Securities and Exchange Commission (SEC) through their EDGAR system from 1994 until December 14, 2024. The data is extracted from quarterly master files and includes key information about company filings such as CIK numbers, company names, form types, and filing dates.
Curated by: Arthur (arthur@cicero.chat)
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC_filings_1994_2024.SecureVibeBench
SecureVibeBench: First Secure Vibe Coding Benchmark
SecureVibeBench is a benchmark consisting of 105 C/C++ secure coding tasks sourced from 41 projects in OSS-Fuzz for code agents. It is designed to evaluate secure vibe coding by reconstructing real-world scenarios where human developers introduced vulnerabilities.
Paper: SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
Repository: iCSawyer/SecureVibeBench
Venue:… See the full description on the dataset page: https://huggingface.co/datasets/iCSawyer/SecureVibeBench.cvefixes-security-ir-graphrag
CVEfixes Security IR GraphRAG
This is a content-addressed, remotely routable Security IR release of the exact pinned CVEfixes snapshot. It packages the byte-identical original Parquet data together with a searchable corpus, BM25 postings, CUDA-generated vectors, typed graph nodes and edges, bounded adjacency indexes, and a source-CID-to-original-row lookup.
All entries are inert, non-authoritative evidence. Candidate and formal-logic rows cannot grant execution authority; exact… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-ir-graphrag.chunk_97
Dataset Card for "chunk_97"
More Information needed
Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.secopxcyber-security
Cybersecurity Instruction-Tuning Dataset
A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning,
built from 198 distinct sources spanning offensive security, blue-team
operations, vulnerability intelligence, cloud/AWS security, malware analysis,
digital forensics, and more. Every record is normalized to the standard
messages chat format and deduplicated at both file and record level.
⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.sec-8k-events
SEC Form 8-K Corporate Events
Every Form 8-K filed since the modern item taxonomy took effect — and, for each
one, the second the SEC accepted it, which is not the date printed on it.
1 761 353 filings · 3 676 835 item-level events · 23 August 2004 to today
The pipeline lives in recipe/ at the same revision as the data.
See PIPELINE.md for the method.
The problem this dataset exists to solve
Apple filed its June-quarter results on 30 July 2026. Here is the filing… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/sec-8k-events.security-kg
Security Knowledge Graph Triples
Security data from 24 sources represented as Subject-Predicate-Object (SPO) triples in Parquet format, ready for knowledge-graph construction, graph-ML, RAG pipelines, and threat-intelligence analysis.
Sources: ATT&CK · CAPEC · CWE · CVE · CPE · D3FEND · ATLAS · CAR · ENGAGE · F3 · EPSS · KEV · Vulnrichment · GHSA · Sigma · ExploitDB · MISP Galaxies · LOLBAS · LOLDrivers · Atomic Red Team · NIST 800-53 · Nuclei · EUVD · OSV
Last updated:… See the full description on the dataset page: https://huggingface.co/datasets/s0u9ata/security-kg.SecureCodeV2_qwen-2.5-7b-instruct_tokenized_vulnerablelekiwi_second_floor_0915_environmentThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi",
"total_episodes": 50,
"total_frames": 14471,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yuzhous/lekiwi_second_floor_0915_environment.zenodo-second-hand-fashion-v3
Second-Hand Fashion Dataset — wide (one row per garment)
Repack of Zenodo record 10.5281/zenodo.13788681 (Nauman et al., RISE + Wargön Innovation + Myrorna, CC-BY-4.0) into a one-row-per-garment wide layout so the HF dataset viewer shows every attribute — three images plus 25 metadata columns — on a single row.
Previous v3 releases stored one row per (garment, view) with satellite tables that had to be joined manually. That layout is preserved in git history if you need it; the… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/zenodo-second-hand-fashion-v3.lekiwi_second_floor_0916_environmentThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi",
"total_episodes": 50,
"total_frames": 15084,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yuzhous/lekiwi_second_floor_0916_environment.
