datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dojo_sector_precomputed
Languages: 简体中文 · English
dojo_sector_precomputed — Precomputed Sector Analytics
Overview
Derived sector analytics: L3 constituent snapshots, daily cap-weighted sector index levels, and per-constituent daily returns. Built offline from taxonomy, mappings, quotes, and stock K-lines.
Files
File
Description
manifest.json
Generation metadata: version, window start, row counts, latest trade dates
constituents.parquet
L3 constituent… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_precomputed.arxiv-complete
arXiv Complete Corpus
A snapshot of arXiv's metadata, version history, submission files and rendered
documents. It covers 3,148,796 papers and includes file contents, paths, sizes
and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface;
files come from the GCS mirror, S3 source archives and direct PDF fetches.
This release holds a PDF for 99.47% of papers and 99.54% of versions reported
with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/TeraflopAI/SEC-EDGAR.security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models.
These traces focus on security audits of opensource software.
Sharing traces with Swival
Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session:
swival "Fix the login bug" --trace-dir traces/
Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.dojo_sector_info
Languages: 简体中文 · English
dojo_sector_info — Sector Taxonomy
Overview
Three-level sector taxonomy (L1 industry → L2 chain → L3 sub-sector) with bilingual Chinese/English names and definitions.
Files
File
Description
data.parquet
Taxonomy tree (one L1 row each; L2/L3 nested in children)
Key Fields
Field
Description
id
L1 sector ID
name / name_alias
L1 English name / Chinese alias
description /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_info.dojo_sector_symbol_relations
Languages: 简体中文 · English
dojo_sector_symbol_relations — Stock–Sector Mapping
Overview
Maps each stock to L1/L2/L3 sector paths with primary and secondary assignments. One row per (ticker, market) pair.
Files
File
Description
data.parquet
Full stock ↔ sector relations
Key Fields
Field
Description
ticker
Stock symbol
market
us, cn, or hk
primary
JSON object — primary sector path
secondary
JSON array —… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_symbol_relations.secqa
SecQA
SecQA is a specialized dataset created for the evaluation of Large Language Models (LLMs) in the domain of computer security.
It consists of multiple-choice questions, generated using GPT-4 and the
Computer Systems Security: Planning for Success textbook,
aimed at assessing the understanding and application of LLMs' knowledge in computer security.
Dataset Details
Dataset Description
SecQA is an innovative dataset designed to benchmark the… See the full description on the dataset page: https://huggingface.co/datasets/zefang-liu/secqa.SEC
SEC Annual Reports (Form 10-K) 1993-2024
Dataset Overview
This dataset comprises SEC annual reports (Form 10-K) for the years 1993 to 2024, providing comprehensive coverage of publicly traded companies' financial and business information. The reports are stored in Parquet format, ensuring efficient storage and quick access. This dataset was meticulously compiled using the EDGAR-Crawler toolkit, which facilitates the extraction and processing of SEC filings from the EDGAR… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/SEC.financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system.
Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences.
Sentiment labels are provided on a per filing basis from the market reaction around the filing data.
Additional metadata for each filing is included in the dataset.sec-filingsfinancial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system.
Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences.
Sentiment labels are provided on a per filing basis from the market reaction around the filing data.
Additional metadata for each filing is included in the dataset.404mini
404-GEN Mini 3D
This dataset contains over 20,000 3D assets generated with text prompts using 3D Gaussian Splatting, designed for text-to-3D generation tasks. This is a sample of a much larger dataset comprised of 21.5M assets and 40TB in size, available by request at https://dataset.404.xyz
Dataset Description
Dataset Summary
404-GEN Mini 3D is a collection of over 20,000 3D assets generated from text prompts on Bittensor Subnet 17, providing mid-… See the full description on the dataset page: https://huggingface.co/datasets/Olague-Secret/404mini.security_instruct_mcq_2481cyber-security
Cybersecurity AI Knowledge Base — PhD-Level Dataset
Overview
This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security.
Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms
Purpose
Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.secret-scan-remediation-trajectories
Secret Scan Remediation Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/secret-scan-remediation-trajectories.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC-EDGAR.sec-10k-markdown-uncompressed
📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents)
Dataset Summary
This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025).
The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.DTap-Bench-Agent-Trajectories
DecodingTrust-Agent Platform
A Controllable and Interactive Red-Teaming Platform for AI Agents.
This is the full collection of the agent trajectories produced from evaluating the DTap-Bench from DecodingTrust-Agent Platform (DTAP),
spanning 14 real-world domains and 50+ simulation environments that replicate widely-used
systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task
ships the configuration the evaluator needs to spin up the… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DTap-Bench-Agent-Trajectories.sec-contracts-financial-extraction-instructions
S&P 500 SEC Financial Extraction Instructions
Dataset Summary
7,683 instruction-tuning examples for training LLMs to extract structured financial data from SEC filings. Covers two filing types across S&P 500 companies:
Split
Examples
Filing Type
Description
train
3,430
Exhibit 10 + DEF 14A
Positive examples with validated outputs
corrective
4,253
Exhibit 10 + DEF 14A
Corrective, rescued, and negative examples
Exhibit 10 — Material Contracts (2… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-financial-extraction-instructions.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/Jeremydh911/SEC-EDGAR.the-secrets-of-ceos-book-2k
The-Secrets-Of-Ceos-Book-2k
Made with ❤️ using 🦥 Unsloth Studio
Beta2x was generated with Unsloth Recipe Studio. It contains 2,000 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("filipwx/the-secrets-of-ceos-book-2k", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 2,000
📋 Columns: 4
📋 Schema & Statistics
Column
Type
Column Type
Unique… See the full description on the dataset page: https://huggingface.co/datasets/filipwx/the-secrets-of-ceos-book-2k.SEC-bench
Data Instances
instance_id: (str) - A unique identifier for the instance
repo: (str) - The repository name including the owner
project_name: (str) - The name of the project without owner
lang: (str) - The programming language of the repository
work_dir: (str) - Working directory path
sanitizer: (str) - The type of sanitizer used for testing (e.g., Address, Memory, Undefined)
bug_description: (str) - Description of the vulnerability
base_commit: (str) - The base commit hash where the… See the full description on the dataset page: https://huggingface.co/datasets/SEC-bench/SEC-bench.comptia_security_pluse_701DecodingTrust-Agent-Platform
DecodingTrust-Agent Platform
A Controllable and Interactive Red-Teaming Platform for AI Agents.
This is the per-task dataset for the DecodingTrust-Agent Platform (DTAP),
spanning 14 real-world domains and 50+ simulation environments that replicate widely-used
systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task
ships the configuration the evaluator needs to spin up the sandbox, run an agent, and verify the
outcome — config.yaml (task… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust-Agent-Platform.eliciting-secret-knowledge-resultssecurecode-web
SecureCode Web: Traditional Web & Application Security Dataset
Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance
Paper | GitHub | Dataset | Model Collection | Blog Post
What's new in v2.6
v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had
shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/scthornton/securecode-web.generated-passport-faces-aditya-second-halfCode_Vulnerability_Security_DPO
Cybernative.ai Code Vulnerability and Security Dataset
Dataset Description
The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen on… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative/Code_Vulnerability_Security_DPO.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/baridhi/SEC-EDGAR.airport-security
