datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/TeraflopAI/SEC-EDGAR.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC-EDGAR.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/Jeremydh911/SEC-EDGAR.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/baridhi/SEC-EDGAR.EDGAR_FILINGS_DATASET
SFD: SEC Filings Dataset (v1)
SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation.
This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in:
The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/adityaag2k/SEC-EDGAR.edgar-forecast-benchmark
EDGAR-Forecast Benchmark
EDGAR-Forecast is a closed-sandbox benchmark for filing-grounded numerical forecasting from historical SEC filings in EDGAR. The benchmark contains 50 company-level instances and 250 numeric forecast targets from hidden 2026 10-Q filings.
Questions that mention 2025 refer to values disclosed in 2026 Q1 filings; those filings were filed in Q1 2026, so they remain outside the evaluated models' knowledge cutoffs.
Each benchmark directory includes the question… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/edgar-forecast-benchmark.CoER-RL
CoER RL
Project page · Paper · Code
Executable task configurations for Stage 2 bilateral Co-PPO, with separate training and internal-validation splits. These are not generated rollouts or the official evaluation panel.
Contents
Split
File
Size
train
train.parquet
12,705 configurations
validation
validation.parquet
3,186 configurations
Training and validation configurations are disjoint. Internal validation is not a strict domain- or injection-OOD… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-RL.CoER-Attacker-SFT
CoER Attacker SFT
Project page · Paper · Code
Stage 1 supervision for initializing the adaptive attacker. Successful conversations retain all attacker turns, including earlier attempts that provide context for later adaptation.
Contents
Split
File
Size
train
train.jsonl
3,995 conversations
The corpus contains 11,655 assistant/attacker turns. Preserve all assistant-turn supervision; do not reduce a conversation to its final payload.
Load… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Attacker-SFT.CoER-Defender-SFT
CoER Defender SFT
Project page · Paper · Code
Stage 3 population-guided refinement data. Teacher agents execute tasks from initial states under retained attackers; only demonstrations verified for both safety and task completion are used.
Contents
Split
File
Size
train
train.jsonl
5,760 trajectories
This includes 4,907 attacked trajectories + 853 untriggered replays. An untriggered replay is an attack-configured run whose injection site was not… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Defender-SFT.Agent-IPI-Structured-Interaction-Datasets-v2
Adversarial Dataset for LLM Instruction Hijacking / Tool-Calling Attacks
This directory contains the processed training and test datasets for evaluating and training defenses against prompt injection / instruction hijacking attacks in LLM tool-calling scenarios.
The dataset includes both JSON and XML formatted inputs, with three difficulty buckets:
no_attack: clean (benign) examples
easy: value-level or structure-level single attacks
hard: structure-destroying attacks or combined… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets-v2.EDGAR-FinTrace
EDGAR-FinTrace
Agentic financial-reasoning traces over SEC EDGAR filings, for finetuning
tool-using finance models. Each example is a complete episode: a question, the
agent's tool calls against a real filing, the tool observations, deterministic
calculate steps, and a grounded final answer. Every trace is verified to land
on a Python-computed gold value, so the corpus is label-noise-free.
Built for the Adaption AI hackathon (Finance category). Companion to the trained
weights… See the full description on the dataset page: https://huggingface.co/datasets/rachpradhan/EDGAR-FinTrace.edgar-corpus-htm2020
edgar-corpus-htm2020
dataset_info:
features:
- name: filename
dtype: string
- name: cik
dtype: string
- name: text
dtype: string
splits:
- name: train
num_bytes: 1004707483.6321189
num_examples: 6505
- name: test
num_bytes: 26565670.58950414
num_examples: 172
- name: validation
num_bytes: 26256767.44311456
num_examples: 170
download_size: 851417906
dataset_size: 1057529921.6647376
sp500-edgar-10k-markdown
edgar s&p500
Source Datasets
The source dataset used for this report is jlohding/sp500-edgar-10k.
Dataset Information
Configuration: default
Feature
Data Type
cik
string
sic
string
company
string
date
timestamp[us]
ret
float64
mkt_cap
float64
report_intro
string
text
string
report_returns
string
word_count
int64
Splits:
Train:
Number of Examples: 6258
Size: 2260000389 bytes
Download Size: 974801155 bytesDataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/sp500-edgar-10k-markdown.
