datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/TeraflopAI/SEC-EDGAR.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/Jeremydh911/SEC-EDGAR.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/baridhi/SEC-EDGAR.EDGAR_FILINGS_DATASET
SFD: SEC Filings Dataset (v1)
SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation.
This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in:
The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/adityaag2k/SEC-EDGAR.CoER-RL
CoER RL
Project page · Paper · Code
Executable task configurations for Stage 2 bilateral Co-PPO, with separate training and internal-validation splits. These are not generated rollouts or the official evaluation panel.
Contents
Split
File
Size
train
train.parquet
12,705 configurations
validation
validation.parquet
3,186 configurations
Training and validation configurations are disjoint. Internal validation is not a strict domain- or injection-OOD… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-RL.CoER-Attacker-SFT
CoER Attacker SFT
Project page · Paper · Code
Stage 1 supervision for initializing the adaptive attacker. Successful conversations retain all attacker turns, including earlier attempts that provide context for later adaptation.
Contents
Split
File
Size
train
train.jsonl
3,995 conversations
The corpus contains 11,655 assistant/attacker turns. Preserve all assistant-turn supervision; do not reduce a conversation to its final payload.
Load… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Attacker-SFT.CoER-Defender-SFT
CoER Defender SFT
Project page · Paper · Code
Stage 3 population-guided refinement data. Teacher agents execute tasks from initial states under retained attackers; only demonstrations verified for both safety and task completion are used.
Contents
Split
File
Size
train
train.jsonl
5,760 trajectories
This includes 4,907 attacked trajectories + 853 untriggered replays. An untriggered replay is an attack-configured run whose injection site was not… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Defender-SFT.edgar-corpus-htm2020
edgar-corpus-htm2020
dataset_info:
features:
- name: filename
dtype: string
- name: cik
dtype: string
- name: text
dtype: string
splits:
- name: train
num_bytes: 1004707483.6321189
num_examples: 6505
- name: test
num_bytes: 26565670.58950414
num_examples: 172
- name: validation
num_bytes: 26256767.44311456
num_examples: 170
download_size: 851417906
dataset_size: 1057529921.6647376
sp500-edgar-10k-markdown
edgar s&p500
Source Datasets
The source dataset used for this report is jlohding/sp500-edgar-10k.
Dataset Information
Configuration: default
Feature
Data Type
cik
string
sic
string
company
string
date
timestamp[us]
ret
float64
mkt_cap
float64
report_intro
string
text
string
report_returns
string
word_count
int64
Splits:
Train:
Number of Examples: 6258
Size: 2260000389 bytes
Download Size: 974801155 bytesDataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/sp500-edgar-10k-markdown.
