datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
10-K_sec_filings
Dataset Card for "10-K_sec_filings"
Dataset of 93.5K 10K SEC EDGAR filings since 1999 year. This dataset contains a lot of bad parsed filings and also empty rows
More Information needed
example-sec-filingsai-sec-10k-filingsEDGAR_FILINGS_DATASET_2022_2026H1EDGAR_FILINGS_DATASET
SFD: SEC Filings Dataset (v1)
SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation.
This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in:
The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.ai-sec-10k-filings-since-2020SEC_filings_1994_2024
Dataset Card for SEC EDGAR Filings Master Index
Dataset Details
Dataset Description
This dataset contains metadata for all submissions to the Securities and Exchange Commission (SEC) through their EDGAR system from 1994 until December 14, 2024. The data is extracted from quarterly master files and includes key information about company filings such as CIK numbers, company names, form types, and filing dates.
Curated by: Arthur (arthur@cicero.chat)
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC_filings_1994_2024.EDGAR_FILINGS_DATASET_2016_2021sec-filings-forward-return-2026
SEC Filings → Forward-Return (US large-cap, 2000–2026)
Leak-free, time-ordered SEC filings (8-K / 10-Q / 10-K) with objective forward-return
labels for ~610 current US large-cap names, through 2026. A benchmark ("can filing text
predict forward returns?"), not just a corpus.
⚖️ Licensing & provenance (read first — this is the honest part)
This repo deliberately separates two provenance classes:
part
source
license
Filing text + accession,cik,ticker,form… See the full description on the dataset page: https://huggingface.co/datasets/BatuhanECB/sec-filings-forward-return-2026.sec-filings-qa-instruct
SEC Filings Instruction-Tuning Dataset (Llama-3 Format)
This dataset contains 5,000 curated, instruction-formatted question-answering pairs derived from corporate SEC filings (Forms 10-K and 10-Q). It is structured specifically for parameter-efficient instruction fine-tuning (SFT/QLoRA) of Small Language Models using the standard Llama-3 ChatML template.
Dataset Details
Origin Source: Curated subset extracted from nvidia/Nemotron-SpecializedDomains-Finance-v1.… See the full description on the dataset page: https://huggingface.co/datasets/lateesha-bhatia/sec-filings-qa-instruct.fcc-ngso-filings
FCC NGSO Satellite Filings
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Structured metadata for the major FCC NGSO (non-geostationary satellite orbit) constellation filings — the authoritative public record of who has asked the US Federal Communications Commission for permission to launch and operate which mega-constellations. Includes Starlink Gen1 and Gen2, Amazon Project Kuiper, OneWeb, Telesat Lightspeed, and other… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/fcc-ngso-filings.kl3m-index-edgar-filingskl3m-index-edgar-filings-8-kfinsight-sec-filings
FinSight — SEC EDGAR Filings
Cleaned plain-text 10-K (annual) and 10-Q (quarterly) filings from the
US SEC EDGAR system for 20 large publicly-traded companies across 6 sectors.
Created as part of the FinSight project —
a financial research AI assistant combining BERT fine-tuning, RAG, and
multi-agent systems.
Stats
Records: 97
Companies: 20 (AAPL, MSFT, GOOGL, AMZN, META, NVDA, TSLA, JPM, BAC, GS,
JNJ, PFE, UNH, WMT, PG, KO, MCD, XOM, CVX, CAT)
Forms: 10-K, 10-Q… See the full description on the dataset page: https://huggingface.co/datasets/musk1209/finsight-sec-filings.sp500-sec-edgar-filings
S&P 500 Corporate Filings (SEC EDGAR 10-K & 8-K Archive)
Overview
This dataset contains clean, structured public data exported directly from production runs of Apify actors.
It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines.
Source Actor: captainhandsome/sec-edgar-filings-search
Dataset Page: Public sample and schema
Preconfigured Run Task: captainhandsome/apple-microsoft-10k-filings
Records… See the full description on the dataset page: https://huggingface.co/datasets/joeygambino/sp500-sec-edgar-filings.kl3m-index-edgar-filings-10-ksec-filingssec-filings-snippetskl3m-index-edgar-filings-sfilings-10kkl3m-index-edgar-filings-10-qsec_filingsfilings-rag-indexsec-filings-sample
SEC EDGAR Filings — Sample Dataset
A sample of 1,000 recent 8-K filings from SEC EDGAR, containing filing metadata and document references.
Dataset Description
The full SEC dataset includes millions of filings across all form types (10-K, 10-Q, 8-K, S-1, DEF 14A, etc.) from all publicly traded US companies. This sample contains 1,000 recent 8-K (material event) filings for evaluation.
Source
Publisher: U.S. Securities and Exchange Commission
URL:… See the full description on the dataset page: https://huggingface.co/datasets/carrierone/sec-filings-sample.stk-sec-filingsadaption-brazilian-regulatory-filings
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-brazilian_regulatory_filings
This dataset contains samples of Brazilian regulatory filings (Fatos Relevantes and Market Notices) from publicly traded companies, presented in both Portuguese and English. Each sample includes a classification task where the text is analyzed to determine the event type, status, scope relative to a target entity, and market signal sentiment. The content… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazilian-regulatory-filings.bankless_ROLLUP_BTC_1T_Market_Cap__More_ETH_ETF_Filings__STRK_Airdrop_Pushback8k_extracted_filingsai-sec-10k-filings-since-2020gr_athex_company_filings_processedThis data originates from https://www.athexgroup.gr/el/market-data/financial-data
It is mainly yearly and semesterly company filings, totalling 5937 company filings.
100 of these company filings were put aside for the "test" split to be used in the Greek OCR task, mainly recent ones from 2023 and 2024.
"tokens" column contains an integer which is the token count for that specific row, using the GPT-4 tokenizer.
The total amount of tokens using the Llama-3.1-8B tokenizer are: 0.45 B, using the… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/gr_athex_company_filings_processed.
