datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sec-filings10-K_sec_filings
Dataset Card for "10-K_sec_filings"
Dataset of 93.5K 10K SEC EDGAR filings since 1999 year. This dataset contains a lot of bad parsed filings and also empty rows
More Information needed
example-sec-filingsai-sec-10k-filingsai-sec-10k-filings-since-2020SEC_filings_1994_2024
Dataset Card for SEC EDGAR Filings Master Index
Dataset Details
Dataset Description
This dataset contains metadata for all submissions to the Securities and Exchange Commission (SEC) through their EDGAR system from 1994 until December 14, 2024. The data is extracted from quarterly master files and includes key information about company filings such as CIK numbers, company names, form types, and filing dates.
Curated by: Arthur (arthur@cicero.chat)
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC_filings_1994_2024.sec-filings-forward-return-2026
SEC Filings → Forward-Return (US large-cap, 2000–2026)
Leak-free, time-ordered SEC filings (8-K / 10-Q / 10-K) with objective forward-return
labels for ~610 current US large-cap names, through 2026. A benchmark ("can filing text
predict forward returns?"), not just a corpus.
⚖️ Licensing & provenance (read first — this is the honest part)
This repo deliberately separates two provenance classes:
part
source
license
Filing text + accession,cik,ticker,form… See the full description on the dataset page: https://huggingface.co/datasets/BatuhanECB/sec-filings-forward-return-2026.finsight-sec-filings
FinSight — SEC EDGAR Filings
Cleaned plain-text 10-K (annual) and 10-Q (quarterly) filings from the
US SEC EDGAR system for 20 large publicly-traded companies across 6 sectors.
Created as part of the FinSight project —
a financial research AI assistant combining BERT fine-tuning, RAG, and
multi-agent systems.
Stats
Records: 97
Companies: 20 (AAPL, MSFT, GOOGL, AMZN, META, NVDA, TSLA, JPM, BAC, GS,
JNJ, PFE, UNH, WMT, PG, KO, MCD, XOM, CVX, CAT)
Forms: 10-K, 10-Q… See the full description on the dataset page: https://huggingface.co/datasets/musk1209/finsight-sec-filings.sec-filingssec-filings-qa-instruct
SEC Filings Instruction-Tuning Dataset (Llama-3 Format)
This dataset contains 5,000 curated, instruction-formatted question-answering pairs derived from corporate SEC filings (Forms 10-K and 10-Q). It is structured specifically for parameter-efficient instruction fine-tuning (SFT/QLoRA) of Small Language Models using the standard Llama-3 ChatML template.
Dataset Details
Origin Source: Curated subset extracted from nvidia/Nemotron-SpecializedDomains-Finance-v1.… See the full description on the dataset page: https://huggingface.co/datasets/lateesha-bhatia/sec-filings-qa-instruct.sp500-sec-edgar-filings
S&P 500 Corporate Filings (SEC EDGAR 10-K & 8-K Archive)
Overview
This dataset contains clean, structured public data exported directly from production runs of Apify actors.
It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines.
Source Actor: captainhandsome/sec-edgar-filings-search
Dataset Page: Public sample and schema
Preconfigured Run Task: captainhandsome/apple-microsoft-10k-filings
Records… See the full description on the dataset page: https://huggingface.co/datasets/joeygambino/sp500-sec-edgar-filings.sec-filings-snippetssec-filings-sample
SEC EDGAR Filings — Sample Dataset
A sample of 1,000 recent 8-K filings from SEC EDGAR, containing filing metadata and document references.
Dataset Description
The full SEC dataset includes millions of filings across all form types (10-K, 10-Q, 8-K, S-1, DEF 14A, etc.) from all publicly traded US companies. This sample contains 1,000 recent 8-K (material event) filings for evaluation.
Source
Publisher: U.S. Securities and Exchange Commission
URL:… See the full description on the dataset page: https://huggingface.co/datasets/carrierone/sec-filings-sample.zoryntiq-sec-filings
Zoryntiq SEC Filings Dataset
A clean, LLM-ready dataset of SEC EDGAR filings from recently IPO'd and pre-IPO companies.
Dataset Summary
5,179 text chunks extracted from 261 high-signal SEC filings (S-1, 10-K, 10-Q, 8-K, DRS, and more). Each chunk is ~1,500 words with 150-word overlap, cleaned and normalized for LLM training and financial NLP tasks.
What's inside
Registration statements (S-1, S-1/A, DRS) — full IPO prospectuses including business descriptions… See the full description on the dataset page: https://huggingface.co/datasets/zorynthiq/zoryntiq-sec-filings.sec_filingsstk-sec-filingssec-filings
SEC EDGAR filings
Every SEC filing indexed by company (CIK), form type, and date. Includes 10-K/10-Q/8-K, Form 4 insider transactions, Form 144 notice-of-sale, Schedule 13D/G, 13F holdings, NPORT-P fund holdings, XBRL financial facts.
Live API
This dataset is served via a live REST API at api.ai-analytics.org. The card you're reading exists so HuggingFace's index can route AI agents + researchers to the canonical source.
API endpoint:… See the full description on the dataset page: https://huggingface.co/datasets/emperor-mew/sec-filings.verilex-sec-filings
SEC EDGAR Filings
Structured SEC filings (10-K, 10-Q, 8-K, all form types). Updated every 15 minutes.
API: https://api.verilexdata.com/api/v1/sec/filings ($0.018/query via x402)
Free: /api/v1/sec/sample, /api/v1/sec/stats
Website | MCP
{"accession_number":"0001104659-24-098765","company_name":"Lenovo Group Limited","form_type":"20-F","filed_date":"2024-08-14"}
Provided by Verilex Data (Optimal Reality LLC).
ai-sec-10k-filings-since-2020
