zorynthiq/zoryntiq-sec-filings
Zoryntiq SEC Filings Dataset A clean, LLM-ready dataset of SEC EDGAR filings from recently IPO'd and pre-IPO companies. Dataset Summary 5,179 text chunks extracted from 261 high-signal SEC filings (S-1, 10-K, 10-Q, 8-K, DRS, and more). Each chunk is ~1,500 words with 150-word overlap, cleaned and normalized for LLM training and financial NLP tasks. What's inside Registration statements (S-1, S-1/A, DRS) — full IPO prospectuses including business… See the full description on the dataset page: https://huggingface.co/datasets/zorynthiq/zoryntiq-sec-filings.
146
Upload 2 files
initial commit
