CoolFace
Datasetpublic

zorynthiq/zoryntiq-sec-filings

Zoryntiq SEC Filings Dataset A clean, LLM-ready dataset of SEC EDGAR filings from recently IPO'd and pre-IPO companies. Dataset Summary 5,179 text chunks extracted from 261 high-signal SEC filings (S-1, 10-K, 10-Q, 8-K, DRS, and more). Each chunk is ~1,500 words with 150-word overlap, cleaned and normalized for LLM training and financial NLP tasks. What's inside Registration statements (S-1, S-1/A, DRS) — full IPO prospectuses including business… See the full description on the dataset page: https://huggingface.co/datasets/zorynthiq/zoryntiq-sec-filings.

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
1likes46downloads

zorynthiq/zoryntiq-sec-filings · main · files are served by the source, never re-hosted here