zorynthiq/zoryntiq-sec-filings
Zoryntiq SEC Filings Dataset A clean, LLM-ready dataset of SEC EDGAR filings from recently IPO'd and pre-IPO companies. Dataset Summary 5,179 text chunks extracted from 261 high-signal SEC filings (S-1, 10-K, 10-Q, 8-K, DRS, and more). Each chunk is ~1,500 words with 150-word overlap, cleaned and normalized for LLM training and financial NLP tasks. What's inside Registration statements (S-1, S-1/A, DRS) — full IPO prospectuses including business… See the full description on the dataset page: https://huggingface.co/datasets/zorynthiq/zoryntiq-sec-filings.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face