CoolFace
Datasetpublic

Jeremydh911/SEC-EDGAR

Datamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/Jeremydh911/SEC-EDGAR.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes2.8kdownloads
1 commits on main
ae291e55mo ago

Duplicate from TeraflopAI/SEC-EDGAR

Jeremydh911, conceptofmind