CoolFace
Datasetpublic

kapilrao/SEC-EDGAR

Datamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC-EDGAR.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
1likes4.4kdownloads

kapilrao/SEC-EDGAR · main · files are served by the source, never re-hosted here