CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TeraflopAI /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/TeraflopAI/SEC-EDGAR.texttext-generation1M<n<10M47 likes16k downloads5mo agoHugging Face02Jeremydh911 /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/Jeremydh911/SEC-EDGAR.texttext-generation1M<n<10M0 likes2.8k downloads5mo agoHugging Face03baridhi /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/baridhi/SEC-EDGAR.texttext-generation1M<n<10M0 likes1.9k downloads5mo agoHugging Face04anonymous-md /EDGAR_FILINGS_DATASET SFD: SEC Filings Dataset (v1) SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation. This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in: The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.tabulartext-generation1M<n<10M2 likes895 downloads5mo agoHugging Face05adityaag2k /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/adityaag2k/SEC-EDGAR.texttext-generation1M<n<10M0 likes859 downloads5mo agoHugging Face06Z-Edgar /CoER-RL CoER RL Project page · Paper · Code Executable task configurations for Stage 2 bilateral Co-PPO, with separate training and internal-validation splits. These are not generated rollouts or the official evaluation panel. Contents Split File Size train train.parquet 12,705 configurations validation validation.parquet 3,186 configurations Training and validation configurations are disjoint. Internal validation is not a strict domain- or injection-OOD… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-RL.texttext-generation10K<n<100K0 likes106 downloads4d agoHugging Face07Z-Edgar /CoER-Attacker-SFT CoER Attacker SFT Project page · Paper · Code Stage 1 supervision for initializing the adaptive attacker. Successful conversations retain all attacker turns, including earlier attempts that provide context for later adaptation. Contents Split File Size train train.jsonl 3,995 conversations The corpus contains 11,655 assistant/attacker turns. Preserve all assistant-turn supervision; do not reduce a conversation to its final payload. Load… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Attacker-SFT.texttext-generation1K<n<10K0 likes98 downloads4d agoHugging Face08Z-Edgar /CoER-Defender-SFT CoER Defender SFT Project page · Paper · Code Stage 3 population-guided refinement data. Teacher agents execute tasks from initial states under retained attackers; only demonstrations verified for both safety and task completion are used. Contents Split File Size train train.jsonl 5,760 trajectories This includes 4,907 attacked trajectories + 853 untriggered replays. An untriggered replay is an attack-configured run whose injection site was not… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Defender-SFT.texttext-generation1K<n<10K0 likes91 downloads4d agoHugging Face09pszemraj /edgar-corpus-htm2020 edgar-corpus-htm2020 dataset_info: features: - name: filename dtype: string - name: cik dtype: string - name: text dtype: string splits: - name: train num_bytes: 1004707483.6321189 num_examples: 6505 - name: test num_bytes: 26565670.58950414 num_examples: 172 - name: validation num_bytes: 26256767.44311456 num_examples: 170 download_size: 851417906 dataset_size: 1057529921.6647376 texttext-generation10K<n<100K0 likes56 downloads9mo agoHugging Face10BEE-spoke-data /sp500-edgar-10k-markdowngated edgar s&p500 Source Datasets The source dataset used for this report is jlohding/sp500-edgar-10k. Dataset Information Configuration: default Feature Data Type cik string sic string company string date timestamp[us] ret float64 mkt_cap float64 report_intro string text string report_returns string word_count int64 Splits: Train: Number of Examples: 6258 Size: 2260000389 bytes Download Size: 974801155 bytesDataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/sp500-edgar-10k-markdown.tabulartext-generation10K<n<100K6 likes33 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.