CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TeraflopAI /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/TeraflopAI/SEC-EDGAR.texttext-generation1M<n<10M47 likes16k downloads5mo agoHugging Face02kapilrao /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC-EDGAR.text-generation1M<n<10M1 likes4.1k downloads5mo agoHugging Face03Jeremydh911 /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/Jeremydh911/SEC-EDGAR.texttext-generation1M<n<10M0 likes2.8k downloads5mo agoHugging Face04baridhi /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/baridhi/SEC-EDGAR.texttext-generation1M<n<10M0 likes1.7k downloads5mo agoHugging Face05anonymous-md /EDGAR_FILINGS_DATASET SFD: SEC Filings Dataset (v1) SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation. This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in: The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.tabulartext-generation1M<n<10M2 likes861 downloads5mo agoHugging Face06adityaag2k /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/adityaag2k/SEC-EDGAR.texttext-generation1M<n<10M0 likes854 downloads5mo agoHugging Face07sfd-anonymous /edgar-forecast-benchmark EDGAR-Forecast Benchmark EDGAR-Forecast is a closed-sandbox benchmark for filing-grounded numerical forecasting from historical SEC filings in EDGAR. The benchmark contains 50 company-level instances and 250 numeric forecast targets from hidden 2026 10-Q filings. Questions that mention 2025 refer to values disclosed in 2026 Q1 filings; those filings were filed in Q1 2026, so they remain outside the evaluated models' knowledge cutoffs. Each benchmark directory includes the question… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/edgar-forecast-benchmark.text-generation0 likes384 downloads5mo agoHugging Face08Z-Edgar /CoER-RL CoER RL Project page · Paper · Code Executable task configurations for Stage 2 bilateral Co-PPO, with separate training and internal-validation splits. These are not generated rollouts or the official evaluation panel. Contents Split File Size train train.parquet 12,705 configurations validation validation.parquet 3,186 configurations Training and validation configurations are disjoint. Internal validation is not a strict domain- or injection-OOD… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-RL.texttext-generation10K<n<100K0 likes100 downloads4d agoHugging Face09Z-Edgar /CoER-Attacker-SFT CoER Attacker SFT Project page · Paper · Code Stage 1 supervision for initializing the adaptive attacker. Successful conversations retain all attacker turns, including earlier attempts that provide context for later adaptation. Contents Split File Size train train.jsonl 3,995 conversations The corpus contains 11,655 assistant/attacker turns. Preserve all assistant-turn supervision; do not reduce a conversation to its final payload. Load… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Attacker-SFT.texttext-generation1K<n<10K0 likes95 downloads4d agoHugging Face10Z-Edgar /CoER-Defender-SFT CoER Defender SFT Project page · Paper · Code Stage 3 population-guided refinement data. Teacher agents execute tasks from initial states under retained attackers; only demonstrations verified for both safety and task completion are used. Contents Split File Size train train.jsonl 5,760 trajectories This includes 4,907 attacked trajectories + 853 untriggered replays. An untriggered replay is an attack-configured run whose injection site was not… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Defender-SFT.texttext-generation1K<n<10K0 likes87 downloads4d agoHugging Face11Z-Edgar /Agent-IPI-Structured-Interaction-Datasets-v2 Adversarial Dataset for LLM Instruction Hijacking / Tool-Calling Attacks This directory contains the processed training and test datasets for evaluating and training defenses against prompt injection / instruction hijacking attacks in LLM tool-calling scenarios. The dataset includes both JSON and XML formatted inputs, with three difficulty buckets: no_attack: clean (benign) examples easy: value-level or structure-level single attacks hard: structure-destroying attacks or combined… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets-v2.text-generation100K<n<1M1 likes65 downloads8mo agoHugging Face12rachpradhan /EDGAR-FinTrace EDGAR-FinTrace Agentic financial-reasoning traces over SEC EDGAR filings, for finetuning tool-using finance models. Each example is a complete episode: a question, the agent's tool calls against a real filing, the tool observations, deterministic calculate steps, and a grounded final answer. Every trace is verified to land on a Python-computed gold value, so the corpus is label-noise-free. Built for the Adaption AI hackathon (Finance category). Companion to the trained weights… See the full description on the dataset page: https://huggingface.co/datasets/rachpradhan/EDGAR-FinTrace.text-generation0 likes53 downloads4mo agoHugging Face13pszemraj /edgar-corpus-htm2020 edgar-corpus-htm2020 dataset_info: features: - name: filename dtype: string - name: cik dtype: string - name: text dtype: string splits: - name: train num_bytes: 1004707483.6321189 num_examples: 6505 - name: test num_bytes: 26565670.58950414 num_examples: 172 - name: validation num_bytes: 26256767.44311456 num_examples: 170 download_size: 851417906 dataset_size: 1057529921.6647376 texttext-generation10K<n<100K0 likes51 downloads9mo agoHugging Face14BEE-spoke-data /sp500-edgar-10k-markdowngated edgar s&p500 Source Datasets The source dataset used for this report is jlohding/sp500-edgar-10k. Dataset Information Configuration: default Feature Data Type cik string sic string company string date timestamp[us] ret float64 mkt_cap float64 report_intro string text string report_returns string word_count int64 Splits: Train: Number of Examples: 6258 Size: 2260000389 bytes Download Size: 974801155 bytesDataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/sp500-edgar-10k-markdown.tabulartext-generation10K<n<100K6 likes31 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.