CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TheTokenFactory /sec-contracts-financial-extraction-instructions S&P 500 SEC Financial Extraction Instructions Dataset Summary 7,683 instruction-tuning examples for training LLMs to extract structured financial data from SEC filings. Covers two filing types across S&P 500 companies: Split Examples Filing Type Description train 3,430 Exhibit 10 + DEF 14A Positive examples with validated outputs corrective 4,253 Exhibit 10 + DEF 14A Corrective, rescued, and negative examples Exhibit 10 — Material Contracts (2… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-financial-extraction-instructions.texttext-generation10K<n<100K1 likes3.2k downloads6mo agoHugging Face02chenghao /sec-material-contracts Material Contracts (Exhibit 10) from SEC/EDGAR Because sometimes you need 1,141,632 examples of corporate legalese to train your next model ☕ Dataset Summary Picture this: 1,141,632 material contracts (Exhibit 10) painstakingly collected from sec.gov's EDGAR database. We're talking about legal agreements spanning from 1994 to 2025 Q1, sourced from 10-K, 10-Q, and 8-K filings. Think of Exhibit 10 as the treasure trove where companies hide their most important legal… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts.texttext-generation1M<n<10M3 likes1.4k downloads1y agoHugging Face03mteb /legalbench_consumer_contracts_qa LegalBenchConsumerContractsQA An MTEB dataset Massive Text Embedding Benchmark The dataset includes questions and answers related to contracts. Task category t2t Domains Legal, Written Reference https://huggingface.co/datasets/nguha/legalbench/viewer/consumer_contracts_qa How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["LegalBenchConsumerContractsQA"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/legalbench_consumer_contracts_qa.texttext-retrievaln<1K1 likes1.2k downloads1y agoHugging Face04chenghao /sec-material-contracts-qa800+ EDGAR contracts with PDF images and key information extracted by the OpenAI GPT-4o model. The key information is defined as follows: class KeyInformation(BaseModel): agreement_date : str = Field(description="Agreement signing date of the contract. (date)") effective_date : str = Field(description="Effective date of the contract. (date)") expiration_date : str = Field(description="Service end date or expiration date of the contract. (date)") party_address : str =… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts-qa.tabularvisual-question-answeringn<1K3 likes930 downloads2y agoHugging Face05mwritescode /slither-audited-smart-contractsThis dataset contains source code and deployed bytecode for Solidity Smart Contracts that have been verified on Etherscan.io, along with a classification of their vulnerabilities according to the Slither static analysis framework.texttext-classification100K<n<1M61 likes738 downloads4y agoHugging Face06ginntonicfun /cuad-pdf-contractsThis dataset contains the contract pdf files from https://huggingface.co/datasets/theatticusproject/cuad Files present in https://huggingface.co/datasets/theatticusproject/cuad/tree/main/CUAD_v1/full_contract_pdf are extracted and then uploaded in single bucket without folder heiraarchies. About CUAD: CONTRACT UNDERSTANDING ATTICUS DATASET Contract Understanding Atticus Dataset (CUAD) v1 is a corpus of more than 13,000 labels in 510 commercial legal contracts that have been manually labeled… See the full description on the dataset page: https://huggingface.co/datasets/ginntonicfun/cuad-pdf-contracts.documentn<1K0 likes621 downloads4mo agoHugging Face07nhankins /legal_contractstext10K<n<100K5 likes499 downloads3y agoHugging Face08Lucius-Morningstar /mailroom-cuad-contracts Mailroom Eval: Cuad Contracts Mirror of the Braintrust evaluation dataset mailroom-cuad-contracts from the llm-entity-extraction experiment loop (llm-mailroom legal document pipeline). Field Value Rows 50 Source script stream_cuad_to_bt.py Braintrust dataset id c55ac7f0-56ff-4a2c-b968-7f382ce7daea Braintrust project id 02fb28b9-60e2-40b6-a68a-b72ee0b237ad Exported (UTC) 2026-08-22T05:03:53+00:00 Export sha256… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/mailroom-cuad-contracts.imagetext-classificationn<1K0 likes455 downloads1mo agoHugging Face09geodesic-research /environment-contracts geodesic-research/environment-contracts Local-pipeline snapshot published via --push-from-local (GH #52). All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision. Pipeline run params hash: dd13a8843fce63fefa4e70c743f85c6cba30904523b6097a96bfe98100c298eb Configs in this snapshot: conversation Per-run provenance: _pipeline_state/dd13a8843fce63fefa4e70c743f85c6cba30904523b6097a96bfe98100c298eb.json (in this repo) and each… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/environment-contracts.textn<1K1 likes445 downloads4mo agoHugging Face10RBudzynski /Contracts-Ethereum-Cryptocurrency-Data Contracts-Ethereum-Cryptocurrency-Data Hive-partitioned Parquet export of BlockDB contracts (Ethereum). Load from datasets import load_dataset ds = load_dataset("RBudzynski/Contracts-Ethereum-Cryptocurrency-Data", split="train") Files live under data/year=YYYY/month=MM/part-NNNN.parquet. Range: 2015-08 .. 2026-06 (UTC calendar months). Schema column type block_timestamp timestamp block_number int64 tx_index int32 trace_address… See the full description on the dataset page: https://huggingface.co/datasets/RBudzynski/Contracts-Ethereum-Cryptocurrency-Data.tabular100M<n<1B0 likes441 downloads26d agoHugging Face11Zellic /all-ethereum-contracts All ethereum contracts This dataset contains all deployed Ethereum contracts as of block 21850000 (February 15th, 2025), bytecodes of the contracts, and the block numbers the contracts were deployed. Contract bytecodes are stored as a hash of the bytecode, and another dataset is provided mapping bytecode hashes to bytecodes. This is to reduce the size of the dataset, as many contracts have identical bytecodes. This dataset was exported from a PostgreSQL database into CSV format.… See the full description on the dataset page: https://huggingface.co/datasets/Zellic/all-ethereum-contracts.imagen<1K7 likes372 downloads1y agoHugging Face12ymoslem /us_contractstext100K<n<1M0 likes342 downloads2y agoHugging Face13albertvillanova /legal_contractsThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.text100K<n<1M54 likes316 downloads5y agoHugging Face14chenghao /sec-material-contracts-qa-splittedMixed and filtered version of chenghao/sec-material-contracts-qa and jordyvl/DUDE_subset_100val. imagevisual-question-answering1K<n<10K8 likes300 downloads2y agoHugging Face15apehex /evm_contracts EVM Contracts Description EVMC (Ethereum Virtual Machine Contracts) is a collection of smart contracts from the ETH blockchain. In particular, each sample holds the creation and runtime bytecodes. When available, the sources are also included. Metadata homepage: https://github.com/apehex/feedblocks version: 1.0.1 HEX Dataset Config Split Size Samples Blocks 'hex-ethereum' 'train' 2.8 GB 1,294,247 19,493,000 - 20,292,000 'hex-ethereum'… See the full description on the dataset page: https://huggingface.co/datasets/apehex/evm_contracts.tabular1M<n<10M0 likes291 downloads2y agoHugging Face16ZipLime /us-government-contracts US Federal Contract Awards Federal contract actions paid to publicly traded companies — and, for each one, the day the public could first see it, which for the Department of Defense is three months after the contract was signed. 12 371 350 contract transactions · 9 552 299 awards · 1 694 listed companies · FY2005 to today · $4.41tn obligated 70% of those rows are the Department of Defense and the Army Corps of Engineers, and not one of them is dated less than ninety-two days… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/us-government-contracts.tabulartabular-regression10M<n<100M0 likes239 downloads7d agoHugging Face17narcolepticchicken /sec-contracts-2015-2025 SEC Contracts 2015–2025 Mini corpus of contract text extracted from SEC filings (EDGAR). Includes clause‑level rows with metadata for classification and QA experiments. texttext-classification10K<n<100K0 likes180 downloads1y agoHugging Face18zzsi /usaspending_2024_all_contractstabular1M<n<10M4 likes174 downloads2y agoHugging Face19momilla /flattened_contractstabular100K<n<1M1 likes150 downloads4y agoHugging Face20hugsid /legal-contractstext10K<n<100K3 likes126 downloads2y agoHugging Face21TheTokenFactory /sec-contracts-corrective-extraction S&P 500 SEC Financial Extractions - Corrective Dataset Dataset Summary 4,253 corrective instruction-tuning examples designed to teach LLMs what the base model gets wrong when extracting structured financial data from SEC filings. Covers both Exhibit 10 material contracts and DEF 14A proxy statements from S&P 500 companies. This is a companion to TheTokenFactory/sec-contracts-financial-extraction-instructions, which contains the positive training examples. Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-corrective-extraction.texttext-generation10K<n<100K0 likes126 downloads6mo agoHugging Face22arthrod /contractscrub-repro ContractScrub reproduction (synthetic variant) Reproduction of ContractScrub: A benchmark for final review of legal contracts (arXiv:2608.20204, Thomson Reuters Foundational Research). The official gold (tri-fair-lab/contract_scrub) was still unreleased at reproduction time, so this repo contains a method-faithful pipeline + a synthetic benchmark with ground truth true by construction — not the paper's leaderboard numbers. Contents SCORING.md — detailed… See the full description on the dataset page: https://huggingface.co/datasets/arthrod/contractscrub-repro.0 likes119 downloads7d agoHugging Face23MongoDB /supply_chain_contracts_dataset_smalltextn<1K2 likes112 downloads2y agoHugging Face24andreaaltomani /reinsurance-contracts-classificationThis dataset contains the full text and the classification of 6671 reinsurance-related documents extracted from SEC filings. The documents include reinsurance contracts, but also ancillary documents, such as amendments, endorsements, extensions, and other documents referring to reinsurance agreement. The documents have been classified by three language models: qwen3-235b-a22b-2507, gpt-oss-120b and gemini-2.5-flash-lite. Manual inspection of the classification results reveals that qwen3 is… See the full description on the dataset page: https://huggingface.co/datasets/andreaaltomani/reinsurance-contracts-classification.texttext-classification1K<n<10K2 likes98 downloads1y agoHugging Face25Mr-Bridge /nba-salary-cap-contracts-2016-2026 NBA Salary Cap & Contracts, 2016-2026 An analysis-ready dataset of NBA salaries, payrolls and the salary-cap system over ten seasons (2015-16 baseline through 2025-26, plus a mechanical 2026-27 to 2031-32 projection). It combines player-salary snapshots with the institutional thresholds that govern them: cap, luxury tax, aprons, minimums and the rookie scale. Published by MrBridge. It backs the study NBA Salaries and Contracts, 2016-2026 on mr-bridge.com. Files… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/nba-salary-cap-contracts-2016-2026.1K<n<10K0 likes89 downloads2mo agoHugging Face26santoshtyss /us_contractstext100K<n<1M5 likes78 downloads3y agoHugging Face27kozue13 /Contractstextsummarization10K<n<100K0 likes76 downloads2y agoHugging Face28AlfredPros /smart-contracts-instructions Smart Contracts Instructions A dataset containing 6,003 GPT-generated human instruction and Solidity source code data pairs. GPT models used to make this data are GPT-3.5 turbo, GPT-3.5 turbo 16k context, and GPT-4. Solidity source codes are used from mwritescode's Slither Audited Smart Contracts (https://huggingface.co/datasets/mwritescode/slither-audited-smart-contracts). Distributions of the GPT models used to make this dataset: GPT-3.5 Turbo: 5,276 GPT-3.5 Turbo 16k Context:… See the full description on the dataset page: https://huggingface.co/datasets/AlfredPros/smart-contracts-instructions.textquestion-answering1K<n<10K6 likes69 downloads2y agoHugging Face29tri-fair-lab /contract_scrub Coming soon! 1 likes61 downloads1mo agoHugging Face30lawinsider /uk_ner_contracts_spacy Dataset Description Legal Contracts Dataset for Training SpaCy NER Model This repository contains a specially curated dataset consisting of legal contracts. It is designed for the purpose of training a Named Entity Recognition (NER) model using SpaCy, with the aim to recognize and classify four types of entities in the text: Contract Type, Clause Title, Clause Number, Definition Title The dataset includes a broad variety of legal contracts, covering diverse domains such as… See the full description on the dataset page: https://huggingface.co/datasets/lawinsider/uk_ner_contracts_spacy.tabulartoken-classification10K<n<100K2 likes50 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.