CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TheTokenFactory /sec-contracts-financial-extraction-instructions S&P 500 SEC Financial Extraction Instructions Dataset Summary 7,683 instruction-tuning examples for training LLMs to extract structured financial data from SEC filings. Covers two filing types across S&P 500 companies: Split Examples Filing Type Description train 3,430 Exhibit 10 + DEF 14A Positive examples with validated outputs corrective 4,253 Exhibit 10 + DEF 14A Corrective, rescued, and negative examples Exhibit 10 — Material Contracts (2… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-financial-extraction-instructions.texttext-generation10K<n<100K1 likes3.2k downloads6mo agoHugging Face02ajaxdavis /donto-qwen3.8-27b-predicate-extraction-data Donto-Qwen3.8 Predicate Extraction Data V15 This repository is the complete public data and evidence companion to ajaxdavis/donto-qwen3.8-27b-predicate-extractor. It contains the canonical V15 extraction training/validation corpus, the validator corpus, the optional D1-repeat ablation, the once-sealed 100-document graph-first gold suite, exact tool schemas, generator/evaluator source, hashes, and audit reports. Why this dataset exists Donto is designed for… See the full description on the dataset page: https://huggingface.co/datasets/ajaxdavis/donto-qwen3.8-27b-predicate-extraction-data.texttext-generation1K<n<10K0 likes155 downloads29d agoHugging Face03TheTokenFactory /sec-contracts-corrective-extraction S&P 500 SEC Financial Extractions - Corrective Dataset Dataset Summary 4,253 corrective instruction-tuning examples designed to teach LLMs what the base model gets wrong when extracting structured financial data from SEC filings. Covers both Exhibit 10 material contracts and DEF 14A proxy statements from S&P 500 companies. This is a companion to TheTokenFactory/sec-contracts-financial-extraction-instructions, which contains the positive training examples. Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-corrective-extraction.texttext-generation10K<n<100K0 likes126 downloads6mo agoHugging Face04Ionio-ai /ecommerce-search-extraction Ionio E-commerce Search Query Extraction Built with: simula — schema-driven synthetic data generation with auditable taxonomy lineage. An English synthetic dataset for training and evaluating systems that translate natural-language shopping requests into narrow, atomic, database-queryable JSON. It contains 10,985 accepted examples from a 13,000-attempt generation run. No accepted rows were trimmed from this release. Each example pairs a realistic typed or spoken shopper query… See the full description on the dataset page: https://huggingface.co/datasets/Ionio-ai/ecommerce-search-extraction.texttext-generation10K<n<100K2 likes103 downloads1mo agoHugging Face05cometadata /funding-entity-extraction-dataset-mix Funding Entity Extraction Dataset Mix Training and evaluation corpus for funding-entity extraction from full text. This dataset is the data mix used to fine-tune cometadata/funding-extraction-llama-3.1-8b-instruct-artifact-data-mix-grpo-mixed-reward on top of meta-llama/Llama-3.1-8B-Instruct in two stages: SFT on the mix described below, followed by GRPO with a hierarchical F0.5 reward. Loading the dataset from datasets import load_dataset # Default config (degraded… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-entity-extraction-dataset-mix.texttoken-classification10K<n<100K0 likes99 downloads5mo agoHugging Face06necrasov-ilya /ru-invoice-extraction-benchmark Набор для извлечения данных из русскоязычных счетов 50 синтетических русскоязычных счетов с эталонными ответами в JSON. Набор не привязан к определённой модели и предназначен для проверки систем извлечения структурированных данных из текста. Разделы Раздел Документы Назначение development 30 разработка шаблонов и примеров validation 10 выбор настроек test 10 итоговая оценка Раздел test зафиксирован для версии 1. Его нельзя использовать для… See the full description on the dataset page: https://huggingface.co/datasets/necrasov-ilya/ru-invoice-extraction-benchmark.tabulartext-generationn<1K0 likes77 downloads12d agoHugging Face07cometadata /funding-extraction-artifact-data-mix-grpo-mixed-reward Funding Extraction Training Data Training, evaluation, and test data for fine-tuning LLMs to extract structured funding metadata (funder names, award IDs, funding schemes, award titles) from academic paper funding statements. Dataset Structure data/ ├── full/ # Complete unsplit dataset │ ├── train.jsonl # 5,264 real Crossref funding statements │ └── synthetic.jsonl # 10,124 LLM-generated funding statements ├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-extraction-artifact-data-mix-grpo-mixed-reward.texttext-generationn<1K0 likes67 downloads5mo agoHugging Face08mi-obra /miobra-synthetic-construction-material-extraction-v1 Mi Obra Synthetic Construction Material Extraction v1.0 English Dataset Description This dataset contains 10,000 synthetic Spanish construction-material titles paired with structured entity-extraction targets. It is designed for Argentine construction terminology and controlled experiments in fine-tuning, teaching, information extraction, and structured generation. Language: Spanish (es) Regional context: Argentina Rows: 10,000 License: CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/mi-obra/miobra-synthetic-construction-material-extraction-v1.texttext-generation10K<n<100K0 likes64 downloads26d agoHugging Face09stindardlogic /data-extraction-sft-100k Data Extraction SFT (100K) 100,000 ShareGPT conversations demonstrating structured information extraction from unstructured text. Each example takes a real-world document (invoice, contract, resume, research abstract, meeting notes, log files) and extracts the relevant information into JSON, markdown tables, or other structured formats. Motivation Information extraction is one of the highest-value NLP tasks in enterprise settings. Common model failures include:… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-extraction-sft-100k.texttext-generation100K<n<1M0 likes61 downloads2mo agoHugging Face10nassimjp /Pashto-Brain-Extraction-Dataset 🧠 Pashto Brain Extraction Dataset A small experimental Pashto reasoning dataset designed to extract and preserve useful model reasoning/planning while discarding the final answer. Keep the brain 🧠 — throw away the mouth 🗣️ 🎯 Purpose A language model may understand a question and produce useful reasoning while still generating poor, unnatural, or grammatically incorrect Pashto in its final answer. Instead of throwing away the entire generation, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Brain-Extraction-Dataset.texttext-generationn<1K0 likes42 downloads9d agoHugging Face11TheTokenFactory /sec-extraction-multitask-v4 SEC Extraction Multitask v4 Instruction-tuning dataset for fine-tuning a small language model (e.g. Gemma 4 E2B) to extract structured data from SEC filings across three verticals: Exhibit 10 (contracts) — financial terms from executive employment, credit agreements, indemnification, licensing, and similar filings DEF 14A (proxy statements) — executive compensation, governance items, say-on-pay MD&A (10-K / 10-Q Management's Discussion & Analysis) — operating metrics, segment… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-extraction-multitask-v4.texttext-generation1K<n<10K0 likes39 downloads5mo agoHugging Face12Aulvem /japanese-invoice-receipt-extraction-eval 証憑 · Shōhyō — 日本語 証憑(請求書・領収書・支払通知書)構造化抽出ベンチマーク(無料サンプル) 証憑(しょうひょう)=取引の事実を証明する書類(請求書・領収書など)の会計用語。 「文書テキスト → JSON 抽出」パイプラインの精度を測るための 正解付き評価データセット の無料サンプルです。 config 文書タイプ 無料サンプル invoice 請求書 20 件 receipt 領収書 10 件 payment_notice 支払通知書 / 仕入明細書 15 件 いずれもインボイス制度(適格請求書等保存方式)に対応。 既存のHF日本語帳票データはOCR・画像系が中心です。本データは 画像でなくテキスト→JSON を対象にし、和暦・軽減税率・源泉徴収・収入印紙税・相殺控除のロジックを正解側で算術検証してあります。 実在の企業名・個人情報は含みません(すべて合成)。本サンプルは無料・評価/検証用途で配布します。 このサンプルの位置づけ(凍結版)… See the full description on the dataset page: https://huggingface.co/datasets/Aulvem/japanese-invoice-receipt-extraction-eval.texttext-generationn<1K0 likes35 downloads2mo agoHugging Face13cometadata /llama-3.1-8b-funding-extraction-sft-ablations LLaMA 3.1 8B Funding Extraction SFT Ablations Ablation study results for LoRA SFT of Meta LLaMA 3.1 8B Instruct on structured funding metadata extraction from scholarly text. The model extracts four fields: funder_name, award_ids, funding_scheme, and award_title. Key findings Factor Best config Avg F1 Overall best synthetic, twostage (2+1 epochs), LoRA r=64, lr=3e-5 0.588 Data type Synthetic >> non-synthetic (+0.126 avg F1) — LoRA rank r=64 > r=32 > r=16 —… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/llama-3.1-8b-funding-extraction-sft-ablations.tabulartoken-classificationn<1K0 likes34 downloads6mo agoHugging Face14Siyu2Zhou /ICKG-immunology-triple-extraction-sft ICKG 免疫学知识三元组抽取 SFT 数据集 本数据集用于从 PubMed 免疫学摘要中抽取生物医学知识三元组的指令微调(SFT)。每条样本是一段对话(system / user / assistant),assistant 即为该摘要抽取出的三元组 JSON 数组。配套的微调 adapter 见 Siyu2Zhou/Baichuan-M2-32B-QLoRA-immunology-triples。 数据规模 切分 文件 样本数 train train.jsonl 4,500 validation val.jsonl 250 test test.jsonl 250 合计 5,000 篇摘要 三元组总数 52,597,平均 10.5 条/篇(最少 3、最多 30)。 5,000 篇按「关系覆盖 + 三元组密度」分层抽样(A/B/C 三档 = 2000/2000/1000),并做关系再平衡(associated_with ≥ 35%、increases ≤… See the full description on the dataset page: https://huggingface.co/datasets/Siyu2Zhou/ICKG-immunology-triple-extraction-sft.texttext-generation1K<n<10K0 likes29 downloads3mo agoHugging Face15nielsr /paper-url-extraction-v1 Papers With Code URL Extraction A representative dataset for training and evaluating tool-using agents that find the official GitHub repository and project page for an AI research paper. It was prepared for the pwc-url-extraction-v1 Prime/verifiers environment. Splits Split Rows train 4,000 validation 500 test 500 Rows were sampled with seed 13 from up to 120,000 Papers With Code candidates, stratified by paper year and known URL state.… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/paper-url-extraction-v1.tabulartext-generation1K<n<10K0 likes29 downloads2mo agoHugging Face16acidtib /dispensary-product-name-extraction Dispensary Product Name Extraction Raw Dutchie POS listing names paired with a canonical, human-corrected clean name in "Strain - Type" form. Powers a fine-tuned Gemma 3 270M generative model that rewrites messy listing names into consistent product names, see the paired model repo, acidtib/dispensary-product-name-gemma3. Fields rawName: the raw listing name as synced from Dutchie. brand: brand name, or empty string if the listing had none. category /… See the full description on the dataset page: https://huggingface.co/datasets/acidtib/dispensary-product-name-extraction.texttext-generation1K<n<10K0 likes18 downloads1mo agoHugging Face17MJ16 /pharmacoeconomic-evidence-extraction-dataset Pharmacoeconomic Evidence Extraction Dataset License This dataset is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. You may use, share, and adapt the dataset provided that appropriate credit is given to the dataset authors. For the full license terms, see the CC BY 4.0 license. Overview This dataset contains 250 expert-annotated records for research on automated extraction of structured pharmacoeconomic and… See the full description on the dataset page: https://huggingface.co/datasets/MJ16/pharmacoeconomic-evidence-extraction-dataset.texttext-generationn<1K0 likes13 downloads1mo agoHugging Face18joaomdaltoe /medicament-extraction Medicament Extraction Este dataset contém raciocínios estruturados e extrações automáticas de princípios ativos, concentrações e formas farmacêuticas, com base em descrições regulatórias de medicamentos. Estrutura dos Subsets all — Todas as amostras processadas correct — Casos com extração correta (acerto total) incorrect — Casos com extração incorreta (erro em pelo menos um atributo) texttext-generation10K<n<100K0 likes6 downloads9mo agoHugging Face19ReDiX /reasoning-data-extractiongatedtexttext-generation10K<n<100K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.