datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sec-contracts-financial-extraction-instructions
S&P 500 SEC Financial Extraction Instructions
Dataset Summary
7,683 instruction-tuning examples for training LLMs to extract structured financial data from SEC filings. Covers two filing types across S&P 500 companies:
Split
Examples
Filing Type
Description
train
3,430
Exhibit 10 + DEF 14A
Positive examples with validated outputs
corrective
4,253
Exhibit 10 + DEF 14A
Corrective, rescued, and negative examples
Exhibit 10 — Material Contracts (2… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-financial-extraction-instructions.donto-qwen3.8-27b-predicate-extraction-data
Donto-Qwen3.8 Predicate Extraction Data V15
This repository is the complete public data and evidence companion to
ajaxdavis/donto-qwen3.8-27b-predicate-extractor.
It contains the canonical V15 extraction training/validation corpus, the
validator corpus, the optional D1-repeat ablation, the once-sealed 100-document
graph-first gold suite, exact tool schemas, generator/evaluator source, hashes,
and audit reports.
Why this dataset exists
Donto is designed for… See the full description on the dataset page: https://huggingface.co/datasets/ajaxdavis/donto-qwen3.8-27b-predicate-extraction-data.sec-contracts-corrective-extraction
S&P 500 SEC Financial Extractions - Corrective Dataset
Dataset Summary
4,253 corrective instruction-tuning examples designed to teach LLMs what the base model gets wrong when extracting structured financial data from SEC filings. Covers both Exhibit 10 material contracts and DEF 14A proxy statements from S&P 500 companies.
This is a companion to TheTokenFactory/sec-contracts-financial-extraction-instructions, which contains the positive training examples.
Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-corrective-extraction.ecommerce-search-extraction
Ionio E-commerce Search Query Extraction
Built with: simula — schema-driven synthetic data generation with auditable taxonomy lineage.
An English synthetic dataset for training and evaluating systems that translate natural-language
shopping requests into narrow, atomic, database-queryable JSON. It contains 10,985 accepted
examples from a 13,000-attempt generation run. No accepted rows were trimmed from this release.
Each example pairs a realistic typed or spoken shopper query… See the full description on the dataset page: https://huggingface.co/datasets/Ionio-ai/ecommerce-search-extraction.funding-entity-extraction-dataset-mix
Funding Entity Extraction Dataset Mix
Training and evaluation corpus for funding-entity extraction from full text. This dataset is the data mix used to fine-tune cometadata/funding-extraction-llama-3.1-8b-instruct-artifact-data-mix-grpo-mixed-reward on top of meta-llama/Llama-3.1-8B-Instruct in two stages: SFT on the mix described below, followed by GRPO with a hierarchical F0.5 reward.
Loading the dataset
from datasets import load_dataset
# Default config (degraded… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-entity-extraction-dataset-mix.ru-invoice-extraction-benchmark
Набор для извлечения данных из русскоязычных счетов
50 синтетических русскоязычных счетов с эталонными ответами в JSON. Набор не привязан к определённой модели и предназначен для проверки систем извлечения структурированных данных из текста.
Разделы
Раздел
Документы
Назначение
development
30
разработка шаблонов и примеров
validation
10
выбор настроек
test
10
итоговая оценка
Раздел test зафиксирован для версии 1. Его нельзя использовать для… See the full description on the dataset page: https://huggingface.co/datasets/necrasov-ilya/ru-invoice-extraction-benchmark.funding-extraction-artifact-data-mix-grpo-mixed-reward
Funding Extraction Training Data
Training, evaluation, and test data for fine-tuning LLMs to extract structured funding metadata (funder names, award IDs, funding schemes, award titles) from academic paper funding statements.
Dataset Structure
data/
├── full/ # Complete unsplit dataset
│ ├── train.jsonl # 5,264 real Crossref funding statements
│ └── synthetic.jsonl # 10,124 LLM-generated funding statements
├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-extraction-artifact-data-mix-grpo-mixed-reward.miobra-synthetic-construction-material-extraction-v1
Mi Obra Synthetic Construction Material Extraction v1.0
English
Dataset Description
This dataset contains 10,000 synthetic Spanish construction-material titles
paired with structured entity-extraction targets. It is designed for Argentine
construction terminology and controlled experiments in fine-tuning, teaching,
information extraction, and structured generation.
Language: Spanish (es)
Regional context: Argentina
Rows: 10,000
License: CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/mi-obra/miobra-synthetic-construction-material-extraction-v1.data-extraction-sft-100k
Data Extraction SFT (100K)
100,000 ShareGPT conversations demonstrating structured information extraction from unstructured text. Each example takes a real-world document (invoice, contract, resume, research abstract, meeting notes, log files) and extracts the relevant information into JSON, markdown tables, or other structured formats.
Motivation
Information extraction is one of the highest-value NLP tasks in enterprise settings. Common model failures include:… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-extraction-sft-100k.Pashto-Brain-Extraction-Dataset
🧠 Pashto Brain Extraction Dataset
A small experimental Pashto reasoning dataset designed to extract and preserve useful model reasoning/planning while discarding the final answer.
Keep the brain 🧠 — throw away the mouth 🗣️
🎯 Purpose
A language model may understand a question and produce useful reasoning while still generating poor, unnatural, or grammatically incorrect Pashto in its final answer.
Instead of throwing away the entire generation, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Brain-Extraction-Dataset.sec-extraction-multitask-v4
SEC Extraction Multitask v4
Instruction-tuning dataset for fine-tuning a small language model (e.g. Gemma 4 E2B) to extract structured data from SEC filings across three verticals:
Exhibit 10 (contracts) — financial terms from executive employment, credit agreements, indemnification, licensing, and similar filings
DEF 14A (proxy statements) — executive compensation, governance items, say-on-pay
MD&A (10-K / 10-Q Management's Discussion & Analysis) — operating metrics, segment… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-extraction-multitask-v4.japanese-invoice-receipt-extraction-eval
証憑 · Shōhyō — 日本語 証憑(請求書・領収書・支払通知書)構造化抽出ベンチマーク(無料サンプル)
証憑(しょうひょう)=取引の事実を証明する書類(請求書・領収書など)の会計用語。
「文書テキスト → JSON 抽出」パイプラインの精度を測るための 正解付き評価データセット の無料サンプルです。
config
文書タイプ
無料サンプル
invoice
請求書
20 件
receipt
領収書
10 件
payment_notice
支払通知書 / 仕入明細書
15 件
いずれもインボイス制度(適格請求書等保存方式)に対応。
既存のHF日本語帳票データはOCR・画像系が中心です。本データは 画像でなくテキスト→JSON を対象にし、和暦・軽減税率・源泉徴収・収入印紙税・相殺控除のロジックを正解側で算術検証してあります。
実在の企業名・個人情報は含みません(すべて合成)。本サンプルは無料・評価/検証用途で配布します。
このサンプルの位置づけ(凍結版)… See the full description on the dataset page: https://huggingface.co/datasets/Aulvem/japanese-invoice-receipt-extraction-eval.llama-3.1-8b-funding-extraction-sft-ablations
LLaMA 3.1 8B Funding Extraction SFT Ablations
Ablation study results for LoRA SFT of Meta LLaMA 3.1 8B Instruct on structured funding metadata extraction from scholarly text.
The model extracts four fields: funder_name, award_ids, funding_scheme, and award_title.
Key findings
Factor
Best config
Avg F1
Overall best
synthetic, twostage (2+1 epochs), LoRA r=64, lr=3e-5
0.588
Data type
Synthetic >> non-synthetic (+0.126 avg F1)
—
LoRA rank
r=64 > r=32 > r=16
—… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/llama-3.1-8b-funding-extraction-sft-ablations.ICKG-immunology-triple-extraction-sft
ICKG 免疫学知识三元组抽取 SFT 数据集
本数据集用于从 PubMed 免疫学摘要中抽取生物医学知识三元组的指令微调(SFT)。每条样本是一段对话(system / user / assistant),assistant 即为该摘要抽取出的三元组 JSON 数组。配套的微调 adapter 见 Siyu2Zhou/Baichuan-M2-32B-QLoRA-immunology-triples。
数据规模
切分
文件
样本数
train
train.jsonl
4,500
validation
val.jsonl
250
test
test.jsonl
250
合计
5,000 篇摘要
三元组总数 52,597,平均 10.5 条/篇(最少 3、最多 30)。
5,000 篇按「关系覆盖 + 三元组密度」分层抽样(A/B/C 三档 = 2000/2000/1000),并做关系再平衡(associated_with ≥ 35%、increases ≤… See the full description on the dataset page: https://huggingface.co/datasets/Siyu2Zhou/ICKG-immunology-triple-extraction-sft.paper-url-extraction-v1
Papers With Code URL Extraction
A representative dataset for training and evaluating tool-using agents that
find the official GitHub repository and project page for an AI research paper.
It was prepared for the pwc-url-extraction-v1 Prime/verifiers environment.
Splits
Split
Rows
train
4,000
validation
500
test
500
Rows were sampled with seed 13 from up to 120,000 Papers With Code candidates,
stratified by paper year and known URL state.… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/paper-url-extraction-v1.dispensary-product-name-extraction
Dispensary Product Name Extraction
Raw Dutchie POS listing names paired with a canonical, human-corrected
clean name in "Strain - Type" form. Powers a fine-tuned Gemma 3 270M
generative model that rewrites messy listing names into consistent
product names, see the paired model repo,
acidtib/dispensary-product-name-gemma3.
Fields
rawName: the raw listing name as synced from Dutchie.
brand: brand name, or empty string if the listing had none.
category /… See the full description on the dataset page: https://huggingface.co/datasets/acidtib/dispensary-product-name-extraction.pharmacoeconomic-evidence-extraction-dataset
Pharmacoeconomic Evidence Extraction Dataset
License
This dataset is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. You may use, share, and adapt the dataset provided that appropriate credit is given to the dataset authors.
For the full license terms, see the CC BY 4.0 license.
Overview
This dataset contains 250 expert-annotated records for research on automated extraction of structured pharmacoeconomic and… See the full description on the dataset page: https://huggingface.co/datasets/MJ16/pharmacoeconomic-evidence-extraction-dataset.medicament-extraction
Medicament Extraction
Este dataset contém raciocínios estruturados e extrações automáticas de princípios ativos, concentrações e formas farmacêuticas, com base em descrições regulatórias de medicamentos.
Estrutura dos Subsets
all — Todas as amostras processadas
correct — Casos com extração correta (acerto total)
incorrect — Casos com extração incorreta (erro em pelo menos um atributo)
reasoning-data-extraction
