datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
donto-qwen3.8-27b-predicate-extraction-data
Donto-Qwen3.8 Predicate Extraction Data V15
This repository is the complete public data and evidence companion to
ajaxdavis/donto-qwen3.8-27b-predicate-extractor.
It contains the canonical V15 extraction training/validation corpus, the
validator corpus, the optional D1-repeat ablation, the once-sealed 100-document
graph-first gold suite, exact tool schemas, generator/evaluator source, hashes,
and audit reports.
Why this dataset exists
Donto is designed for… See the full description on the dataset page: https://huggingface.co/datasets/ajaxdavis/donto-qwen3.8-27b-predicate-extraction-data.funding-extraction-artifact-data-mix-grpo-mixed-reward
Funding Extraction Training Data
Training, evaluation, and test data for fine-tuning LLMs to extract structured funding metadata (funder names, award IDs, funding schemes, award titles) from academic paper funding statements.
Dataset Structure
data/
├── full/ # Complete unsplit dataset
│ ├── train.jsonl # 5,264 real Crossref funding statements
│ └── synthetic.jsonl # 10,124 LLM-generated funding statements
├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-extraction-artifact-data-mix-grpo-mixed-reward.feature_extraction_training_datadata-extraction-sft-100k
Data Extraction SFT (100K)
100,000 ShareGPT conversations demonstrating structured information extraction from unstructured text. Each example takes a real-world document (invoice, contract, resume, research abstract, meeting notes, log files) and extracts the relevant information into JSON, markdown tables, or other structured formats.
Motivation
Information extraction is one of the highest-value NLP tasks in enterprise settings. Common model failures include:… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-extraction-sft-100k.ecotopia-extraction-data
Ecotopia Extraction Data
Training dataset for the Ecotopia promise extraction model. Contains mayor speeches paired with structured JSON extractions of political promises and contradiction detection.
Dataset Details
Size: 200 examples (160 train / 40 validation)
Format: Conversational (system/user/assistant messages)
Task: Extract promises (text, type, impact) and detect contradictions from free-text speeches
Links
Extract Model
GitHub Repo
feature_extraction_training_data_unilever_v1.1feature_extraction_training_data_unileverfunding-extraction-sft-data
Funding Extraction SFT Data
Training data for supervised fine-tuning of funding statement extraction models. Given a funding statement from a scholarly work, the task is to extract structured funder information including funder names, award IDs, funding schemes, and award titles.
Dataset Structure
Files
File
Examples
Description
train.jsonl
1,316
Real funding statements from Crossref metadata (with DOIs)
synthetic.jsonl
2,531
Synthetically… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-extraction-sft-data.feature_extraction_training_data_function_callinguniliver_product_extraction_training_data_v2.1data-extraction-testreasoning-data-extraction
