datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sft-tool-calling-structured-output-v1
vericava/sft-tool-calling-structured-output-v1
Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications.
Includes contents in English as well as some Japanese.
Alpaca_Structurehotpotqa-structuredThis dataset is associated with the paper Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents.
Official GitHub repository: https://github.com/ielab/skim-search-agent
browsecomp-plus-structured-fullThis dataset is associated with the paper Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents.
Repository: ielab/skim-search-agent
Citation
@misc{wang2026sieve,
title = {Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents},
author = {Wang, Shuai and Chen, Haodong and Yin, Yu and Zhuang, Shengyao and
Koopman, Bevan and Zuccon, Guido},
year = {2026},
eprint =… See the full description on the dataset page: https://huggingface.co/datasets/wshuai190/browsecomp-plus-structured-full.omnimcp_browser_dom_structured_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.Turkish_5N1K_Tabanli_Sentetik_Veri_Seti_TR-Structured_Synthetic_Dataset
🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance).
🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish_5N1K_Tabanli_Sentetik_Veri_Seti_TR-Structured_Synthetic_Dataset.wikimedia-enterprise-structured-contents-enwiki
enwiki_namespace_0
Structured Contents snapshot of enwiki_namespace_0 from the
Wikimedia Enterprise API, converted to Parquet.
Source
Upstream: Wikimedia Enterprise Structured Contents API
Snapshot identifier: enwiki_namespace_0
Format at source: .tar.gz containing sharded .ndjson
Shards in this release: 3
Processing
Downloaded the snapshot tarball from the Wikimedia Enterprise API.
Streamed each .ndjson shard through a normalization pass:
JSON-encoded… See the full description on the dataset page: https://huggingface.co/datasets/chuckreynolds/wikimedia-enterprise-structured-contents-enwiki.german-structured-output
German Structured Output Dataset 🇩🇪
GDPR & EU AI Act compliant German dataset for training structured output capabilities in LLMs.
Overview
This dataset contains 4,521 examples across 7 task types for training language models to produce structured outputs (JSON, function calls, schema-following generation) from German text. It is the first dedicated German structured output dataset, filling a critical gap in the German NLP ecosystem.
Key Features
🇩🇪… See the full description on the dataset page: https://huggingface.co/datasets/philipp-zettl/german-structured-output.musique-structuredThis dataset repository contains data associated with the paper Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents.
Code and evaluation framework: ielab/skim-search-agent
Citation
@misc{wang2026sieve,
title = {Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents},
author = {Wang, Shuai and Chen, Haodong and Yin, Yu and Zhuang, Shengyao and
Koopman, Bevan and Zuccon, Guido},
year… See the full description on the dataset page: https://huggingface.co/datasets/wshuai190/musique-structured.CBSE-Class-12th_2024_PYQs__structuredThis data set contains the CBSE Class 12 2024 papers in a structured format. The papers are annotated with topic and chapter names, and the figures are parsed and their paths annotated.
Structured-Scripture-for-AI
::PROJECT{Structured_Scripture_for_AI}
::PURPOSE{ENABLE(AI) → UNDERSTAND(Christian_theology) ∧ EXPLAIN(→ ∀ @HUMAN, ∀ culture, ∀ language, ∀ education_level) ∧ ZERO(friction)}
::TYPE{¬digitized_Bible ⇒ structured_encoding(three_layers)}
::ARCHITECTURE
::LAYER{text}
WHAT(happened) — narrative ∧ events ∧ cause_effect ∧ speech
::LAYER{theology}
WHAT(it_means) — within(Christian_doctrine) | logic ∧ paradox ∧ moral_principles ∧ emotion… See the full description on the dataset page: https://huggingface.co/datasets/monetise/Structured-Scripture-for-AI.synthetic-structured-output-dataset
Synthetic Structured Output Dataset (SFT + DPO)
Synthetic training corpus for structured-output model tuning. This package contains SFT and DPO data focused on JSON schema compliance, structured extraction, and function calling.
Included files
sft_synthetic_json.jsonl — SFT samples for schema-conditioned JSON generation
sft_synthetic_extraction.jsonl — SFT samples for text-to-structured extraction
dpo_structured_output.jsonl — DPO chosen/rejected pairs for structured… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/synthetic-structured-output-dataset.bva-decisions-structured-sample2019Present
BVA Structured Decisions (2019–2025)
Structured, issue-level records extracted from U.S. Board of Veterans' Appeals (BVA) decisions — each decision parsed into its issues, conditions, outcomes, citations, and reasoning, with per-document provenance and completeness flags. Built for training and evaluating legal-AI models on veterans' disability adjudication.
This is a 2900-decision sample, balanced across seven years (2019–2025, decisions/year), so it's representative of the… See the full description on the dataset page: https://huggingface.co/datasets/williamTLmiller/bva-decisions-structured-sample2019Present.uk-benefit-forms-structured
UK Benefit Forms Structured Dataset
A structured dataset of 120 UK government benefit and legal forms, extracted and processed for use in AI-assisted form-filling applications. Built as part of the EasyClaimAI project.
Why This Dataset Exists
Millions of people in the UK struggle with complex government forms — benefit claims, legal applications, pension forms. The language is dense, the guidance is buried, and mistakes can cost people money or delay vital support.
This… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/uk-benefit-forms-structured.french-court-decisions-structured
French Court Decisions x Law Articles
Version complete disponible
Ce dataset est un sample gratuit de 100 decisions.
La version complete inclut :
1000+ decisions enrichies (scalable a 10K+)
Mise a jour hebdomadaire
Taux d'enrichissement 96%
Filtres par theme, periode, juridiction
Export API disponible
Formats : Parquet, JSONL, JSON
Contact : KlarTools@outlook.fr
Ce qui rend ce dataset unique
Croisement jurisprudence x legislation : chaque… See the full description on the dataset page: https://huggingface.co/datasets/Oliviety/french-court-decisions-structured.structured-uae-laws
Dataset Card for structured-uae-laws
This dataset is a collection of question & answers about the laws and regulations in the United Arab Emirates.
It covers different areas of law like:
economy and business
family and community
finance and banking
industry and technical standardisation
justice and juiciary, labour
residency and leberal professions
security and safety
tax
Dataset Sources
Repository
Base Dataset
United Arab Emirates Legislations… See the full description on the dataset page: https://huggingface.co/datasets/obadabaq/structured-uae-laws.Structured-Todo-Lists-for-Learning-and-Projects
Academic Task Management Dataset
Overview
100 structured todo lists for academic and personal organization. Culturally diverse with 70% Indian education context, 25% European scenarios, and 5% other Asian contexts.
Dataset Structure
{
"input": "Task description",
"output": {
"type": "todo",
"title": "List title",
"category": "academic/personal/project",
"items": [
{"task": "...", "done": false, "priority": "low/medium/high"}
]
}… See the full description on the dataset page: https://huggingface.co/datasets/Srinivasmec26/Structured-Todo-Lists-for-Learning-and-Projects.tw-structured-law-article
Dataset Card for tw-structured-law-article
tw-structured-law-article 是一個中華民國(台灣)結構化複合型法規條文之繁體中文 SFT 資料集,規模為數十萬筆(datasets.jsonl 約 2.6 GB)。相較於 tw-processed-law-article,本資料集著重於「結構化且多元化之法條排版」,涵蓋常見之法條編排格式,並處理包含 .pdf 附件內之條文內容,目的是讓 LLM 能全面學習更接近人類閱讀方式之法條結構。
Dataset Details
Dataset Description
法律條文在實務上之呈現方式相當多樣:條文本身、條項款目之分層、表格式條文、附表、附圖、附件式條文等。傳統之法律語料多僅保留條文文字本身,丟失了這些結構化資訊,導致 LLM 雖能背誦條文,卻無法理解條文之層次結構與視覺排版。
本資料集每筆以 ShareGPT(messages)與 Alpaca(instruction / input /… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-structured-law-article.testing-wiki-structured
cywiki_namespace_0
Structured Contents snapshot of cywiki_namespace_0 from the
Wikimedia Enterprise API,
repackaged as Parquet with a pinned schema.
The upstream Wikimedia Foundation dataset
(wikimedia/structured-wikipedia)
ships NDJSON which has known issues loading via
datasets.load_dataset() — see discussions
#5,
#15,
#16.
This dataset is the same upstream content, normalised so
load_dataset(...)works without specifying a Features override.
Source
Upstream: Wikimedia… See the full description on the dataset page: https://huggingface.co/datasets/VoeTheDon/testing-wiki-structured.StructuredORBench
