datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-RL-Instruction-Following-Structured-Outputs-v2
Dataset Description:
Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema.
Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.sharegpt-structured-output-json
ShareGPT-Formatted Dataset for Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.structured-cpt
Structured CPT - JSON + SQL pretrain documents
SmolLM2-1.7B continued-pretraining shard of structured documents. Each document
is a <task> / <input> / <output> block whose <output> is a canonical
JSON object, terminated by the SmolLM2 end-of-text token ``.
Sources:
source
description
rows
shards
repeat
sql_bmc2
b-mc2 sql-create-context -> JSON (4 keys, stub explanation)
392,885
1
5
sql_gretelai
gretelai synthetic_text_to_sql -> JSON (4 keys)
529,255
1
5… See the full description on the dataset page: https://huggingface.co/datasets/domofon/structured-cpt.task210_logic2text_structured_text_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task210_logic2text_structured_text_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task210_logic2text_structured_text_generation.openclassgen-structured-v1
OpenClassGen Structured v1
Derived from mrahman2025/OpenClassGen (Rahman et al. 2025, arXiv:2504.15564).
License: CC BY 2.0 (same as upstream). Keep repository_name and file_path when redistributing.
Underlying GitHub repos may carry additional software licenses.
gold_code is upstream human_written_code.
We add parsed fields, body-span indices, and a Variant-3 prompt/target pair (v3_prompt_text / v3_target_text).
No unit tests. Splits are repository-disjoint (train /… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/openclassgen-structured-v1.task128_scan_structured_text_generation_command_action_short
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.task1566_propara_structured_text_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1566_propara_structured_text_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1566_propara_structured_text_generation.omnimcp_browser_dom_structured_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.task130_scan_structured_text_generation_command_action_long
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task130_scan_structured_text_generation_command_action_long
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task130_scan_structured_text_generation_command_action_long.wikimedia-enterprise-structured-contents-enwiki
enwiki_namespace_0
Structured Contents snapshot of enwiki_namespace_0 from the
Wikimedia Enterprise API, converted to Parquet.
Source
Upstream: Wikimedia Enterprise Structured Contents API
Snapshot identifier: enwiki_namespace_0
Format at source: .tar.gz containing sharded .ndjson
Shards in this release: 3
Processing
Downloaded the snapshot tarball from the Wikimedia Enterprise API.
Streamed each .ndjson shard through a normalization pass:
JSON-encoded… See the full description on the dataset page: https://huggingface.co/datasets/chuckreynolds/wikimedia-enterprise-structured-contents-enwiki.german-structured-output
German Structured Output Dataset 🇩🇪
GDPR & EU AI Act compliant German dataset for training structured output capabilities in LLMs.
Overview
This dataset contains 4,521 examples across 7 task types for training language models to produce structured outputs (JSON, function calls, schema-following generation) from German text. It is the first dedicated German structured output dataset, filling a critical gap in the German NLP ecosystem.
Key Features
🇩🇪… See the full description on the dataset page: https://huggingface.co/datasets/philipp-zettl/german-structured-output.Nemotron-RL-Instruction-Following-Structured-Outputs-v2
Dataset Description:
Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema.
Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.TCGA_Reports_ja_structured_qwen38_27b
TCGA Reports Japanese Structured Dataset with Qwen3.8-27B
The Cancer Genome Atlas(TCGA)由来の英語病理報告書を日本語へ翻訳し、その日本語病理報告書から主要な病理情報を9項目へ構造化したデータセットです。
既存の morizon/TCGA_Reports_ja_structured と同じ100症例を使用し、生成モデルを Qwen/Qwen3.8-27B に変更して再生成しています。
日本語訳には morizon/TCGA_Reports_ja_qwen38_27b と同じ生成結果を使用しています。
元データ
本データセットでは、The Cancer Genome Atlas(TCGA)の病理報告書をもとに作成されたTCGA-Reportsを使用しています。
TCGAは、複数のがん種についてゲノム情報や臨床情報などを収集した大規模ながん研究プロジェクトです。… See the full description on the dataset page: https://huggingface.co/datasets/morizon/TCGA_Reports_ja_structured_qwen38_27b.classeval-structured-v1
ClassEval Structured v1
Derived from FudanSELab/ClassEval (Du et al. 2023, arXiv:2308.01861).
License: CC BY-NC 4.0 (upstream data license). Non-commercial use only.
One row per (task_id, variant) with variant in {1,2,3} (100 tasks × 3 = 300 rows; Hub split test).
solution_code, test, and methods_info_json come from upstream.
We add rendered prompts/targets and stratification fields.
Missing bodies use ....
Variants:
Signatures and docstrings kept; every method body is ....… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/classeval-structured-v1.structured_data_merged_v2v5_0222
Dataset Card for structured_data_merged_v2v5_0222
Dataset Details
Dataset Description
structured_data_merged_v2v5_0222 is a dataset for Supervised Fine-Tuning (SFT) focused on structured data format conversion tasks — specifically, interconversion among JSON, XML, YAML, TOML, and CSV.
It was created by deduplicating and merging the following two existing datasets:
u-10bei/structured_data_with_cot_dataset_512_v2 (train split only)… See the full description on the dataset page: https://huggingface.co/datasets/takami2022/structured_data_merged_v2v5_0222.openmathinstruct-2_structured-1000
OpenMathInstruct-2 Structured (CoT)
OpenMathInstruct-2 の solution を gpt-oss-120b で再生成し、<analyze> <plan> <verify> を含む構造化 CoT を付与した SFT 用データセットです。
目的: 数学タスク向けの長手順 CoT を安定して生成するための教師あり微調整。
ライセンス: CC-BY-4.0(元データのライセンスに従います)。
データ概要
言語: 英語
レコード数: 1,446
形式: JSONL / Parquet(Hugging Face datasets 形式)
カラム
フィールド
型
説明
question
string
OpenMathInstruct-2 の問題文
answer
string
<think> ブロック内に <analyze>, <plan>, <verify>, <reason> を埋め込んだ構造化 CoT と最終回答
category… See the full description on the dataset page: https://huggingface.co/datasets/mssfj/openmathinstruct-2_structured-1000.task126_scan_structured_text_generation_command_action_all
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task126_scan_structured_text_generation_command_action_all
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task126_scan_structured_text_generation_command_action_all.combined-structured-dataset
Combined Structured Output Dataset
このデータセットは、構造化出力生成タスクのための統合データセットです。
Dataset Details
Total Samples: 3,425
Source Datasets:
u-10bei/structured_data_with_cot_dataset_v2 (2,500 samples)
daichira/structured-3k-mix-sft (3,000 samples)
Preprocessing:
Format normalization to 3-turn (system/user/assistant)
Deduplication by user content
Quality filtering
Format Distribution
JSON: 685 (20.0%)
YAML: 485 (14.2%)
TOML: 685 (20.0%)
XML: 885 (25.8%)
CSV: 685… See the full description on the dataset page: https://huggingface.co/datasets/kurota0612/combined-structured-dataset.testing-wiki-structured
cywiki_namespace_0
Structured Contents snapshot of cywiki_namespace_0 from the
Wikimedia Enterprise API,
repackaged as Parquet with a pinned schema.
The upstream Wikimedia Foundation dataset
(wikimedia/structured-wikipedia)
ships NDJSON which has known issues loading via
datasets.load_dataset() — see discussions
#5,
#15,
#16.
This dataset is the same upstream content, normalised so
load_dataset(...)works without specifying a Features override.
Source
Upstream: Wikimedia… See the full description on the dataset page: https://huggingface.co/datasets/VoeTheDon/testing-wiki-structured.structured_poem_interpretation_corpus_staleMasking policy: For rows with source == "poetry_foundation", the poem and
interpretation fields are set to null to respect content licensing. Public-domain
entries (source == "public_domain_poetry") include full text. All categorical annotations
(emotions, primary_emotion, sentiment, themes, themes_50) and metadata remain available.
harmonicbench-planir-main-structured
HARMONICBench PlanIR Structured Main Dataset
This repository is a Hugging Face dataset-friendly structured export derived from the local outputs/fixed/main directory in the HARMONICBench unified package.
Included tables
plans/train.parquet: primary aggregate table converted from plan_runs_all.jsonl.
domains/*.parquet: per-domain plan runs, selected samples, and domain5 image descriptions.
summaries/*: key JSON/CSV/JSONL summary artifacts.
artifacts/roundtrip_recovery*/*:… See the full description on the dataset page: https://huggingface.co/datasets/guhhhgu/harmonicbench-planir-main-structured.nemotron-gym-structured-outputs-v3
laion/nemotron-gym-structured-outputs-v3
Harbor task-binary dataset (53,870 tasks) converted from nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2
(part of the nvidia/Nemotron-Post-Training-v3 collection).
Each row is a valid Harbor
task binary: columns path (str) and task_binary (gzip tar). Converted with the
OpenThoughts-Agent data.nemotron_gym framework.
Grading: JSON/YAML/TOML schema validation; XML/CSV structural (well-formed + required keys).
nemotron-gym-structured-outputs-v4
laion/nemotron-gym-structured-outputs-v4
Harbor task-binary dataset (53,870 tasks) converted from nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2
(part of nvidia/Nemotron-Post-Training-v3).
Columns path (str) + task_binary (gzip tar). Converted with the
OpenThoughts-Agent data.nemotron_gym framework.
Grading: JSON/YAML/TOML schema validation; XML/CSV structural.
What changed vs the prior version
This version fixes the answer-delivery contract for… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-structured-outputs-v4.TOON-Unstructured-Structured
TOON-Unstructured-Structured
This dataset is a validated and cleaned version of the originalMasterControlAIML/JSON-Unstructured-Structured.
It has been reformatted using the official Token-Oriented Object Notation (TOON) specification —a compact, token-efficient data serialization format optimized for LLM-ready structured data.All records have been verified for JSON integrity and TOON-decoding consistency.
Overview
Field
Description
text
Original text… See the full description on the dataset page: https://huggingface.co/datasets/yasserrmd/TOON-Unstructured-Structured.
