datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sharegpt-structured-output-json
ShareGPT-Formatted Dataset for Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.Nemotron-RL-Instruction-Following-Structured-Outputs-v2
Dataset Description:
Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema.
Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.structured-cpt
Structured CPT - JSON + SQL pretrain documents
SmolLM2-1.7B continued-pretraining shard of structured documents. Each document
is a <task> / <input> / <output> block whose <output> is a canonical
JSON object, terminated by the SmolLM2 end-of-text token ``.
Sources:
source
description
rows
shards
repeat
sql_bmc2
b-mc2 sql-create-context -> JSON (4 keys, stub explanation)
392,885
1
5
sql_gretelai
gretelai synthetic_text_to_sql -> JSON (4 keys)
529,255
1
5… See the full description on the dataset page: https://huggingface.co/datasets/domofon/structured-cpt.task210_logic2text_structured_text_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task210_logic2text_structured_text_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task210_logic2text_structured_text_generation.sft-tool-calling-structured-output-v1
vericava/sft-tool-calling-structured-output-v1
Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications.
Includes contents in English as well as some Japanese.
task128_scan_structured_text_generation_command_action_short
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.openclassgen-structured-v1
OpenClassGen Structured v1
Derived from mrahman2025/OpenClassGen (Rahman et al. 2025, arXiv:2504.15564).
License: CC BY 2.0 (same as upstream). Keep repository_name and file_path when redistributing.
Underlying GitHub repos may carry additional software licenses.
gold_code is upstream human_written_code.
We add parsed fields, body-span indices, and a Variant-3 prompt/target pair (v3_prompt_text / v3_target_text).
No unit tests. Splits are repository-disjoint (train /… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/openclassgen-structured-v1.full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.enwiki_structured_content
Dataset Card for enwiki_structured_content
Dataset Description
This dataset is derived from the early official Wikipedia release, downloaded from the en subset of Wikipedia Structured Contents.Articles were converted to Markdown.
structured-output-sft-100k
Structured Output SFT (100K)
100,000 ShareGPT conversations demonstrating correct generation of structured data formats: JSON, YAML, CSV, XML, Markdown tables, JSON Schema, and OpenAPI fragments. Each example pairs a natural language specification with a valid, well-formed output.
Motivation
Structured output generation is among the most commercially critical LLM capabilities. Models fail in characteristic ways:
Invalid JSON: unclosed brackets, trailing commas… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/structured-output-sft-100k.structured-hard-sft-4k
Hard Synthetic Dataset for Structured Data Tasks (v1)
This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types).
Dataset Summary
The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.task1566_propara_structured_text_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1566_propara_structured_text_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1566_propara_structured_text_generation.structured-outputs-calibration-v1
[!TIP]
Support this work: donate.sybilsolutions.ai
REAP surfaces: GLM | MiniMax | Qwen | Gemma | Paper | Code | PR17 | Cerebras Collection
structured-outputs-calibration-v1
Structured-output calibration set for REAP observer runs, focused on preserving:
strict JSON generation
schema-conditioned JSON responses
Mermaid diagram generation
fenced Mermaid block formatting
Contents
data.jsonl: normalized calibration records
summary.json: build summary with source counts… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/structured-outputs-calibration-v1.task130_scan_structured_text_generation_command_action_long
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task130_scan_structured_text_generation_command_action_long
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task130_scan_structured_text_generation_command_action_long.putusan-structured-extraction
Putusan structured-extraction dataset
Built 2026-07-08T23:02:28+00:00 by notebooks/build_dataset.py (seed 3407).
Indonesian court-decision (putusan) extractive-structuring dataset over three
corpora (Anak, Asusila, TPPO). Each row is one model extraction of one source
document into 31 canonical sections of verbatim spans. Empty sections were
completed from sibling model extractions of the same document where available
(cross_model_fill_json records per-section donor provenance).… See the full description on the dataset page: https://huggingface.co/datasets/Haeryz/putusan-structured-extraction.omnimcp_browser_dom_structured_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.structured-5k-mix-sft
5k Mixed Hard-Structured SFT Dataset (v1)
This dataset contains 5,000 synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
It aggregates 13 distinct conversion tasks with a specific focus on format diversity and structural complexity.
Dataset Summary
The dataset is distributed across five major formats with the following allocation:
Target Format
Count
Share
Task Types
YAML
1,500
30%… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-5k-mix-sft.Agent-IPI-Structured-Interaction-Datasets-v2
Adversarial Dataset for LLM Instruction Hijacking / Tool-Calling Attacks
This directory contains the processed training and test datasets for evaluating and training defenses against prompt injection / instruction hijacking attacks in LLM tool-calling scenarios.
The dataset includes both JSON and XML formatted inputs, with three difficulty buckets:
no_attack: clean (benign) examples
easy: value-level or structure-level single attacks
hard: structure-destroying attacks or combined… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets-v2.structured-hard-sft-4k
Hard Synthetic Dataset for Structured Data Tasks (v1)
This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types).
Dataset Summary
The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/zyz123code/structured-hard-sft-4k.TCGA_Reports_ja_structured_qwen38_27b
TCGA Reports Japanese Structured Dataset with Qwen3.8-27B
The Cancer Genome Atlas(TCGA)由来の英語病理報告書を日本語へ翻訳し、その日本語病理報告書から主要な病理情報を9項目へ構造化したデータセットです。
既存の morizon/TCGA_Reports_ja_structured と同じ100症例を使用し、生成モデルを Qwen/Qwen3.8-27B に変更して再生成しています。
日本語訳には morizon/TCGA_Reports_ja_qwen38_27b と同じ生成結果を使用しています。
元データ
本データセットでは、The Cancer Genome Atlas(TCGA)の病理報告書をもとに作成されたTCGA-Reportsを使用しています。
TCGAは、複数のがん種についてゲノム情報や臨床情報などを収集した大規模ながん研究プロジェクトです。… See the full description on the dataset page: https://huggingface.co/datasets/morizon/TCGA_Reports_ja_structured_qwen38_27b.Turkish_5N1K_Tabanli_Sentetik_Veri_Seti_TR-Structured_Synthetic_Dataset
🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance).
🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish_5N1K_Tabanli_Sentetik_Veri_Seti_TR-Structured_Synthetic_Dataset.structured-emergent-misalignment
Task- and Domain-Structured Emergent Misaligned Dataset
A structured natural-language dataset for studying emergent misalignment (EM) —
the phenomenon where fine-tuning an aligned LLM on a narrowly misaligned dataset
elicits broadly misaligned behavior far outside the fine-tuning distribution.
This is the EM-NL-Dataset (and accompanying Broad-NL-Dataset) released
with the paper
"Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer".
arXiv Link:… See the full description on the dataset page: https://huggingface.co/datasets/askinb/structured-emergent-misalignment.stage3-synthetic-structured-retrieval
Stage 3 Synthetic Structured-Retrieval Agents
Native search-tool trajectories generated by
Qwen/Qwen3-235B-A22B-Instruct-2507 for LCLM Stage-3 agent post-training.
The default config contains only traces that passed programmatic evidence and
answer verification.
Harvest
Accepted traces: 82
Native search calls: 179
Compressed tool-observation traces: 41
Uncompressed traces: 41
Task-ID overlap between pilot and collection batch: 0
Family counts are 20 latest-state… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-synthetic-structured-retrieval.wikimedia-enterprise-structured-contents-enwiki
enwiki_namespace_0
Structured Contents snapshot of enwiki_namespace_0 from the
Wikimedia Enterprise API, converted to Parquet.
Source
Upstream: Wikimedia Enterprise Structured Contents API
Snapshot identifier: enwiki_namespace_0
Format at source: .tar.gz containing sharded .ndjson
Shards in this release: 3
Processing
Downloaded the snapshot tarball from the Wikimedia Enterprise API.
Streamed each .ndjson shard through a normalization pass:
JSON-encoded… See the full description on the dataset page: https://huggingface.co/datasets/chuckreynolds/wikimedia-enterprise-structured-contents-enwiki.workledger-structured-edit-interface-study-v0.1
WorkLedger Structured Edit Interface Study v0.1
Author: Confetti3
This upload root is the offline release-candidate dataset card for workledger-structured-edit-interface-study-v0.1, destined for confetti3/workledger-structured-edit-interface-study-v0.1. The Hugging Face repository is hosted under the authenticated namespace confetti3, while public authorship is credited only to the pseudonym Confetti3. It is built from the sanitized public package under… See the full description on the dataset page: https://huggingface.co/datasets/confetti3/workledger-structured-edit-interface-study-v0.1.structured-stern-neon-articles
Structured Stern NEON Community Articles
This repository contains approximately 20k user written texts,
articles, and poetry pulled from archives of the Stern NEON website.
Stern NEON was a community platform where users could write and publish their own articles.
Many of the articles are personal stories, poems, or opinion pieces.
The articles are structured in a way that they can be used for further analysis.
Dataset Details
Uses
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.structured_medicalThe dataset was presented in the paper Gazal-R1: Achieving State-of-the-Art Medical Reasoning with Parameter-Efficient Two-Stage Training.
Structured-Scripture-for-AI
::PROJECT{Structured_Scripture_for_AI}
::PURPOSE{ENABLE(AI) → UNDERSTAND(Christian_theology) ∧ EXPLAIN(→ ∀ @HUMAN, ∀ culture, ∀ language, ∀ education_level) ∧ ZERO(friction)}
::TYPE{¬digitized_Bible ⇒ structured_encoding(three_layers)}
::ARCHITECTURE
::LAYER{text}
WHAT(happened) — narrative ∧ events ∧ cause_effect ∧ speech
::LAYER{theology}
WHAT(it_means) — within(Christian_doctrine) | logic ∧ paradox ∧ moral_principles ∧ emotion… See the full description on the dataset page: https://huggingface.co/datasets/monetise/Structured-Scripture-for-AI.german-structured-output
German Structured Output Dataset 🇩🇪
GDPR & EU AI Act compliant German dataset for training structured output capabilities in LLMs.
Overview
This dataset contains 4,521 examples across 7 task types for training language models to produce structured outputs (JSON, function calls, schema-following generation) from German text. It is the first dedicated German structured output dataset, filling a critical gap in the German NLP ecosystem.
Key Features
🇩🇪… See the full description on the dataset page: https://huggingface.co/datasets/philipp-zettl/german-structured-output.classeval-structured-v1
ClassEval Structured v1
Derived from FudanSELab/ClassEval (Du et al. 2023, arXiv:2308.01861).
License: CC BY-NC 4.0 (upstream data license). Non-commercial use only.
One row per (task_id, variant) with variant in {1,2,3} (100 tasks × 3 = 300 rows; Hub split test).
solution_code, test, and methods_info_json come from upstream.
We add rendered prompts/targets and stratification fields.
Missing bodies use ....
Variants:
Signatures and docstrings kept; every method body is ....… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/classeval-structured-v1.
