datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
structured-wikipedia
Dataset Card for Wikimedia Structured Wikipedia
Quick Links
Wikimedia Enterprise
Structured Contents Documentation
Data Dictionary
Wikimedia Attribution Framework
Meta-Wiki Discussion
Dataset Summary
Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API.
This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/structured-wikipedia.sharegpt-structured-output-json
ShareGPT-Formatted Dataset for Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.Nemotron-RL-Instruction-Following-Structured-Outputs-v2
Dataset Description:
Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema.
Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.a3-rl-laion_nemotron-gym-instruction-following-structuredstructured-file-audit-benchmark
Paper Data Release
This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them.
Contents
datasets/
Benchmark data and per-task manifests for the three paper-facing splits.
datasets/sc_flat/data
SC-Flat is derived from DaBench, augmented with a replayable perturbation
injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.Nemotron-RL-instruction_following-structured_outputs
Dataset Description:
The Nemotron-RL-instruction_following-structured_outputs dataset tests the ability of the model to follow output formatting instructions under schema constraints under the JSON format. Each problem consists of three components: The document, output formatting Instruction (Schema), and question. The dataset varies the difficulty of each problem by varying the location of instructions, the comprehensiveness of instructions, the complexity of the schema, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-instruction_following-structured_outputs.pseudo-camera-10k-structured-json
pseudo-camera-10k, structured JSON captions
The 9,997 training images from bghira/pseudo-camera-10k, recaptioned into the structured JSON caption schema that Ideogram 4 consumes. The images are unchanged: free photographs from world class photographers, Lanczos-resized so the shorter edge is 1024px, nothing upsampled.
The original dataset carries short CogVLM prose captions. This one replaces them with one JSON object per image describing the scene at three levels: an overall… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/pseudo-camera-10k-structured-json.nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces.structured-wikipedia
Dataset Card for Wikimedia Structured Wikipedia
Quick Links
Wikimedia Enterprise
Structured Contents Documentation
Data Dictionary
Wikimedia Attribution Framework
Meta-Wiki Discussion
Dataset Summary
Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API.
This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/Aregay01/structured-wikipedia.structured-cpt
Structured CPT - JSON + SQL pretrain documents
SmolLM2-1.7B continued-pretraining shard of structured documents. Each document
is a <task> / <input> / <output> block whose <output> is a canonical
JSON object, terminated by the SmolLM2 end-of-text token ``.
Sources:
source
description
rows
shards
repeat
sql_bmc2
b-mc2 sql-create-context -> JSON (4 keys, stub explanation)
392,885
1
5
sql_gretelai
gretelai synthetic_text_to_sql -> JSON (4 keys)
529,255
1
5… See the full description on the dataset page: https://huggingface.co/datasets/domofon/structured-cpt.nemotron-gym-instruction-following-structured-minimax-m27-131k-tracesturkish-structured-summarization-1.5m
Turkish Structured Summarization 1.5M v2
Üç cümlelik kurgusal operasyon kayıtları ve kısa Türkçe özetleri.
Doğrulanmış boyut
Train: 1,470,000
Validation: 15,000
Test: 15,000
Toplam: 1,500,000
Ana görev sütunları: id, document, summary, domain
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-structured-summarization-1.5m.eu-ai-act-structured
EU AI Act, structured
Regulation (EU) 2024/1689 (the Artificial Intelligence Act) as tables: every article, recital, annex and definition, 677 obligations coded by actor, risk tier, application date and penalty basis, plus milestones, national competent authorities and fine tiers.
Built 2026-09-08 by SafeLegalAI (Cognesio LLP) from the official English texts served by the Publications Office of the European Union (Cellar): the consolidated text as of 27 July 2026 (CELEX… See the full description on the dataset page: https://huggingface.co/datasets/safelegalaidata/eu-ai-act-structured.task210_logic2text_structured_text_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task210_logic2text_structured_text_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task210_logic2text_structured_text_generation.structured_imagesMLR_structured_trajectory
Reasoning Trajectories with Step-Level Annotations
This dataset contains structured reasoning trajectories introduced in (ICLR 2026) Enhancing Language Model Reasoning with Structured Multi-Level Modeling.
Compared with the full-trajectory release, this dataset is a cleaned and segmented version designed for research on hierarchical reasoning, trajectory supervision, and multi-step policy training. Each example contains the original prompt and response fields together with a… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/MLR_structured_trajectory.philosophy-classics-structured
Classical Decision Frameworks — Philosophy Dataset
Structured public domain philosophical texts focused on decision-making, leadership,
and organizational ethics. All content is in the public domain.
Content
Works from classical philosophy structured for AI analysis:
Stoic decision principles (Marcus Aurelius, Epictetus, Seneca)
Political philosophy (Machiavelli, Aristotle)
Virtue ethics (Aristotle, Plato)
Sources
All works published before 1928… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/philosophy-classics-structured.phased-self-discover-mistral-structured-5-shot-bbh-evalsft-tool-calling-structured-output-v1
vericava/sft-tool-calling-structured-output-v1
Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications.
Includes contents in English as well as some Japanese.
ko-en-structured-translations
Korean–English Multistyle Parallel Corpus
한국어 사용자에게 익숙한 표현 기반의 다도메인·다문체 한–영 병렬 코퍼스
소개(Introduction)
저는 머신러닝, 인공지능 수업을 진행하는 강사입니다.Seq2Seq, Attention, Transformer 등 자연어처리(NLP) 수업을 진행하며한국 학습자에게 자연스럽고 익숙한 한–영 번역 데이터셋의 부족을 경험했습니다.
기존 공개 데이터셋은
도메인 다양성이 부족하거나
문체가 한국 사용자에게 자연스럽지 않거나
전반적으로 문장의 퀄리티가 매우 부족하여
학습한 번역 모델의 실제 성능이 기대만큼 나오지 않는 문제가 있었습니다.
이 문제를 해결하기 위해, 딥러닝 강사로서 langchain을 사용하여 직접 고품질 병렬 데이터를 자동으로 생성·정제하여 구성한 데이터셋입니다.
한국어 사용자에게 익숙한 표현을 중심으로 다양한 문체, 문장 구조를… See the full description on the dataset page: https://huggingface.co/datasets/strongminsu/ko-en-structured-translations.tt-structured-content
Dataset Summary
This dataset contains structured textual content in Markdown format extracted from Tatar-language documents, originally in EPUB and PDF formats. The documents include books and other long-form content with rich formatting. The dataset is intended to provide clean, structured, and semantically meaningful content to support natural language processing tasks, content modeling, and research in Tatar language technologies.
The extracted Markdown preserves key… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tt-structured-content.MedCase-Structured
MedCase-Structured
Dataset for Paper MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings
Structured FHIR R4 representations of clinical reasoning cases, derived from the
MedCaseReasoning dataset (Wu et al., 2025). Each case pairs a free-text
clinical presentation with a machine-readable FHIR bundle and a held-out
ground-truth diagnosis, supporting evaluation of clinical information
extraction, terminology coding… See the full description on the dataset page: https://huggingface.co/datasets/system-technologies/MedCase-Structured.Nemotron-Agents-Tools-and-Structured-Task-Execution-prompt-only
Agents, Tools and Structured Task Execution Prompt-Only
This dataset combines prompt-only datasets by capability theme for distillation experiments.
It contains 675,882 unique prompts from 808,884 raw rows;
133,002 exact canonical duplicates were removed.
Rows retain the canonical prompt-extraction columns and add source_repo_id for provenance.
Deduplication uses normalized system_prompt, prompt, tools, and schema_str, with the first
row in manifest order retained. Original… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-Agents-Tools-and-Structured-Task-Execution-prompt-only.cyber-evidence-kev-structured
Cyber Security Evidence Dataset — CISA KEV Structured CC0 Layer
This configuration is the structured CISA Known Exploited Vulnerabilities (KEV) layer of the broader Cyber Security Evidence Dataset project. It contains 1,687 deterministic records generated from the official CISA KEV database snapshot.
What is included
The records contain the official KEV database fields: CVE identifier, vendor/project, product, vulnerability name, short description, dates… See the full description on the dataset page: https://huggingface.co/datasets/frangelbarrera/cyber-evidence-kev-structured.task128_scan_structured_text_generation_command_action_short
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.openclassgen-structured-v1
OpenClassGen Structured v1
Derived from mrahman2025/OpenClassGen (Rahman et al. 2025, arXiv:2504.15564).
License: CC BY 2.0 (same as upstream). Keep repository_name and file_path when redistributing.
Underlying GitHub repos may carry additional software licenses.
gold_code is upstream human_written_code.
We add parsed fields, body-span indices, and a Variant-3 prompt/target pair (v3_prompt_text / v3_target_text).
No unit tests. Splits are repository-disjoint (train /… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/openclassgen-structured-v1.structured3d-spatiallm
Structured3D-SpatialLM Dataset
Structured3D dataset preprocessed in SpatialLM format for layout estimation with LLMs.
Overview
This dataset is derived from Structured3D 3,500 synthetic house designs created by professional designers, preprocessed and formatted specifically for SpatialLM training.
Point clouds and layouts are derived from the RoomFormer data preprocessing script.
Data Extraction
Point clouds and layouts are compressed in zip files. To… See the full description on the dataset page: https://huggingface.co/datasets/ysmao/structured3d-spatiallm.hotpotqa-structuredThis dataset is associated with the paper Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents.
Official GitHub repository: https://github.com/ielab/skim-search-agent
structured-wikipedia
Dataset Card for Wikimedia Structured Wikipedia
Quick Links
Wikimedia Enterprise
Structured Contents Documentation
Data Dictionary
Wikimedia Attribution Framework
Meta-Wiki Discussion
Dataset Summary
Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API.
This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/NewCarbon37/structured-wikipedia.full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.
