datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
structured-wikipedia
Dataset Card for Wikimedia Structured Wikipedia
Quick Links
Wikimedia Enterprise
Structured Contents Documentation
Data Dictionary
Wikimedia Attribution Framework
Meta-Wiki Discussion
Dataset Summary
Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API.
This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/structured-wikipedia.Nemotron-RL-Instruction-Following-Structured-Outputs-v2
Dataset Description:
Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema.
Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.sharegpt-structured-output-json
ShareGPT-Formatted Dataset for Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.nomad_structure
Dataset Details
Dataset Description
A subset from NOMAD dataset, which is a database of DFT computed results of materials.
This subset consists of cif structures of around 0.5 million bulk stable materials and their geometric and structural information.
All materials in this dataset are modeled using Density Functional Theory using GGA functional.
Curated by:
License: CC BY 4.0
Dataset Sources
original data source
Citation
BibTeX:… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nomad_structure.neurips-spectraThe dataset from Albert's et al, downloaded from zenodo. It's on here for easier access and organisation.
a3-rl-laion_nemotron-gym-instruction-following-structuredbeacon-secondary-structure
BEACON — Secondary_structure_prediction
RNA secondary-structure prediction data with nucleotide-level pair matrices.
Official data from the shared BEACON/RNABenchmark Drive folder:
https://drive.google.com/drive/folders/19ddrwI8ycvIxkgSV3gDo_VunLofYd4-6?hl=en.
This repository is the standardized Hugging Face publication of the official
task data. The data/ directory is the canonical viewer-friendly layer, and
the original file contents and source names are preserved for… See the full description on the dataset page: https://huggingface.co/datasets/jiahaozhang2003/beacon-secondary-structure.nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces.structured-cpt
Structured CPT - JSON + SQL pretrain documents
SmolLM2-1.7B continued-pretraining shard of structured documents. Each document
is a <task> / <input> / <output> block whose <output> is a canonical
JSON object, terminated by the SmolLM2 end-of-text token ``.
Sources:
source
description
rows
shards
repeat
sql_bmc2
b-mc2 sql-create-context -> JSON (4 keys, stub explanation)
392,885
1
5
sql_gretelai
gretelai synthetic_text_to_sql -> JSON (4 keys)
529,255
1
5… See the full description on the dataset page: https://huggingface.co/datasets/domofon/structured-cpt.bank-statement-structure-recognition
Synthetic Bank Statement Table Structure Dataset
A synthetically generated collection of bank statement images with pixel-perfect, automatically-produced bounding box annotations for table structure recognition (TSR).
🔑 In one sentence: fake bank statements + auto-generated YOLO labels for every table cell, built so you can train table-detection models (TATR, DETR, YOLO) without manual annotation.
At a Glance
Task
Object Detection → Table… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/bank-statement-structure-recognition.turkish-structured-summarization-1.5m
Turkish Structured Summarization 1.5M v2
Üç cümlelik kurgusal operasyon kayıtları ve kısa Türkçe özetleri.
Doğrulanmış boyut
Train: 1,470,000
Validation: 15,000
Test: 15,000
Toplam: 1,500,000
Ana görev sütunları: id, document, summary, domain
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-structured-summarization-1.5m.eu-ai-act-structured
EU AI Act, structured
Regulation (EU) 2024/1689 (the Artificial Intelligence Act) as tables: every article, recital, annex and definition, 677 obligations coded by actor, risk tier, application date and penalty basis, plus milestones, national competent authorities and fine tiers.
Built 2026-09-08 by SafeLegalAI (Cognesio LLP) from the official English texts served by the Publications Office of the European Union (Cellar): the consolidated text as of 27 July 2026 (CELEX… See the full description on the dataset page: https://huggingface.co/datasets/safelegalaidata/eu-ai-act-structured.structured-wikipedia
Dataset Card for Wikimedia Structured Wikipedia
Quick Links
Wikimedia Enterprise
Structured Contents Documentation
Data Dictionary
Wikimedia Attribution Framework
Meta-Wiki Discussion
Dataset Summary
Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API.
This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/Aregay01/structured-wikipedia.nemotron-gym-instruction-following-structured-minimax-m27-131k-tracesuniref50-sorted-structure-tokentask210_logic2text_structured_text_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task210_logic2text_structured_text_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task210_logic2text_structured_text_generation.structured_imagesphased-self-discover-mistral-structured-5-shot-bbh-evalopenclassgen-structured-v1
OpenClassGen Structured v1
Derived from mrahman2025/OpenClassGen (Rahman et al. 2025, arXiv:2504.15564).
License: CC BY 2.0 (same as upstream). Keep repository_name and file_path when redistributing.
Underlying GitHub repos may carry additional software licenses.
gold_code is upstream human_written_code.
We add parsed fields, body-span indices, and a Variant-3 prompt/target pair (v3_prompt_text / v3_target_text).
No unit tests. Splits are repository-disjoint (train /… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/openclassgen-structured-v1.structure_wildfire_damage_classification
Dataset Card for Structures Damaged by Wildfire
Homepage: Image Dataset of Structures Damaged by Wildfire in California 2020-2022
Dataset Summary
The dataset contains over 18,000 images of homes damaged by wildfire between 2020 and 2022 in California, USA, captured by the California Department of Forestry and Fire Protection (Cal Fire) during the damage assessment process. The dataset spans across more than 18 wildfire events, including the 2020 August Complex Fire, the… See the full description on the dataset page: https://huggingface.co/datasets/kevincluo/structure_wildfire_damage_classification.cyber-evidence-kev-structured
Cyber Security Evidence Dataset — CISA KEV Structured CC0 Layer
This configuration is the structured CISA Known Exploited Vulnerabilities (KEV) layer of the broader Cyber Security Evidence Dataset project. It contains 1,687 deterministic records generated from the official CISA KEV database snapshot.
What is included
The records contain the official KEV database fields: CVE identifier, vendor/project, product, vulnerability name, short description, dates… See the full description on the dataset page: https://huggingface.co/datasets/frangelbarrera/cyber-evidence-kev-structured.PDB-Monomeric-Structure-ESMFold2
PDB-Monomeric-Structure-ESMFold2
Monomeric, protein-only PDB structure dataset for minimum ESMFold2-style
training. Each row is one eligible single-chain biological assembly with a
canonical amino-acid sequence input and all-atom protein labels in atom37.
Labels
atom37_positions: residue x 37 x 3 coordinates, with zeros for missing atoms.
atom37_mask: residue x 37 resolved-atom mask.
aatype, residue_index, auth_seq_id, insertion_code, residue_name, ca_mask.… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/PDB-Monomeric-Structure-ESMFold2.tt-structured-content
Dataset Summary
This dataset contains structured textual content in Markdown format extracted from Tatar-language documents, originally in EPUB and PDF formats. The documents include books and other long-form content with rich formatting. The dataset is intended to provide clean, structured, and semantically meaningful content to support natural language processing tasks, content modeling, and research in Tatar language technologies.
The extracted Markdown preserves key… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tt-structured-content.task128_scan_structured_text_generation_command_action_short
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.structured-wikipedia
Dataset Card for Wikimedia Structured Wikipedia
Quick Links
Wikimedia Enterprise
Structured Contents Documentation
Data Dictionary
Wikimedia Attribution Framework
Meta-Wiki Discussion
Dataset Summary
Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API.
This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/NewCarbon37/structured-wikipedia.Alpaca_StructureToucan-1.5M-structured-Qwenstructured-vitalsbank-statement-structure-recognition
Synthetic Bank Statement Table Structure Dataset
A synthetically generated collection of bank statement images with pixel-perfect, automatically-produced bounding box annotations for table structure recognition (TSR).
🔑 In one sentence: fake bank statements + auto-generated YOLO labels for every table cell, built so you can train table-detection models (TATR, DETR, YOLO) without manual annotation.
At a Glance
Task
Object Detection → Table… See the full description on the dataset page: https://huggingface.co/datasets/sajid1235/bank-statement-structure-recognition.TCGA_Reports_ja_structured-filtered
🧠 TCGA 日本語翻訳・構造化データセット
このデータセットは、The Cancer Genome Atlas (TCGA) により公開された英語の病理報告書をもとに、大規模言語モデル(LLM)を用いて 日本語翻訳 および 情報抽出による構造化 を行ったものです。
📘 概要
原データ:
Mendeley Data — TCGA Pathology Reports (Version 1)https://data.mendeley.com/datasets/hyg5xkznpx/1
本データセットは上記を基にし、以下の2種類の加工を行っています。
カラム名
内容
生成方法
question_en
英語の病理報告書原文
TCGA オリジナル
question_ja
英語報告書の日本語訳
LLMによる翻訳
answer_ja
構造化データ(JSON形式)
LLMによる情報抽出
🧾 データ構成… See the full description on the dataset page: https://huggingface.co/datasets/morizon/TCGA_Reports_ja_structured-filtered.
