datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
structured-wikipedia
Dataset Card for Wikimedia Structured Wikipedia
Quick Links
Wikimedia Enterprise
Structured Contents Documentation
Data Dictionary
Wikimedia Attribution Framework
Meta-Wiki Discussion
Dataset Summary
Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API.
This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/structured-wikipedia.ndl-core-structured-data
NDL Core – Structured Data
Overview
NDL Core – Structured Data is a curated collection of structured UK public sector datasets, converted into Apache Parquet format for efficient analytics and machine learning workflows.
This repository is part of the broader NDL Core Corpus, which combines both textual and structured data sourced from authoritative UK government and public sector platforms.
Textual sources (e.g. GOV.UK, Hansard, legislation.gov.uk) are hosted separately… See the full description on the dataset page: https://huggingface.co/datasets/theodi/ndl-core-structured-data.brain-structureA collection of T1-weighted .nii.gz structural MRI scans in a BIDS-like arrangement,
with JSON sidecar metadata indicating train/validation/test splits.tcren_structures
isalgo/tcren_structures
TCR:peptide:MHC structure sets and benchmarks for TCRen2 (structure-based prediction of TCR
recognition). Fetch with tcren / the manuscript scripts/bootstrap_data.py.
Contents rule: structures as .gz/.tar.gz (LFS) and .txt/.md descriptions only —
no notebooks, figures, or analysis tables.
Layout
folder
task
contents
Native2026/, Canonical2026/
derivation / ergodicity
non-redundant TCR:pMHC structures (.gz)
Native2022/… See the full description on the dataset page: https://huggingface.co/datasets/isalgo/tcren_structures.sharegpt-structured-output-json
ShareGPT-Formatted Dataset for Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.viral-protein-structuresFolded/Extracted structures from PDB, AF2, ESMAtlas, and additional structures folded via AlphaFold2 on the Kempner Institute H100 GPUs.
Nemotron-RL-Instruction-Following-Structured-Outputs-v2
Dataset Description:
Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema.
Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.a3-rl-laion_nemotron-gym-instruction-following-structuredpxr-structure-pose-pool
PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands)
Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2)
structure-prediction challenge, released openly with per-pose labels so the community can
reuse the compute already spent — and, we hope, crack the problem this data makes visible.
What's here
poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand
(resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.nomad_structure
Dataset Details
Dataset Description
A subset from NOMAD dataset, which is a database of DFT computed results of materials.
This subset consists of cif structures of around 0.5 million bulk stable materials and their geometric and structural information.
All materials in this dataset are modeled using Density Functional Theory using GGA functional.
Curated by:
License: CC BY 4.0
Dataset Sources
original data source
Citation
BibTeX:… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nomad_structure.structured-file-audit-benchmark
Paper Data Release
This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them.
Contents
datasets/
Benchmark data and per-task manifests for the three paper-facing splits.
datasets/sc_flat/data
SC-Flat is derived from DaBench, augmented with a replayable perturbation
injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.neurips-spectraThe dataset from Albert's et al, downloaded from zenodo. It's on here for easier access and organisation.
MISATO_MDNemotron-RL-instruction_following-structured_outputs
Dataset Description:
The Nemotron-RL-instruction_following-structured_outputs dataset tests the ability of the model to follow output formatting instructions under schema constraints under the JSON format. Each problem consists of three components: The document, output formatting Instruction (Schema), and question. The dataset varies the difficulty of each problem by varying the location of instructions, the comprehensiveness of instructions, the complexity of the schema, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-instruction_following-structured_outputs.secondary_structure_prediction3dvlm-structured3d_subset
Structured3D Subset (3DVLM)
A small, fast-to-download slice of the Structured3D synthetic indoor dataset,
converted to a uniform posed-RGB-D format for quick model test-runs. This is a
subset: 100 scenes (randomly sampled, seed 0) from collection 00, using
the pre-rendered full (furnished) perspective views. Across the 100 scenes
there are 2,198 frames (3–49 per scene).
These are photorealistic synthetic renders with perfect dense ground-truth
depth and exact camera poses — no… See the full description on the dataset page: https://huggingface.co/datasets/helioom/3dvlm-structured3d_subset.beacon-secondary-structure
BEACON — Secondary_structure_prediction
RNA secondary-structure prediction data with nucleotide-level pair matrices.
Official data from the shared BEACON/RNABenchmark Drive folder:
https://drive.google.com/drive/folders/19ddrwI8ycvIxkgSV3gDo_VunLofYd4-6?hl=en.
This repository is the standardized Hugging Face publication of the official
task data. The data/ directory is the canonical viewer-friendly layer, and
the original file contents and source names are preserved for… See the full description on the dataset page: https://huggingface.co/datasets/jiahaozhang2003/beacon-secondary-structure.OMat24ImagePulseV2-Edit-Structure
ImagePulseV2 Dataset - Image Structure
The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It comprises multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Structure.pseudo-camera-10k-structured-json
pseudo-camera-10k, structured JSON captions
The 9,997 training images from bghira/pseudo-camera-10k, recaptioned into the structured JSON caption schema that Ideogram 4 consumes. The images are unchanged: free photographs from world class photographers, Lanczos-resized so the shorter edge is 1024px, nothing upsampled.
The original dataset carries short CogVLM prose captions. This one replaces them with one JSON object per image describing the scene at three levels: an overall… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/pseudo-camera-10k-structured-json.structure-heavy-token-quality-datasetdatasets-nanophotonic-structuresnemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces.structured-wikipedia
Dataset Card for Wikimedia Structured Wikipedia
Quick Links
Wikimedia Enterprise
Structured Contents Documentation
Data Dictionary
Wikimedia Attribution Framework
Meta-Wiki Discussion
Dataset Summary
Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API.
This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/Aregay01/structured-wikipedia.structured-cpt
Structured CPT - JSON + SQL pretrain documents
SmolLM2-1.7B continued-pretraining shard of structured documents. Each document
is a <task> / <input> / <output> block whose <output> is a canonical
JSON object, terminated by the SmolLM2 end-of-text token ``.
Sources:
source
description
rows
shards
repeat
sql_bmc2
b-mc2 sql-create-context -> JSON (4 keys, stub explanation)
392,885
1
5
sql_gretelai
gretelai synthetic_text_to_sql -> JSON (4 keys)
529,255
1
5… See the full description on the dataset page: https://huggingface.co/datasets/domofon/structured-cpt.protein_structure_NER_model_v3.1
Overview
This data was used to train model:
https://huggingface.co/PDBEurope/BiomedNLP-PubMedBERT-ProteinStructure-NER-v3.1
There are 20 different entity types in this dataset:
"bond_interaction", "chemical", "complex_assembly", "evidence", "experimental_method", "gene",
"mutant", "oligomeric_state", "protein", "protein_state", "protein_type", "ptm", "residue_name",
"residue_name_number","residue_number", "residue_range", "site", "species", "structure_element",
"taxonomy_domain"… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_structure_NER_model_v3.1.bank-statement-structure-recognition
Synthetic Bank Statement Table Structure Dataset
A synthetically generated collection of bank statement images with pixel-perfect, automatically-produced bounding box annotations for table structure recognition (TSR).
🔑 In one sentence: fake bank statements + auto-generated YOLO labels for every table cell, built so you can train table-detection models (TATR, DETR, YOLO) without manual annotation.
At a Glance
Task
Object Detection → Table… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/bank-statement-structure-recognition.protein_structure_NER_model_v2.1
Overview
This data was used to train model:
https://huggingface.co/PDBEurope/BiomedNLP-PubMedBERT-ProteinStructure-NER-v2.1
There are 20 different entity types in this dataset:
"bond_interaction", "chemical", "complex_assembly", "evidence", "experimental_method", "gene",
"mutant", "oligomeric_state", "protein", "protein_state", "protein_type", "ptm", "residue_name",
"residue_name_number","residue_number", "residue_range", "site", "species", "structure_element",
"taxonomy_domain"… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_structure_NER_model_v2.1.nemotron-gym-instruction-following-structured-minimax-m27-131k-tracesturkish-structured-summarization-1.5m
Turkish Structured Summarization 1.5M v2
Üç cümlelik kurgusal operasyon kayıtları ve kısa Türkçe özetleri.
Doğrulanmış boyut
Train: 1,470,000
Validation: 15,000
Test: 15,000
Toplam: 1,500,000
Ana görev sütunları: id, document, summary, domain
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-structured-summarization-1.5m.
