datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-RL-instruction_following-structured_outputs
Dataset Description:
The Nemotron-RL-instruction_following-structured_outputs dataset tests the ability of the model to follow output formatting instructions under schema constraints under the JSON format. Each problem consists of three components: The document, output formatting Instruction (Schema), and question. The dataset varies the difficulty of each problem by varying the location of instructions, the comprehensiveness of instructions, the complexity of the schema, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-instruction_following-structured_outputs.structure-heavy-token-quality-datasetMLR_structured_trajectory
Reasoning Trajectories with Step-Level Annotations
This dataset contains structured reasoning trajectories introduced in (ICLR 2026) Enhancing Language Model Reasoning with Structured Multi-Level Modeling.
Compared with the full-trajectory release, this dataset is a cleaned and segmented version designed for research on hierarchical reasoning, trajectory supervision, and multi-step policy training. Each example contains the original prompt and response fields together with a… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/MLR_structured_trajectory.philosophy-classics-structured
Classical Decision Frameworks — Philosophy Dataset
Structured public domain philosophical texts focused on decision-making, leadership,
and organizational ethics. All content is in the public domain.
Content
Works from classical philosophy structured for AI analysis:
Stoic decision principles (Marcus Aurelius, Epictetus, Seneca)
Political philosophy (Machiavelli, Aristotle)
Virtue ethics (Aristotle, Plato)
Sources
All works published before 1928… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/philosophy-classics-structured.sft-tool-calling-structured-output-v1
vericava/sft-tool-calling-structured-output-v1
Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications.
Includes contents in English as well as some Japanese.
protein-secondary-structure-netsurfp
NetSurfP-3.0 Secondary-Structure Splits
This dataset repo contains NetSurfP-derived protein secondary-structure labels
converted for Protein-I-JEPA probe training and evaluation.
Source page: https://services.healthtech.dtu.dk/services/NetSurfP-3.0/5-Dataset.php
Profile: hhblits
Labels are Q3 per-residue labels:
H: helix
E: beta strand
C: coil/other
.: ignored residue for loss and accuracy
Splits
Split
Rows
JSONL
TSV
train
10348
train.jsonl
tsv/train.tsv… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/protein-secondary-structure-netsurfp.structured3d-spatiallm
Structured3D-SpatialLM Dataset
Structured3D dataset preprocessed in SpatialLM format for layout estimation with LLMs.
Overview
This dataset is derived from Structured3D 3,500 synthetic house designs created by professional designers, preprocessed and formatted specifically for SpatialLM training.
Point clouds and layouts are derived from the RoomFormer data preprocessing script.
Data Extraction
Point clouds and layouts are compressed in zip files. To… See the full description on the dataset page: https://huggingface.co/datasets/ysmao/structured3d-spatiallm.full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.enwiki_structured_content
Dataset Card for enwiki_structured_content
Dataset Description
This dataset is derived from the early official Wikipedia release, downloaded from the en subset of Wikipedia Structured Contents.Articles were converted to Markdown.
hanzi-structurestructured-output-sft-100k
Structured Output SFT (100K)
100,000 ShareGPT conversations demonstrating correct generation of structured data formats: JSON, YAML, CSV, XML, Markdown tables, JSON Schema, and OpenAPI fragments. Each example pairs a natural language specification with a valid, well-formed output.
Motivation
Structured output generation is among the most commercially critical LLM capabilities. Models fail in characteristic ways:
Invalid JSON: unclosed brackets, trailing commas… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/structured-output-sft-100k.structured-hard-sft-4k
Hard Synthetic Dataset for Structured Data Tasks (v1)
This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types).
Dataset Summary
The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.MatPESNemotron-RL-Instruction-Following-Structured-Outputs-v2-preferencestructured_data_with_cot_dataset_512_v4
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。
データセットの概要
messages: OpenAIチャット形式 (system, user, assistant)
metadata: format, complexity, schema, estimated_tokens
サポートされるデータ形式
JSON, XML, YAML, TOML, CSV
生成方法
Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。
PCQM4Mv2chemical-structure⚗️ Dataset Summary
A large-scale chemical structure database with 1 million compounds, providing standardized cheminformatics identifiers and taxonomic classification. Designed for structure-based drug discovery, similarity search, and chemical space analysis.
🚀 Key Features
Standard Identifiers: Every record includes InChI, InChIKey, and molecular formula; isomeric SMILES present for 99.8% of records.
Multiple Name Forms: JChem-generated names (97%), IUPAC names (22%), and compound names… See the full description on the dataset page: https://huggingface.co/datasets/PatSnap/chemical-structure.structured-5k-mix-sft
5k Mixed Hard-Structured SFT Dataset (v1)
This dataset contains 5,000 synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
It aggregates 13 distinct conversion tasks with a specific focus on format diversity and structural complexity.
Dataset Summary
The dataset is distributed across five major formats with the following allocation:
Target Format
Count
Share
Task Types
YAML
1,500
30%… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-5k-mix-sft.GEOM1secondary_structureAgent-IPI-Structured-Interaction-Datasets
Dataset Card for Indirect Prompt Injection in Agent Structured Interaction Datasets
Dataset Summary
This dataset contains 470,000 QA pairs designed to study indirect prompt injection in agent-structured interactions. It is split into a training set (80%) and a test set (20%). The dataset is evenly divided into 50% clean-clean QA pairs (no prompt injection) and 50% clean-injected QA pairs (containing prompt injection). The task is to detect and remove prompt injection… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets.structured-hard-sft-4k
Hard Synthetic Dataset for Structured Data Tasks (v1)
This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types).
Dataset Summary
The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/zyz123code/structured-hard-sft-4k.structured_data_with_cot_dataset_512_v5
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。
データセットの概要
messages: OpenAIチャット形式 (system, user, assistant)
metadata: format, complexity, schema, estimated_tokens
サポートされるデータ形式
JSON, XML, YAML, TOML, CSV
生成方法
Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。
v5アップデート:ランダムなスキーマ構造の生成と、最小化(minified)/ソート(sorted)の制約を追加。
1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample
Description
한국어 시험 문제 구조화 분석·가공 데이터로, 약 150만 개의 시험 문제를 포함하고 있습니다. 문제 유형, 문제, 정답, 해설 등의 정보를 포함하며, 과목은 [초등학교] 국어, 수학, 영어, 사회, 과학; [중학교] 국어, 영어, 수학, 과학, 사회; [고등학교] 국어, 영어, 수학, 물리, 화학, 생물, 역사, 지리로 구성되어 있습니다. 문제 유형에는 객관식, 빈칸 채우기, 참·거짓 문제, 단답형 문제 등이 포함됩니다. 본 데이터셋은 대규모 교과 지식 강화 및 학습 데이터 구축 등의 작업에 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/1634?source=Hf.kr
Specifications
Data content
한국어 K12 시험 문제
Amount
약 150만 개의… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-001
ToricBLM dataset state: toricblm-structure-priority-balanced-3day-20260709T185034Z epoch 001
This dataset repo records the exact local training-data state visible to the dynamic epoch launcher.
It intentionally stores manifests and audit records rather than duplicating large Parquet shards.
Special checkpoint: toricblm-structure-priority-balanced-3day-20260709T185034Z_epoch_001_special_structure_current_step_002000.pt
Checkpoint repo: AmelieSchreiber/ToricGT_160M_FoT
Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-001.structured_data_with_cot_dataset_512_v3
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。
データセットの概要
messages: OpenAIチャット形式 (system, user, assistant)
metadata: format, complexity, schema, estimated_tokens
サポートされるデータ形式
JSON, XML, YAML, TOML, CSV
生成方法
Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。
repro-on-structured-state-space-duality-traces
Agent traces
Agent sessions published from a Trackio Logbook.
toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-002
ToricBLM dataset state: toricblm-structure-priority-balanced-3day-20260709T185034Z epoch 002
This dataset repo records the exact local training-data state visible to the dynamic epoch launcher.
It intentionally stores manifests and audit records rather than duplicating large Parquet shards.
Special checkpoint: toricblm-structure-priority-balanced-3day-20260709T185034Z_epoch_002_special_structure_delta_step_002750.pt
Checkpoint repo: AmelieSchreiber/ToricGT_160M_FoT
Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-002.Structured3D
Structured3D-SpatialLM Dataset
Structured3D dataset preprocessed in SpatialLM format for layout estimation with LLMs.
Overview
This dataset is derived from Structured3D 3,500 synthetic house designs created by professional designers, preprocessed and formatted specifically for SpatialLM training.
Point clouds and layouts are derived from the RoomFormer data preprocessing script.
Data Extraction
Point clouds and layouts are compressed in zip files. To… See the full description on the dataset page: https://huggingface.co/datasets/Gen3DF/Structured3D.stage3-synthetic-structured-retrieval
Stage 3 Synthetic Structured-Retrieval Agents
Native search-tool trajectories generated by
Qwen/Qwen3-235B-A22B-Instruct-2507 for LCLM Stage-3 agent post-training.
The default config contains only traces that passed programmatic evidence and
answer verification.
Harvest
Accepted traces: 82
Native search calls: 179
Compressed tool-observation traces: 41
Uncompressed traces: 41
Task-ID overlap between pilot and collection batch: 0
Family counts are 20 latest-state… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-synthetic-structured-retrieval.
