CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes1.2k downloads2y agoHugging Face02nvidia /Nemotron-RL-Instruction-Following-Structured-Outputs-v2 Dataset Description: Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema. Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.texttext-generation10K<n<100K7 likes1.2k downloads4mo agoHugging Face03domofon /structured-cpt Structured CPT - JSON + SQL pretrain documents SmolLM2-1.7B continued-pretraining shard of structured documents. Each document is a <task> / <input> / <output> block whose <output> is a canonical JSON object, terminated by the SmolLM2 end-of-text token ``. Sources: source description rows shards repeat sql_bmc2 b-mc2 sql-create-context -> JSON (4 keys, stub explanation) 392,885 1 5 sql_gretelai gretelai synthetic_text_to_sql -> JSON (4 keys) 529,255 1 5… See the full description on the dataset page: https://huggingface.co/datasets/domofon/structured-cpt.texttext-generation1M<n<10M0 likes298 downloads17d agoHugging Face04Lots-of-LoRAs /task210_logic2text_structured_text_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task210_logic2text_structured_text_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task210_logic2text_structured_text_generation.texttext-generation1K<n<10K0 likes209 downloads2y agoHugging Face05mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes178 downloads7mo agoHugging Face06vericava /sft-tool-calling-structured-output-v1 vericava/sft-tool-calling-structured-output-v1 Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications. Includes contents in English as well as some Japanese. texttext-classification100K<n<1M2 likes167 downloads8mo agoHugging Face07dhruveshpatel /openclassgen-structured-v1 OpenClassGen Structured v1 Derived from mrahman2025/OpenClassGen (Rahman et al. 2025, arXiv:2504.15564). License: CC BY 2.0 (same as upstream). Keep repository_name and file_path when redistributing. Underlying GitHub repos may carry additional software licenses. gold_code is upstream human_written_code. We add parsed fields, body-span indices, and a Variant-3 prompt/target pair (v3_prompt_text / v3_target_text). No unit tests. Splits are repository-disjoint (train /… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/openclassgen-structured-v1.tabulartext-generation100K<n<1M0 likes139 downloads2mo agoHugging Face08Lots-of-LoRAs /task128_scan_structured_text_generation_command_action_short Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.texttext-generation1K<n<10K0 likes129 downloads2y agoHugging Face09stindardlogic /structured-output-sft-100k Structured Output SFT (100K) 100,000 ShareGPT conversations demonstrating correct generation of structured data formats: JSON, YAML, CSV, XML, Markdown tables, JSON Schema, and OpenAPI fragments. Each example pairs a natural language specification with a valid, well-formed output. Motivation Structured output generation is among the most commercially critical LLM capabilities. Models fail in characteristic ways: Invalid JSON: unclosed brackets, trailing commas… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/structured-output-sft-100k.texttext-generation100K<n<1M1 likes109 downloads2mo agoHugging Face10Lots-of-LoRAs /task1566_propara_structured_text_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1566_propara_structured_text_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1566_propara_structured_text_generation.texttext-generationn<1K0 likes90 downloads2y agoHugging Face11Blaze7451 /enwiki_structured_content Dataset Card for enwiki_structured_content Dataset Description This dataset is derived from the early official Wikipedia release, downloaded from the en subset of Wikipedia Structured Contents.Articles were converted to Markdown. texttext-generation1M<n<10M1 likes84 downloads1y agoHugging Face12Lots-of-LoRAs /task130_scan_structured_text_generation_command_action_long Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task130_scan_structured_text_generation_command_action_long Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task130_scan_structured_text_generation_command_action_long.texttext-generation1K<n<10K0 likes82 downloads2y agoHugging Face13daichira /structured-hard-sft-4k Hard Synthetic Dataset for Structured Data Tasks (v1) This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types). Dataset Summary The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.texttext-generation1K<n<10K1 likes79 downloads8mo agoHugging Face14Haeryz /putusan-structured-extraction Putusan structured-extraction dataset Built 2026-07-08T23:02:28+00:00 by notebooks/build_dataset.py (seed 3407). Indonesian court-decision (putusan) extractive-structuring dataset over three corpora (Anak, Asusila, TPPO). Each row is one model extraction of one source document into 31 canonical sections of verbatim spans. Empty sections were completed from sibling model extractions of the same document where available (cross_model_fill_json records per-section donor provenance).… See the full description on the dataset page: https://huggingface.co/datasets/Haeryz/putusan-structured-extraction.tabulartext-generation1K<n<10K0 likes71 downloads2mo agoHugging Face15emgena /omnimcp_browser_dom_structured_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.texttext-generationn<1K0 likes71 downloads6d agoHugging Face16daichira /structured-5k-mix-sft 5k Mixed Hard-Structured SFT Dataset (v1) This dataset contains 5,000 synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. It aggregates 13 distinct conversion tasks with a specific focus on format diversity and structural complexity. Dataset Summary The dataset is distributed across five major formats with the following allocation: Target Format Count Share Task Types YAML 1,500 30%… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-5k-mix-sft.texttext-generation1K<n<10K0 likes52 downloads8mo agoHugging Face17TachyHealth /structured_medicalThe dataset was presented in the paper Gazal-R1: Achieving State-of-the-Art Medical Reasoning with Parameter-Efficient Two-Stage Training. texttext-generation100K<n<1M2 likes50 downloads1y agoHugging Face18zyz123code /structured-hard-sft-4k Hard Synthetic Dataset for Structured Data Tasks (v1) This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types). Dataset Summary The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/zyz123code/structured-hard-sft-4k.texttext-generation1K<n<10K1 likes50 downloads2mo agoHugging Face19dotwee /structured-stern-neon-articles Structured Stern NEON Community Articles This repository contains approximately 20k user written texts, articles, and poetry pulled from archives of the Stern NEON website. Stern NEON was a community platform where users could write and publish their own articles. Many of the articles are personal stories, poems, or opinion pieces. The articles are structured in a way that they can be used for further analysis. Dataset Details Uses This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.tabulartext-classification10K<n<100K0 likes37 downloads8mo agoHugging Face20chuckreynolds /wikimedia-enterprise-structured-contents-enwiki enwiki_namespace_0 Structured Contents snapshot of enwiki_namespace_0 from the Wikimedia Enterprise API, converted to Parquet. Source Upstream: Wikimedia Enterprise Structured Contents API Snapshot identifier: enwiki_namespace_0 Format at source: .tar.gz containing sharded .ndjson Shards in this release: 3 Processing Downloaded the snapshot tarball from the Wikimedia Enterprise API. Streamed each .ndjson shard through a normalization pass: JSON-encoded… See the full description on the dataset page: https://huggingface.co/datasets/chuckreynolds/wikimedia-enterprise-structured-contents-enwiki.texttext-generation100K<n<1M0 likes37 downloads5mo agoHugging Face21philipp-zettl /german-structured-output German Structured Output Dataset 🇩🇪 GDPR & EU AI Act compliant German dataset for training structured output capabilities in LLMs. Overview This dataset contains 4,521 examples across 7 task types for training language models to produce structured outputs (JSON, function calls, schema-following generation) from German text. It is the first dedicated German structured output dataset, filling a critical gap in the German NLP ecosystem. Key Features 🇩🇪… See the full description on the dataset page: https://huggingface.co/datasets/philipp-zettl/german-structured-output.texttext-generation1K<n<10K0 likes37 downloads5mo agoHugging Face22morizon /TCGA_Reports_ja_structured_qwen38_27b TCGA Reports Japanese Structured Dataset with Qwen3.8-27B The Cancer Genome Atlas(TCGA)由来の英語病理報告書を日本語へ翻訳し、その日本語病理報告書から主要な病理情報を9項目へ構造化したデータセットです。 既存の morizon/TCGA_Reports_ja_structured と同じ100症例を使用し、生成モデルを Qwen/Qwen3.8-27B に変更して再生成しています。 日本語訳には morizon/TCGA_Reports_ja_qwen38_27b と同じ生成結果を使用しています。 元データ 本データセットでは、The Cancer Genome Atlas(TCGA)の病理報告書をもとに作成されたTCGA-Reportsを使用しています。 TCGAは、複数のがん種についてゲノム情報や臨床情報などを収集した大規模ながん研究プロジェクトです。… See the full description on the dataset page: https://huggingface.co/datasets/morizon/TCGA_Reports_ja_structured_qwen38_27b.texttext-generationn<1K0 likes37 downloads1mo agoHugging Face23leonli66 /stage3-synthetic-structured-retrieval Stage 3 Synthetic Structured-Retrieval Agents Native search-tool trajectories generated by Qwen/Qwen3-235B-A22B-Instruct-2507 for LCLM Stage-3 agent post-training. The default config contains only traces that passed programmatic evidence and answer verification. Harvest Accepted traces: 82 Native search calls: 179 Compressed tool-observation traces: 41 Uncompressed traces: 41 Task-ID overlap between pilot and collection batch: 0 Family counts are 20 latest-state… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-synthetic-structured-retrieval.tabulartext-generationn<1K0 likes34 downloads1mo agoHugging Face24Arsh9210 /Nemotron-RL-Instruction-Following-Structured-Outputs-v2 Dataset Description: Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema. Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.texttext-generation10K<n<100K0 likes27 downloads2mo agoHugging Face25stindardlogic /json-structured-output-dpo-3k JSON Structured Output DPO Pairs (3K) DPO preference pairs for training LLMs to produce valid, schema-compliant JSON output. Motivation Structured output (JSON mode) is critical for production AI applications — parsers fail, pipelines break, and downstream processing errors when models output malformed JSON, use wrong field names, or wrap responses in markdown. This dataset trains strict schema adherence. Dataset Description 3,000 preference pairs… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/json-structured-output-dpo-3k.texttext-generation1K<n<10K0 likes26 downloads2mo agoHugging Face26dhruveshpatel /classeval-structured-v1 ClassEval Structured v1 Derived from FudanSELab/ClassEval (Du et al. 2023, arXiv:2308.01861). License: CC BY-NC 4.0 (upstream data license). Non-commercial use only. One row per (task_id, variant) with variant in {1,2,3} (100 tasks × 3 = 300 rows; Hub split test). solution_code, test, and methods_info_json come from upstream. We add rendered prompts/targets and stratification fields. Missing bodies use .... Variants: Signatures and docstrings kept; every method body is ....… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/classeval-structured-v1.tabulartext-generationn<1K0 likes25 downloads2mo agoHugging Face27obadabaq /structured-uae-laws Dataset Card for structured-uae-laws This dataset is a collection of question & answers about the laws and regulations in the United Arab Emirates. It covers different areas of law like: economy and business family and community finance and banking industry and technical standardisation justice and juiciary, labour residency and leberal professions security and safety tax Dataset Sources Repository Base Dataset United Arab Emirates Legislations… See the full description on the dataset page: https://huggingface.co/datasets/obadabaq/structured-uae-laws.textquestion-answering1K<n<10K1 likes23 downloads2y agoHugging Face28takami2022 /structured_data_merged_v2v5_0222 Dataset Card for structured_data_merged_v2v5_0222 Dataset Details Dataset Description structured_data_merged_v2v5_0222 is a dataset for Supervised Fine-Tuning (SFT) focused on structured data format conversion tasks — specifically, interconversion among JSON, XML, YAML, TOML, and CSV. It was created by deduplicating and merging the following two existing datasets: u-10bei/structured_data_with_cot_dataset_512_v2 (train split only)… See the full description on the dataset page: https://huggingface.co/datasets/takami2022/structured_data_merged_v2v5_0222.texttext-generation1K<n<10K0 likes23 downloads7mo agoHugging Face29Lots-of-LoRAs /task126_scan_structured_text_generation_command_action_all Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task126_scan_structured_text_generation_command_action_all Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task126_scan_structured_text_generation_command_action_all.texttext-generation1K<n<10K0 likes21 downloads2y agoHugging Face30mssfj /openmathinstruct-2_structured-1000 OpenMathInstruct-2 Structured (CoT) OpenMathInstruct-2 の solution を gpt-oss-120b で再生成し、<analyze> <plan> <verify> を含む構造化 CoT を付与した SFT 用データセットです。 目的: 数学タスク向けの長手順 CoT を安定して生成するための教師あり微調整。 ライセンス: CC-BY-4.0(元データのライセンスに従います)。 データ概要 言語: 英語 レコード数: 1,446 形式: JSONL / Parquet(Hugging Face datasets 形式) カラム フィールド 型 説明 question string OpenMathInstruct-2 の問題文 answer string <think> ブロック内に <analyze>, <plan>, <verify>, <reason> を埋め込んだ構造化 CoT と最終回答 category… See the full description on the dataset page: https://huggingface.co/datasets/mssfj/openmathinstruct-2_structured-1000.texttext-generation1K<n<10K0 likes21 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.