CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-RL-Instruction-Following-Structured-Outputs-v2 Dataset Description: Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema. Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.texttext-generation10K<n<100K7 likes1.1k downloads4mo agoHugging Face02Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes800 downloads2y agoHugging Face03domofon /structured-cpt Structured CPT - JSON + SQL pretrain documents SmolLM2-1.7B continued-pretraining shard of structured documents. Each document is a <task> / <input> / <output> block whose <output> is a canonical JSON object, terminated by the SmolLM2 end-of-text token ``. Sources: source description rows shards repeat sql_bmc2 b-mc2 sql-create-context -> JSON (4 keys, stub explanation) 392,885 1 5 sql_gretelai gretelai synthetic_text_to_sql -> JSON (4 keys) 529,255 1 5… See the full description on the dataset page: https://huggingface.co/datasets/domofon/structured-cpt.texttext-generation1M<n<10M0 likes373 downloads20d agoHugging Face04Lots-of-LoRAs /task210_logic2text_structured_text_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task210_logic2text_structured_text_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task210_logic2text_structured_text_generation.texttext-generation1K<n<10K0 likes189 downloads2y agoHugging Face05dhruveshpatel /openclassgen-structured-v1 OpenClassGen Structured v1 Derived from mrahman2025/OpenClassGen (Rahman et al. 2025, arXiv:2504.15564). License: CC BY 2.0 (same as upstream). Keep repository_name and file_path when redistributing. Underlying GitHub repos may carry additional software licenses. gold_code is upstream human_written_code. We add parsed fields, body-span indices, and a Variant-3 prompt/target pair (v3_prompt_text / v3_target_text). No unit tests. Splits are repository-disjoint (train /… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/openclassgen-structured-v1.tabulartext-generation100K<n<1M0 likes144 downloads2mo agoHugging Face06Lots-of-LoRAs /task128_scan_structured_text_generation_command_action_short Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.texttext-generation1K<n<10K0 likes108 downloads2y agoHugging Face07Lots-of-LoRAs /task1566_propara_structured_text_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1566_propara_structured_text_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1566_propara_structured_text_generation.texttext-generationn<1K0 likes77 downloads2y agoHugging Face08emgena /omnimcp_browser_dom_structured_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.texttext-generationn<1K0 likes76 downloads9d agoHugging Face09Lots-of-LoRAs /task130_scan_structured_text_generation_command_action_long Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task130_scan_structured_text_generation_command_action_long Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task130_scan_structured_text_generation_command_action_long.texttext-generation1K<n<10K0 likes72 downloads2y agoHugging Face10chuckreynolds /wikimedia-enterprise-structured-contents-enwiki enwiki_namespace_0 Structured Contents snapshot of enwiki_namespace_0 from the Wikimedia Enterprise API, converted to Parquet. Source Upstream: Wikimedia Enterprise Structured Contents API Snapshot identifier: enwiki_namespace_0 Format at source: .tar.gz containing sharded .ndjson Shards in this release: 3 Processing Downloaded the snapshot tarball from the Wikimedia Enterprise API. Streamed each .ndjson shard through a normalization pass: JSON-encoded… See the full description on the dataset page: https://huggingface.co/datasets/chuckreynolds/wikimedia-enterprise-structured-contents-enwiki.texttext-generation100K<n<1M0 likes38 downloads5mo agoHugging Face11philipp-zettl /german-structured-output German Structured Output Dataset 🇩🇪 GDPR & EU AI Act compliant German dataset for training structured output capabilities in LLMs. Overview This dataset contains 4,521 examples across 7 task types for training language models to produce structured outputs (JSON, function calls, schema-following generation) from German text. It is the first dedicated German structured output dataset, filling a critical gap in the German NLP ecosystem. Key Features 🇩🇪… See the full description on the dataset page: https://huggingface.co/datasets/philipp-zettl/german-structured-output.texttext-generation1K<n<10K0 likes35 downloads5mo agoHugging Face12Arsh9210 /Nemotron-RL-Instruction-Following-Structured-Outputs-v2 Dataset Description: Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema. Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.texttext-generation10K<n<100K0 likes27 downloads2mo agoHugging Face13morizon /TCGA_Reports_ja_structured_qwen38_27b TCGA Reports Japanese Structured Dataset with Qwen3.8-27B The Cancer Genome Atlas(TCGA)由来の英語病理報告書を日本語へ翻訳し、その日本語病理報告書から主要な病理情報を9項目へ構造化したデータセットです。 既存の morizon/TCGA_Reports_ja_structured と同じ100症例を使用し、生成モデルを Qwen/Qwen3.8-27B に変更して再生成しています。 日本語訳には morizon/TCGA_Reports_ja_qwen38_27b と同じ生成結果を使用しています。 元データ 本データセットでは、The Cancer Genome Atlas(TCGA)の病理報告書をもとに作成されたTCGA-Reportsを使用しています。 TCGAは、複数のがん種についてゲノム情報や臨床情報などを収集した大規模ながん研究プロジェクトです。… See the full description on the dataset page: https://huggingface.co/datasets/morizon/TCGA_Reports_ja_structured_qwen38_27b.texttext-generationn<1K0 likes27 downloads1mo agoHugging Face14dhruveshpatel /classeval-structured-v1 ClassEval Structured v1 Derived from FudanSELab/ClassEval (Du et al. 2023, arXiv:2308.01861). License: CC BY-NC 4.0 (upstream data license). Non-commercial use only. One row per (task_id, variant) with variant in {1,2,3} (100 tasks × 3 = 300 rows; Hub split test). solution_code, test, and methods_info_json come from upstream. We add rendered prompts/targets and stratification fields. Missing bodies use .... Variants: Signatures and docstrings kept; every method body is ....… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/classeval-structured-v1.tabulartext-generationn<1K0 likes24 downloads2mo agoHugging Face15takami2022 /structured_data_merged_v2v5_0222 Dataset Card for structured_data_merged_v2v5_0222 Dataset Details Dataset Description structured_data_merged_v2v5_0222 is a dataset for Supervised Fine-Tuning (SFT) focused on structured data format conversion tasks — specifically, interconversion among JSON, XML, YAML, TOML, and CSV. It was created by deduplicating and merging the following two existing datasets: u-10bei/structured_data_with_cot_dataset_512_v2 (train split only)… See the full description on the dataset page: https://huggingface.co/datasets/takami2022/structured_data_merged_v2v5_0222.texttext-generation1K<n<10K0 likes22 downloads7mo agoHugging Face16mssfj /openmathinstruct-2_structured-1000 OpenMathInstruct-2 Structured (CoT) OpenMathInstruct-2 の solution を gpt-oss-120b で再生成し、<analyze> <plan> <verify> を含む構造化 CoT を付与した SFT 用データセットです。 目的: 数学タスク向けの長手順 CoT を安定して生成するための教師あり微調整。 ライセンス: CC-BY-4.0(元データのライセンスに従います)。 データ概要 言語: 英語 レコード数: 1,446 形式: JSONL / Parquet(Hugging Face datasets 形式) カラム フィールド 型 説明 question string OpenMathInstruct-2 の問題文 answer string <think> ブロック内に <analyze>, <plan>, <verify>, <reason> を埋め込んだ構造化 CoT と最終回答 category… See the full description on the dataset page: https://huggingface.co/datasets/mssfj/openmathinstruct-2_structured-1000.texttext-generation1K<n<10K0 likes21 downloads9mo agoHugging Face17Lots-of-LoRAs /task126_scan_structured_text_generation_command_action_all Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task126_scan_structured_text_generation_command_action_all Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task126_scan_structured_text_generation_command_action_all.texttext-generation1K<n<10K0 likes16 downloads2y agoHugging Face18kurota0612 /combined-structured-dataset Combined Structured Output Dataset このデータセットは、構造化出力生成タスクのための統合データセットです。 Dataset Details Total Samples: 3,425 Source Datasets: u-10bei/structured_data_with_cot_dataset_v2 (2,500 samples) daichira/structured-3k-mix-sft (3,000 samples) Preprocessing: Format normalization to 3-turn (system/user/assistant) Deduplication by user content Quality filtering Format Distribution JSON: 685 (20.0%) YAML: 485 (14.2%) TOML: 685 (20.0%) XML: 885 (25.8%) CSV: 685… See the full description on the dataset page: https://huggingface.co/datasets/kurota0612/combined-structured-dataset.texttext-generation1K<n<10K0 likes13 downloads7mo agoHugging Face19VoeTheDon /testing-wiki-structured cywiki_namespace_0 Structured Contents snapshot of cywiki_namespace_0 from the Wikimedia Enterprise API, repackaged as Parquet with a pinned schema. The upstream Wikimedia Foundation dataset (wikimedia/structured-wikipedia) ships NDJSON which has known issues loading via datasets.load_dataset() — see discussions #5, #15, #16. This dataset is the same upstream content, normalised so load_dataset(...)works without specifying a Features override. Source Upstream: Wikimedia… See the full description on the dataset page: https://huggingface.co/datasets/VoeTheDon/testing-wiki-structured.texttext-generation10K<n<100K0 likes10 downloads5mo agoHugging Face20haining /structured_poem_interpretation_corpus_stalegatedMasking policy: For rows with source == "poetry_foundation", the poem and interpretation fields are set to null to respect content licensing. Public-domain entries (source == "public_domain_poetry") include full text. All categorical annotations (emotions, primary_emotion, sentiment, themes, themes_50) and metadata remain available. texttext-classification10K<n<100K0 likes9 downloads10mo agoHugging Face21guhhhgu /harmonicbench-planir-main-structured HARMONICBench PlanIR Structured Main Dataset This repository is a Hugging Face dataset-friendly structured export derived from the local outputs/fixed/main directory in the HARMONICBench unified package. Included tables plans/train.parquet: primary aggregate table converted from plan_runs_all.jsonl. domains/*.parquet: per-domain plan runs, selected samples, and domain5 image descriptions. summaries/*: key JSON/CSV/JSONL summary artifacts. artifacts/roundtrip_recovery*/*:… See the full description on the dataset page: https://huggingface.co/datasets/guhhhgu/harmonicbench-planir-main-structured.tabulartext-generationn<1K0 likes9 downloads5mo agoHugging Face22open-athena /nemotron-gym-structured-outputs-v3 laion/nemotron-gym-structured-outputs-v3 Harbor task-binary dataset (53,870 tasks) converted from nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2 (part of the nvidia/Nemotron-Post-Training-v3 collection). Each row is a valid Harbor task binary: columns path (str) and task_binary (gzip tar). Converted with the OpenThoughts-Agent data.nemotron_gym framework. Grading: JSON/YAML/TOML schema validation; XML/CSV structural (well-formed + required keys). texttext-generation10K<n<100K0 likes5 downloads22d agoHugging Face23open-athena /nemotron-gym-structured-outputs-v4 laion/nemotron-gym-structured-outputs-v4 Harbor task-binary dataset (53,870 tasks) converted from nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2 (part of nvidia/Nemotron-Post-Training-v3). Columns path (str) + task_binary (gzip tar). Converted with the OpenThoughts-Agent data.nemotron_gym framework. Grading: JSON/YAML/TOML schema validation; XML/CSV structural. What changed vs the prior version This version fixes the answer-delivery contract for… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-structured-outputs-v4.texttext-generation10K<n<100K0 likes4 downloads22d agoHugging Face24yasserrmd /TOON-Unstructured-Structuredgated TOON-Unstructured-Structured This dataset is a validated and cleaned version of the originalMasterControlAIML/JSON-Unstructured-Structured. It has been reformatted using the official Token-Oriented Object Notation (TOON) specification —a compact, token-efficient data serialization format optimized for LLM-ready structured data.All records have been verified for JSON integrity and TOON-decoding consistency. Overview Field Description text Original text… See the full description on the dataset page: https://huggingface.co/datasets/yasserrmd/TOON-Unstructured-Structured.texttext-generation1K<n<10K3 likes3 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.