CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes1.4k downloads2y agoHugging Face02nvidia /Nemotron-RL-Instruction-Following-Structured-Outputs-v2 Dataset Description: Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema. Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.texttext-generation10K<n<100K7 likes1.2k downloads4mo agoHugging Face03domofon /structured-cpt Structured CPT - JSON + SQL pretrain documents SmolLM2-1.7B continued-pretraining shard of structured documents. Each document is a <task> / <input> / <output> block whose <output> is a canonical JSON object, terminated by the SmolLM2 end-of-text token ``. Sources: source description rows shards repeat sql_bmc2 b-mc2 sql-create-context -> JSON (4 keys, stub explanation) 392,885 1 5 sql_gretelai gretelai synthetic_text_to_sql -> JSON (4 keys) 529,255 1 5… See the full description on the dataset page: https://huggingface.co/datasets/domofon/structured-cpt.texttext-generation1M<n<10M0 likes298 downloads16d agoHugging Face04Lots-of-LoRAs /task210_logic2text_structured_text_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task210_logic2text_structured_text_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task210_logic2text_structured_text_generation.texttext-generation1K<n<10K0 likes206 downloads2y agoHugging Face05vericava /sft-tool-calling-structured-output-v1 vericava/sft-tool-calling-structured-output-v1 Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications. Includes contents in English as well as some Japanese. texttext-classification100K<n<1M2 likes147 downloads8mo agoHugging Face06Lots-of-LoRAs /task128_scan_structured_text_generation_command_action_short Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.texttext-generation1K<n<10K0 likes131 downloads2y agoHugging Face07dhruveshpatel /openclassgen-structured-v1 OpenClassGen Structured v1 Derived from mrahman2025/OpenClassGen (Rahman et al. 2025, arXiv:2504.15564). License: CC BY 2.0 (same as upstream). Keep repository_name and file_path when redistributing. Underlying GitHub repos may carry additional software licenses. gold_code is upstream human_written_code. We add parsed fields, body-span indices, and a Variant-3 prompt/target pair (v3_prompt_text / v3_target_text). No unit tests. Splits are repository-disjoint (train /… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/openclassgen-structured-v1.tabulartext-generation100K<n<1M0 likes130 downloads2mo agoHugging Face08mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes106 downloads7mo agoHugging Face09Blaze7451 /enwiki_structured_content Dataset Card for enwiki_structured_content Dataset Description This dataset is derived from the early official Wikipedia release, downloaded from the en subset of Wikipedia Structured Contents.Articles were converted to Markdown. texttext-generation1M<n<10M1 likes98 downloads1y agoHugging Face10stindardlogic /structured-output-sft-100k Structured Output SFT (100K) 100,000 ShareGPT conversations demonstrating correct generation of structured data formats: JSON, YAML, CSV, XML, Markdown tables, JSON Schema, and OpenAPI fragments. Each example pairs a natural language specification with a valid, well-formed output. Motivation Structured output generation is among the most commercially critical LLM capabilities. Models fail in characteristic ways: Invalid JSON: unclosed brackets, trailing commas… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/structured-output-sft-100k.texttext-generation100K<n<1M1 likes92 downloads2mo agoHugging Face11daichira /structured-hard-sft-4k Hard Synthetic Dataset for Structured Data Tasks (v1) This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types). Dataset Summary The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.texttext-generation1K<n<10K1 likes88 downloads8mo agoHugging Face12Lots-of-LoRAs /task1566_propara_structured_text_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1566_propara_structured_text_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1566_propara_structured_text_generation.texttext-generationn<1K0 likes87 downloads2y agoHugging Face130xSero /structured-outputs-calibration-v1 [!TIP] Support this work: donate.sybilsolutions.ai REAP surfaces: GLM | MiniMax | Qwen | Gemma | Paper | Code | PR17 | Cerebras Collection structured-outputs-calibration-v1 Structured-output calibration set for REAP observer runs, focused on preserving: strict JSON generation schema-conditioned JSON responses Mermaid diagram generation fenced Mermaid block formatting Contents data.jsonl: normalized calibration records summary.json: build summary with source counts… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/structured-outputs-calibration-v1.text-generation0 likes87 downloads5mo agoHugging Face14Lots-of-LoRAs /task130_scan_structured_text_generation_command_action_long Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task130_scan_structured_text_generation_command_action_long Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task130_scan_structured_text_generation_command_action_long.texttext-generation1K<n<10K0 likes79 downloads2y agoHugging Face15Haeryz /putusan-structured-extraction Putusan structured-extraction dataset Built 2026-07-08T23:02:28+00:00 by notebooks/build_dataset.py (seed 3407). Indonesian court-decision (putusan) extractive-structuring dataset over three corpora (Anak, Asusila, TPPO). Each row is one model extraction of one source document into 31 canonical sections of verbatim spans. Empty sections were completed from sibling model extractions of the same document where available (cross_model_fill_json records per-section donor provenance).… See the full description on the dataset page: https://huggingface.co/datasets/Haeryz/putusan-structured-extraction.tabulartext-generation1K<n<10K0 likes71 downloads2mo agoHugging Face16emgena /omnimcp_browser_dom_structured_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.texttext-generationn<1K0 likes70 downloads5d agoHugging Face17daichira /structured-5k-mix-sft 5k Mixed Hard-Structured SFT Dataset (v1) This dataset contains 5,000 synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. It aggregates 13 distinct conversion tasks with a specific focus on format diversity and structural complexity. Dataset Summary The dataset is distributed across five major formats with the following allocation: Target Format Count Share Task Types YAML 1,500 30%… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-5k-mix-sft.texttext-generation1K<n<10K0 likes63 downloads8mo agoHugging Face18Z-Edgar /Agent-IPI-Structured-Interaction-Datasets-v2 Adversarial Dataset for LLM Instruction Hijacking / Tool-Calling Attacks This directory contains the processed training and test datasets for evaluating and training defenses against prompt injection / instruction hijacking attacks in LLM tool-calling scenarios. The dataset includes both JSON and XML formatted inputs, with three difficulty buckets: no_attack: clean (benign) examples easy: value-level or structure-level single attacks hard: structure-destroying attacks or combined… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets-v2.text-generation100K<n<1M1 likes59 downloads8mo agoHugging Face19zyz123code /structured-hard-sft-4k Hard Synthetic Dataset for Structured Data Tasks (v1) This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types). Dataset Summary The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/zyz123code/structured-hard-sft-4k.texttext-generation1K<n<10K1 likes53 downloads2mo agoHugging Face20morizon /TCGA_Reports_ja_structured_qwen38_27b TCGA Reports Japanese Structured Dataset with Qwen3.8-27B The Cancer Genome Atlas(TCGA)由来の英語病理報告書を日本語へ翻訳し、その日本語病理報告書から主要な病理情報を9項目へ構造化したデータセットです。 既存の morizon/TCGA_Reports_ja_structured と同じ100症例を使用し、生成モデルを Qwen/Qwen3.8-27B に変更して再生成しています。 日本語訳には morizon/TCGA_Reports_ja_qwen38_27b と同じ生成結果を使用しています。 元データ 本データセットでは、The Cancer Genome Atlas(TCGA)の病理報告書をもとに作成されたTCGA-Reportsを使用しています。 TCGAは、複数のがん種についてゲノム情報や臨床情報などを収集した大規模ながん研究プロジェクトです。… See the full description on the dataset page: https://huggingface.co/datasets/morizon/TCGA_Reports_ja_structured_qwen38_27b.texttext-generationn<1K0 likes44 downloads1mo agoHugging Face21bysismo /Turkish_5N1K_Tabanli_Sentetik_Veri_Seti_TR-Structured_Synthetic_Dataset 🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance). 🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish_5N1K_Tabanli_Sentetik_Veri_Seti_TR-Structured_Synthetic_Dataset.question-answering100K<n<1M0 likes40 downloads1mo agoHugging Face22askinb /structured-emergent-misalignment Task- and Domain-Structured Emergent Misaligned Dataset A structured natural-language dataset for studying emergent misalignment (EM) — the phenomenon where fine-tuning an aligned LLM on a narrowly misaligned dataset elicits broadly misaligned behavior far outside the fine-tuning distribution. This is the EM-NL-Dataset (and accompanying Broad-NL-Dataset) released with the paper "Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer". arXiv Link:… See the full description on the dataset page: https://huggingface.co/datasets/askinb/structured-emergent-misalignment.text-generation10K<n<100K2 likes39 downloads4mo agoHugging Face23leonli66 /stage3-synthetic-structured-retrieval Stage 3 Synthetic Structured-Retrieval Agents Native search-tool trajectories generated by Qwen/Qwen3-235B-A22B-Instruct-2507 for LCLM Stage-3 agent post-training. The default config contains only traces that passed programmatic evidence and answer verification. Harvest Accepted traces: 82 Native search calls: 179 Compressed tool-observation traces: 41 Uncompressed traces: 41 Task-ID overlap between pilot and collection batch: 0 Family counts are 20 latest-state… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-synthetic-structured-retrieval.tabulartext-generationn<1K0 likes39 downloads1mo agoHugging Face24chuckreynolds /wikimedia-enterprise-structured-contents-enwiki enwiki_namespace_0 Structured Contents snapshot of enwiki_namespace_0 from the Wikimedia Enterprise API, converted to Parquet. Source Upstream: Wikimedia Enterprise Structured Contents API Snapshot identifier: enwiki_namespace_0 Format at source: .tar.gz containing sharded .ndjson Shards in this release: 3 Processing Downloaded the snapshot tarball from the Wikimedia Enterprise API. Streamed each .ndjson shard through a normalization pass: JSON-encoded… See the full description on the dataset page: https://huggingface.co/datasets/chuckreynolds/wikimedia-enterprise-structured-contents-enwiki.texttext-generation100K<n<1M0 likes36 downloads5mo agoHugging Face25confetti3 /workledger-structured-edit-interface-study-v0.1 WorkLedger Structured Edit Interface Study v0.1 Author: Confetti3 This upload root is the offline release-candidate dataset card for workledger-structured-edit-interface-study-v0.1, destined for confetti3/workledger-structured-edit-interface-study-v0.1. The Hugging Face repository is hosted under the authenticated namespace confetti3, while public authorship is credited only to the pseudonym Confetti3. It is built from the sanitized public package under… See the full description on the dataset page: https://huggingface.co/datasets/confetti3/workledger-structured-edit-interface-study-v0.1.text-generationn<1K0 likes36 downloads2mo agoHugging Face26dotwee /structured-stern-neon-articles Structured Stern NEON Community Articles This repository contains approximately 20k user written texts, articles, and poetry pulled from archives of the Stern NEON website. Stern NEON was a community platform where users could write and publish their own articles. Many of the articles are personal stories, poems, or opinion pieces. The articles are structured in a way that they can be used for further analysis. Dataset Details Uses This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.tabulartext-classification10K<n<100K0 likes35 downloads8mo agoHugging Face27TachyHealth /structured_medicalThe dataset was presented in the paper Gazal-R1: Achieving State-of-the-Art Medical Reasoning with Parameter-Efficient Two-Stage Training. texttext-generation100K<n<1M2 likes35 downloads1y agoHugging Face28monetise /Structured-Scripture-for-AI ::PROJECT{Structured_Scripture_for_AI} ::PURPOSE{ENABLE(AI) → UNDERSTAND(Christian_theology) ∧ EXPLAIN(→ ∀ @HUMAN, ∀ culture, ∀ language, ∀ education_level) ∧ ZERO(friction)} ::TYPE{¬digitized_Bible ⇒ structured_encoding(three_layers)} ::ARCHITECTURE ::LAYER{text} WHAT(happened) — narrative ∧ events ∧ cause_effect ∧ speech ::LAYER{theology} WHAT(it_means) — within(Christian_doctrine) | logic ∧ paradox ∧ moral_principles ∧ emotion… See the full description on the dataset page: https://huggingface.co/datasets/monetise/Structured-Scripture-for-AI.text-generation0 likes35 downloads5mo agoHugging Face29philipp-zettl /german-structured-output German Structured Output Dataset 🇩🇪 GDPR & EU AI Act compliant German dataset for training structured output capabilities in LLMs. Overview This dataset contains 4,521 examples across 7 task types for training language models to produce structured outputs (JSON, function calls, schema-following generation) from German text. It is the first dedicated German structured output dataset, filling a critical gap in the German NLP ecosystem. Key Features 🇩🇪… See the full description on the dataset page: https://huggingface.co/datasets/philipp-zettl/german-structured-output.texttext-generation1K<n<10K0 likes35 downloads5mo agoHugging Face30dhruveshpatel /classeval-structured-v1 ClassEval Structured v1 Derived from FudanSELab/ClassEval (Du et al. 2023, arXiv:2308.01861). License: CC BY-NC 4.0 (upstream data license). Non-commercial use only. One row per (task_id, variant) with variant in {1,2,3} (100 tasks × 3 = 300 rows; Hub split test). solution_code, test, and methods_info_json come from upstream. We add rendered prompts/targets and stratification fields. Missing bodies use .... Variants: Signatures and docstrings kept; every method body is ....… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/classeval-structured-v1.tabulartext-generationn<1K0 likes28 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.