CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wikimedia /structured-wikipedia Dataset Card for Wikimedia Structured Wikipedia Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta-Wiki Discussion Dataset Summary Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/structured-wikipedia.text10M<n<100M394 likes14k downloads4mo agoHugging Face02Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes1.4k downloads2y agoHugging Face03nvidia /Nemotron-RL-Instruction-Following-Structured-Outputs-v2 Dataset Description: Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema. Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.texttext-generation10K<n<100K7 likes1.2k downloads4mo agoHugging Face04open-athena /a3-rl-laion_nemotron-gym-instruction-following-structuredtext10K<n<100K0 likes1.1k downloads4mo agoHugging Face05anonymous-structured-agent /structured-file-audit-benchmark Paper Data Release This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them. Contents datasets/ Benchmark data and per-task manifests for the three paper-facing splits. datasets/sc_flat/data SC-Flat is derived from DaBench, augmented with a replayable perturbation injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.texttable-question-answering1 likes658 downloads2mo agoHugging Face06nvidia /Nemotron-RL-instruction_following-structured_outputs Dataset Description: The Nemotron-RL-instruction_following-structured_outputs dataset tests the ability of the model to follow output formatting instructions under schema constraints under the JSON format. Each problem consists of three components: The document, output formatting Instruction (Schema), and question. The dataset varies the difficulty of each problem by varying the location of instructions, the comprehensiveness of instructions, the complexity of the schema, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-instruction_following-structured_outputs.text1K<n<10K40 likes563 downloads8mo agoHugging Face07terminusresearch /pseudo-camera-10k-structured-json pseudo-camera-10k, structured JSON captions The 9,997 training images from bghira/pseudo-camera-10k, recaptioned into the structured JSON caption schema that Ideogram 4 consumes. The images are unchanged: free photographs from world class photographers, Lanczos-resized so the shorter edge is 1024px, nothing upsampled. The original dataset carries short CogVLM prose captions. This one replaces them with one JSON object per image describing the scene at three levels: an overall… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/pseudo-camera-10k-structured-json.imagetext-to-image10K<n<100K0 likes332 downloads12d agoHugging Face08open-athena /nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes302 downloads2mo agoHugging Face09Aregay01 /structured-wikipedia Dataset Card for Wikimedia Structured Wikipedia Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta-Wiki Discussion Dataset Summary Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/Aregay01/structured-wikipedia.text10M<n<100M0 likes298 downloads4mo agoHugging Face10domofon /structured-cpt Structured CPT - JSON + SQL pretrain documents SmolLM2-1.7B continued-pretraining shard of structured documents. Each document is a <task> / <input> / <output> block whose <output> is a canonical JSON object, terminated by the SmolLM2 end-of-text token ``. Sources: source description rows shards repeat sql_bmc2 b-mc2 sql-create-context -> JSON (4 keys, stub explanation) 392,885 1 5 sql_gretelai gretelai synthetic_text_to_sql -> JSON (4 keys) 529,255 1 5… See the full description on the dataset page: https://huggingface.co/datasets/domofon/structured-cpt.texttext-generation1M<n<10M0 likes298 downloads16d agoHugging Face11open-athena /nemotron-gym-instruction-following-structured-minimax-m27-131k-tracestext1K<n<10K0 likes245 downloads4mo agoHugging Face12GoktugD /turkish-structured-summarization-1.5m Turkish Structured Summarization 1.5M v2 Üç cümlelik kurgusal operasyon kayıtları ve kısa Türkçe özetleri. Doğrulanmış boyut Train: 1,470,000 Validation: 15,000 Test: 15,000 Toplam: 1,500,000 Ana görev sütunları: id, document, summary, domain Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-structured-summarization-1.5m.textsummarization1M<n<10M0 likes236 downloads1mo agoHugging Face13safelegalaidata /eu-ai-act-structured EU AI Act, structured Regulation (EU) 2024/1689 (the Artificial Intelligence Act) as tables: every article, recital, annex and definition, 677 obligations coded by actor, risk tier, application date and penalty basis, plus milestones, national competent authorities and fine tiers. Built 2026-09-08 by SafeLegalAI (Cognesio LLP) from the official English texts served by the Publications Office of the European Union (Cellar): the consolidated text as of 27 July 2026 (CELEX… See the full description on the dataset page: https://huggingface.co/datasets/safelegalaidata/eu-ai-act-structured.tabular1K<n<10K1 likes232 downloads14d agoHugging Face14Lots-of-LoRAs /task210_logic2text_structured_text_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task210_logic2text_structured_text_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task210_logic2text_structured_text_generation.texttext-generation1K<n<10K0 likes206 downloads2y agoHugging Face15JoyboyBrian /structured_imagesimage100K<n<1M1 likes182 downloads2y agoHugging Face16sxiong /MLR_structured_trajectory Reasoning Trajectories with Step-Level Annotations This dataset contains structured reasoning trajectories introduced in (ICLR 2026) Enhancing Language Model Reasoning with Structured Multi-Level Modeling. Compared with the full-trajectory release, this dataset is a cleaned and segmented version designed for research on hierarchical reasoning, trajectory supervision, and multi-step policy training. Each example contains the original prompt and response fields together with a… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/MLR_structured_trajectory.tabular10K<n<100K1 likes169 downloads3mo agoHugging Face17gmahia /philosophy-classics-structured Classical Decision Frameworks — Philosophy Dataset Structured public domain philosophical texts focused on decision-making, leadership, and organizational ethics. All content is in the public domain. Content Works from classical philosophy structured for AI analysis: Stoic decision principles (Marcus Aurelius, Epictetus, Seneca) Political philosophy (Machiavelli, Aristotle) Virtue ethics (Aristotle, Plato) Sources All works published before 1928… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/philosophy-classics-structured.texttext-classificationn<1K0 likes164 downloads2mo agoHugging Face18sachithgunasekara /phased-self-discover-mistral-structured-5-shot-bbh-evaltext1K<n<10K0 likes158 downloads2y agoHugging Face19vericava /sft-tool-calling-structured-output-v1 vericava/sft-tool-calling-structured-output-v1 Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications. Includes contents in English as well as some Japanese. texttext-classification100K<n<1M2 likes147 downloads8mo agoHugging Face20strongminsu /ko-en-structured-translations Korean–English Multistyle Parallel Corpus 한국어 사용자에게 익숙한 표현 기반의 다도메인·다문체 한–영 병렬 코퍼스 소개(Introduction) 저는 머신러닝, 인공지능 수업을 진행하는 강사입니다.Seq2Seq, Attention, Transformer 등 자연어처리(NLP) 수업을 진행하며한국 학습자에게 자연스럽고 익숙한 한–영 번역 데이터셋의 부족을 경험했습니다. 기존 공개 데이터셋은 도메인 다양성이 부족하거나 문체가 한국 사용자에게 자연스럽지 않거나 전반적으로 문장의 퀄리티가 매우 부족하여 학습한 번역 모델의 실제 성능이 기대만큼 나오지 않는 문제가 있었습니다. 이 문제를 해결하기 위해, 딥러닝 강사로서 langchain을 사용하여 직접 고품질 병렬 데이터를 자동으로 생성·정제하여 구성한 데이터셋입니다. 한국어 사용자에게 익숙한 표현을 중심으로 다양한 문체, 문장 구조를… See the full description on the dataset page: https://huggingface.co/datasets/strongminsu/ko-en-structured-translations.texttranslation1K<n<10K17 likes136 downloads10mo agoHugging Face21yasalma /tt-structured-contentgated Dataset Summary This dataset contains structured textual content in Markdown format extracted from Tatar-language documents, originally in EPUB and PDF formats. The documents include books and other long-form content with rich formatting. The dataset is intended to provide clean, structured, and semantically meaningful content to support natural language processing tasks, content modeling, and research in Tatar language technologies. The extracted Markdown preserves key… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tt-structured-content.text10K<n<100K1 likes133 downloads7mo agoHugging Face22system-technologies /MedCase-Structured MedCase-Structured Dataset for Paper MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings Structured FHIR R4 representations of clinical reasoning cases, derived from the MedCaseReasoning dataset (Wu et al., 2025). Each case pairs a free-text clinical presentation with a machine-readable FHIR bundle and a held-out ground-truth diagnosis, supporting evaluation of clinical information extraction, terminology coding… See the full description on the dataset page: https://huggingface.co/datasets/system-technologies/MedCase-Structured.text1K<n<10K1 likes133 downloads3mo agoHugging Face23jamesdborin /Nemotron-Agents-Tools-and-Structured-Task-Execution-prompt-only Agents, Tools and Structured Task Execution Prompt-Only This dataset combines prompt-only datasets by capability theme for distillation experiments. It contains 675,882 unique prompts from 808,884 raw rows; 133,002 exact canonical duplicates were removed. Rows retain the canonical prompt-extraction columns and add source_repo_id for provenance. Deduplication uses normalized system_prompt, prompt, tools, and schema_str, with the first row in manifest order retained. Original… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-Agents-Tools-and-Structured-Task-Execution-prompt-only.tabular100K<n<1M0 likes132 downloads2mo agoHugging Face24frangelbarrera /cyber-evidence-kev-structured Cyber Security Evidence Dataset — CISA KEV Structured CC0 Layer This configuration is the structured CISA Known Exploited Vulnerabilities (KEV) layer of the broader Cyber Security Evidence Dataset project. It contains 1,687 deterministic records generated from the official CISA KEV database snapshot. What is included The records contain the official KEV database fields: CVE identifier, vendor/project, product, vulnerability name, short description, dates… See the full description on the dataset page: https://huggingface.co/datasets/frangelbarrera/cyber-evidence-kev-structured.texttext-classification1K<n<10K1 likes132 downloads19d agoHugging Face25Lots-of-LoRAs /task128_scan_structured_text_generation_command_action_short Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.texttext-generation1K<n<10K0 likes131 downloads2y agoHugging Face26dhruveshpatel /openclassgen-structured-v1 OpenClassGen Structured v1 Derived from mrahman2025/OpenClassGen (Rahman et al. 2025, arXiv:2504.15564). License: CC BY 2.0 (same as upstream). Keep repository_name and file_path when redistributing. Underlying GitHub repos may carry additional software licenses. gold_code is upstream human_written_code. We add parsed fields, body-span indices, and a Variant-3 prompt/target pair (v3_prompt_text / v3_target_text). No unit tests. Splits are repository-disjoint (train /… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/openclassgen-structured-v1.tabulartext-generation100K<n<1M0 likes130 downloads2mo agoHugging Face27ysmao /structured3d-spatiallm Structured3D-SpatialLM Dataset Structured3D dataset preprocessed in SpatialLM format for layout estimation with LLMs. Overview This dataset is derived from Structured3D 3,500 synthetic house designs created by professional designers, preprocessed and formatted specifically for SpatialLM training. Point clouds and layouts are derived from the RoomFormer data preprocessing script. Data Extraction Point clouds and layouts are compressed in zip files. To… See the full description on the dataset page: https://huggingface.co/datasets/ysmao/structured3d-spatiallm.text1K<n<10K1 likes123 downloads1y agoHugging Face28wshuai190 /hotpotqa-structuredThis dataset is associated with the paper Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents. Official GitHub repository: https://github.com/ielab/skim-search-agent textquestion-answering10K<n<100K0 likes114 downloads2mo agoHugging Face29NewCarbon37 /structured-wikipedia Dataset Card for Wikimedia Structured Wikipedia Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta-Wiki Discussion Dataset Summary Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/NewCarbon37/structured-wikipedia.text10M<n<100M0 likes110 downloads3mo agoHugging Face30mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes106 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.