CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-RL-instruction_following-structured_outputs Dataset Description: The Nemotron-RL-instruction_following-structured_outputs dataset tests the ability of the model to follow output formatting instructions under schema constraints under the JSON format. Each problem consists of three components: The document, output formatting Instruction (Schema), and question. The dataset varies the difficulty of each problem by varying the location of instructions, the comprehensiveness of instructions, the complexity of the schema, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-instruction_following-structured_outputs.text1K<n<10K40 likes553 downloads8mo agoHugging Face02mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes180 downloads7mo agoHugging Face03gmahia /philosophy-classics-structured Classical Decision Frameworks — Philosophy Dataset Structured public domain philosophical texts focused on decision-making, leadership, and organizational ethics. All content is in the public domain. Content Works from classical philosophy structured for AI analysis: Stoic decision principles (Marcus Aurelius, Epictetus, Seneca) Political philosophy (Machiavelli, Aristotle) Virtue ethics (Aristotle, Plato) Sources All works published before 1928… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/philosophy-classics-structured.texttext-classificationn<1K0 likes168 downloads2mo agoHugging Face04vericava /sft-tool-calling-structured-output-v1 vericava/sft-tool-calling-structured-output-v1 Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications. Includes contents in English as well as some Japanese. texttext-classification100K<n<1M2 likes163 downloads8mo agoHugging Face05ysmao /structured3d-spatiallm Structured3D-SpatialLM Dataset Structured3D dataset preprocessed in SpatialLM format for layout estimation with LLMs. Overview This dataset is derived from Structured3D 3,500 synthetic house designs created by professional designers, preprocessed and formatted specifically for SpatialLM training. Point clouds and layouts are derived from the RoomFormer data preprocessing script. Data Extraction Point clouds and layouts are compressed in zip files. To… See the full description on the dataset page: https://huggingface.co/datasets/ysmao/structured3d-spatiallm.text1K<n<10K1 likes138 downloads1y agoHugging Face06sxiong /MLR_structured_trajectory Reasoning Trajectories with Step-Level Annotations This dataset contains structured reasoning trajectories introduced in (ICLR 2026) Enhancing Language Model Reasoning with Structured Multi-Level Modeling. Compared with the full-trajectory release, this dataset is a cleaned and segmented version designed for research on hierarchical reasoning, trajectory supervision, and multi-step policy training. Each example contains the original prompt and response fields together with a… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/MLR_structured_trajectory.tabular10K<n<100K1 likes137 downloads3mo agoHugging Face07stindardlogic /structured-output-sft-100k Structured Output SFT (100K) 100,000 ShareGPT conversations demonstrating correct generation of structured data formats: JSON, YAML, CSV, XML, Markdown tables, JSON Schema, and OpenAPI fragments. Each example pairs a natural language specification with a valid, well-formed output. Motivation Structured output generation is among the most commercially critical LLM capabilities. Models fail in characteristic ways: Invalid JSON: unclosed brackets, trailing commas… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/structured-output-sft-100k.texttext-generation100K<n<1M1 likes113 downloads2mo agoHugging Face08electroglyph /Nemotron-RL-Instruction-Following-Structured-Outputs-v2-preferencetext10K<n<100K0 likes73 downloads3mo agoHugging Face09u-10bei /structured_data_with_cot_dataset_512_v4 structured_data_with_cot_dataset このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。 データセットの概要 messages: OpenAIチャット形式 (system, user, assistant) metadata: format, complexity, schema, estimated_tokens サポートされるデータ形式 JSON, XML, YAML, TOML, CSV 生成方法 Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。 text1K<n<10K1 likes71 downloads8mo agoHugging Face10Z-Edgar /Agent-IPI-Structured-Interaction-Datasets Dataset Card for Indirect Prompt Injection in Agent Structured Interaction Datasets Dataset Summary This dataset contains 470,000 QA pairs designed to study indirect prompt injection in agent-structured interactions. It is split into a training set (80%) and a test set (20%). The dataset is evenly divided into 50% clean-clean QA pairs (no prompt injection) and 50% clean-injected QA pairs (containing prompt injection). The task is to detect and remove prompt injection… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets.text100K<n<1M1 likes66 downloads9mo agoHugging Face11u-10bei /structured_data_with_cot_dataset_512_v5 structured_data_with_cot_dataset このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。 データセットの概要 messages: OpenAIチャット形式 (system, user, assistant) metadata: format, complexity, schema, estimated_tokens サポートされるデータ形式 JSON, XML, YAML, TOML, CSV 生成方法 Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。 v5アップデート:ランダムなスキーマ構造の生成と、最小化(minified)/ソート(sorted)の制約を追加。 text1K<n<10K3 likes55 downloads8mo agoHugging Face12Blaze7451 /enwiki_structured_content Dataset Card for enwiki_structured_content Dataset Description This dataset is derived from the early official Wikipedia release, downloaded from the en subset of Wikipedia Structured Contents.Articles were converted to Markdown. texttext-generation1M<n<10M1 likes51 downloads1y agoHugging Face13daichira /structured-hard-sft-4k Hard Synthetic Dataset for Structured Data Tasks (v1) This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types). Dataset Summary The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.texttext-generation1K<n<10K1 likes51 downloads8mo agoHugging Face14TachyHealth /structured_medicalThe dataset was presented in the paper Gazal-R1: Achieving State-of-the-Art Medical Reasoning with Parameter-Efficient Two-Stage Training. texttext-generation100K<n<1M2 likes50 downloads1y agoHugging Face15zyz123code /structured-hard-sft-4k Hard Synthetic Dataset for Structured Data Tasks (v1) This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types). Dataset Summary The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/zyz123code/structured-hard-sft-4k.texttext-generation1K<n<10K1 likes50 downloads2mo agoHugging Face16Nexdata-kr /1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample Description 한국어 시험 문제 구조화 분석·가공 데이터로, 약 150만 개의 시험 문제를 포함하고 있습니다. 문제 유형, 문제, 정답, 해설 등의 정보를 포함하며, 과목은 [초등학교] 국어, 수학, 영어, 사회, 과학; [중학교] 국어, 영어, 수학, 과학, 사회; [고등학교] 국어, 영어, 수학, 물리, 화학, 생물, 역사, 지리로 구성되어 있습니다. 문제 유형에는 객관식, 빈칸 채우기, 참·거짓 문제, 단답형 문제 등이 포함됩니다. 본 데이터셋은 대규모 교과 지식 강화 및 학습 데이터 구축 등의 작업에 활용할 수 있습니다. 자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/1634?source=Hf.kr Specifications Data content 한국어 K12 시험 문제 Amount 약 150만 개의… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.textn<1K0 likes49 downloads16d agoHugging Face17u-10bei /structured_data_with_cot_dataset_512_v3 structured_data_with_cot_dataset このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。 データセットの概要 messages: OpenAIチャット形式 (system, user, assistant) metadata: format, complexity, schema, estimated_tokens サポートされるデータ形式 JSON, XML, YAML, TOML, CSV 生成方法 Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。 text1K<n<10K0 likes44 downloads8mo agoHugging Face18tomyimkc /repro-on-structured-state-space-duality-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes44 downloads2mo agoHugging Face19Gen3DF /Structured3Dgated Structured3D-SpatialLM Dataset Structured3D dataset preprocessed in SpatialLM format for layout estimation with LLMs. Overview This dataset is derived from Structured3D 3,500 synthetic house designs created by professional designers, preprocessed and formatted specifically for SpatialLM training. Point clouds and layouts are derived from the RoomFormer data preprocessing script. Data Extraction Point clouds and layouts are compressed in zip files. To… See the full description on the dataset page: https://huggingface.co/datasets/Gen3DF/Structured3D.3d1K<n<10K2 likes39 downloads1y agoHugging Face20daichira /structured-5k-mix-sft 5k Mixed Hard-Structured SFT Dataset (v1) This dataset contains 5,000 synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. It aggregates 13 distinct conversion tasks with a specific focus on format diversity and structural complexity. Dataset Summary The dataset is distributed across five major formats with the following allocation: Target Format Count Share Task Types YAML 1,500 30%… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-5k-mix-sft.texttext-generation1K<n<10K0 likes39 downloads8mo agoHugging Face21dotwee /structured-stern-neon-articles Structured Stern NEON Community Articles This repository contains approximately 20k user written texts, articles, and poetry pulled from archives of the Stern NEON website. Stern NEON was a community platform where users could write and publish their own articles. Many of the articles are personal stories, poems, or opinion pieces. The articles are structured in a way that they can be used for further analysis. Dataset Details Uses This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.tabulartext-classification10K<n<100K0 likes38 downloads8mo agoHugging Face22xnileshtiwari /CBSE-Class-12th_2024_PYQs__structuredThis data set contains the CBSE Class 12 2024 papers in a structured format. The papers are annotated with topic and chapter names, and the figures are parsed and their paths annotated. tabularquestion-answeringn<1K2 likes38 downloads2y agoHugging Face23tomyimkc /repro-learning-in-structured-stackelberg-games-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes35 downloads2mo agoHugging Face24Nexdata-AI /1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample Description Korean Test Questions Structured Analysis Processing Data, around 1.5 million questions, contains question types, questions, answers, explanations, etc..For subjects, include [Primary School] Korean, Mathematics, English, Social Studies, Science; [Middle School] Korean, English, Mathematics, Science, Social Studies; [High School] Korean, English, Mathematics, Physics, Chemistry, Biology, History, Geography; question Types indlude single-choice question, fill-in… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.textn<1K0 likes35 downloads1mo agoHugging Face25leonli66 /stage3-synthetic-structured-retrieval Stage 3 Synthetic Structured-Retrieval Agents Native search-tool trajectories generated by Qwen/Qwen3-235B-A22B-Instruct-2507 for LCLM Stage-3 agent post-training. The default config contains only traces that passed programmatic evidence and answer verification. Harvest Accepted traces: 82 Native search calls: 179 Compressed tool-observation traces: 41 Uncompressed traces: 41 Task-ID overlap between pilot and collection batch: 0 Family counts are 20 latest-state… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-synthetic-structured-retrieval.tabulartext-generationn<1K0 likes32 downloads1mo agoHugging Face26structured-reasoning /dev-0913textn<1K0 likes26 downloads1y agoHugging Face27stindardlogic /json-structured-output-dpo-3k JSON Structured Output DPO Pairs (3K) DPO preference pairs for training LLMs to produce valid, schema-compliant JSON output. Motivation Structured output (JSON mode) is critical for production AI applications — parsers fail, pipelines break, and downstream processing errors when models output malformed JSON, use wrong field names, or wrap responses in markdown. This dataset trains strict schema adherence. Dataset Description 3,000 preference pairs… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/json-structured-output-dpo-3k.texttext-generation1K<n<10K0 likes26 downloads2mo agoHugging Face28daichira /structured-orpo-balanced-5fmt-490-v2 Structured ORPO (Balanced 5 Formats, 490 each) This dataset contains 2,450 preference pairs for ORPO training, balanced across five structured formats (JSON, YAML, TOML, XML, CSV) with 490 pairs each. It is derived from u-10bei/structured_data_with_cot_dataset_512_v4 and reconstructed to preserve syntax while emphasizing instruction-following errors. v2 Notes (margin + spec-augmentation) Prompt constraints are injected to make violations explicit (e.g., "Return only JSON… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-orpo-balanced-5fmt-490-v2.text1K<n<10K0 likes25 downloads7mo agoHugging Face29Juliankrg /Web_StructureDataSet_100k Dataset Name: Web Structure Fine-Tuning Dataset for AI Description: This dataset is designed to support the fine-tuning of AI models focused on understanding and generating explanations about web technologies and structures. It contains 100,000 varied examples covering fundamental web concepts such as HTML, CSS, JavaScript, HTTP, APIs, SEO, and responsive design. Each example consists of a prompt and a corresponding completion, providing a comprehensive resource for… See the full description on the dataset page: https://huggingface.co/datasets/Juliankrg/Web_StructureDataSet_100k.text100K<n<1M1 likes24 downloads1y agoHugging Face30ToshiyukiNH /structured_data_with_cot_dataset_512_v2_filtered_3structured_data_with_cot_dataset_512_v2_filtered_3 This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2. Usage from datasets import load_dataset # From local data dataset = load_dataset( 'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_3", split='train' ) print(dataset[0]) How to generate this dataset from the base one from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_3.textn<1K0 likes22 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.