datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.structured_data_with_cot_dataset_512_v4
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。
データセットの概要
messages: OpenAIチャット形式 (system, user, assistant)
metadata: format, complexity, schema, estimated_tokens
サポートされるデータ形式
JSON, XML, YAML, TOML, CSV
生成方法
Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。
brand-structured-data-reference
Brand Structured Data Reference v1.0
This reference maps common public brand facts to structured data concepts that can help people, search engines, and AI systems understand a brand more clearly.
It is intended for independent brands, small businesses, founder-led companies, service providers, local businesses, and early-stage products that need a clearer public identity online.
This is not a ranking guide and it does not guarantee search visibility, rich results, AI… See the full description on the dataset page: https://huggingface.co/datasets/farosio/brand-structured-data-reference.Agent-IPI-Structured-Interaction-Datasets
Dataset Card for Indirect Prompt Injection in Agent Structured Interaction Datasets
Dataset Summary
This dataset contains 470,000 QA pairs designed to study indirect prompt injection in agent-structured interactions. It is split into a training set (80%) and a test set (20%). The dataset is evenly divided into 50% clean-clean QA pairs (no prompt injection) and 50% clean-injected QA pairs (containing prompt injection). The task is to detect and remove prompt injection… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets.structured_data_with_cot_dataset_512_v5
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。
データセットの概要
messages: OpenAIチャット形式 (system, user, assistant)
metadata: format, complexity, schema, estimated_tokens
サポートされるデータ形式
JSON, XML, YAML, TOML, CSV
生成方法
Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。
v5アップデート:ランダムなスキーマ構造の生成と、最小化(minified)/ソート(sorted)の制約を追加。
1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample
Description
한국어 시험 문제 구조화 분석·가공 데이터로, 약 150만 개의 시험 문제를 포함하고 있습니다. 문제 유형, 문제, 정답, 해설 등의 정보를 포함하며, 과목은 [초등학교] 국어, 수학, 영어, 사회, 과학; [중학교] 국어, 영어, 수학, 과학, 사회; [고등학교] 국어, 영어, 수학, 물리, 화학, 생물, 역사, 지리로 구성되어 있습니다. 문제 유형에는 객관식, 빈칸 채우기, 참·거짓 문제, 단답형 문제 등이 포함됩니다. 본 데이터셋은 대규모 교과 지식 강화 및 학습 데이터 구축 등의 작업에 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/1634?source=Hf.kr
Specifications
Data content
한국어 K12 시험 문제
Amount
약 150만 개의… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.structured_data_with_cot_dataset_512_v3
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。
データセットの概要
messages: OpenAIチャット形式 (system, user, assistant)
metadata: format, complexity, schema, estimated_tokens
サポートされるデータ形式
JSON, XML, YAML, TOML, CSV
生成方法
Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。
1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample
Description
Korean Test Questions Structured Analysis Processing Data, around 1.5 million questions, contains question types, questions, answers, explanations, etc..For subjects, include [Primary School] Korean, Mathematics, English, Social Studies, Science; [Middle School] Korean, English, Mathematics, Science, Social Studies; [High School] Korean, English, Mathematics, Physics, Chemistry, Biology, History, Geography; question Types indlude single-choice question, fill-in… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.structured_data_with_cot_dataset_512_v2
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。このデータセットの目的は、自然言語プロンプトに基づいて構造化データを生成できるモデルに対し、CoTを活用してより良い推論と出力品質を達成するための高品質な学習例を提供することです。
データセットの概要
データセットの各エントリには、以下のものが含まれます。
messages: OpenAIチャット形式に従ったメッセージ辞書のリストで、以下で構成されます。
systemメッセージ: 特定のデータ形式の専門家としてのコンテキストを設定します。
userメッセージ: 特定のスキーマと形式で構造化データの生成を要求するプロンプトです。多様なプロンプトテンプレートが使用されています。
assistantメッセージ: 生成された構造化データに続く思考連鎖(CoT)推論が含まれます。… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/structured_data_with_cot_dataset_512_v2.structured_data_with_cot_dataset_512_collected_from_v2v4v5_unique
Dataset creation
This dataset was merged by collecting conversion type records only
from v2, v4, and v5 of u-10bei/structured_data_with_cot_512.
and removed duplicated records.
Total records before removing duplicates:
5530 (Sources: 1485 from v2, 2037 from v4, 2008 from v5)
Type: conversion
Final records after removing duplicated records: 5451 (79 messages duplicated)
Collection method
Records with its type as conversion were collected.
License
The license… See the full description on the dataset page: https://huggingface.co/datasets/Kauhiro/structured_data_with_cot_dataset_512_collected_from_v2v4v5_unique.structured_data_with_cot_dataset_512_crean
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。このデータセットの目的は、自然言語プロンプトに基づいて構造化データを生成できるモデルに対し、CoTを活用してより良い推論と出力品質を達成するための高品質な学習例を提供することです。
データセットの概要
データセットの各エントリには、以下のものが含まれます。
messages: OpenAIチャット形式に従ったメッセージ辞書のリストで、以下で構成されます。
systemメッセージ: 特定のデータ形式の専門家としてのコンテキストを設定します。
userメッセージ: 特定のスキーマと形式で構造化データの生成を要求するプロンプトです。多様なプロンプトテンプレートが使用されています。
assistantメッセージ: 生成された構造化データに続く思考連鎖(CoT)推論が含まれます。… See the full description on the dataset page: https://huggingface.co/datasets/sasa5555/structured_data_with_cot_dataset_512_crean.structured_data_with_cot_dataset_512_collected_from_v2v4v5
Merged dataset from u-10bei/structured_data_with_cot_dataset_512_v2, v4, and v5
Collected conversion type records from v2, v4, and v5 of structured_data_with_cot_512.
Total records: 5530 (Sources: 1485 from v2, 2037 from v4, 2008 from v5)
Type: conversion
Collection method
Records with its type as conversion were collected.
License
The license of this dataset inherits from the license of the original dataset.… See the full description on the dataset page: https://huggingface.co/datasets/Kauhiro/structured_data_with_cot_dataset_512_collected_from_v2v4v5.structured_data_with_cot_dataset_512_v2_filtered_3structured_data_with_cot_dataset_512_v2_filtered_3
This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2.
Usage
from datasets import load_dataset
# From local data
dataset = load_dataset(
'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_3", split='train'
)
print(dataset[0])
How to generate this dataset from the base one
from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_3.10bei_structured_data_with_cot_dataset_512_v2_constraints_added_ver2structured_data_merged_v2v5_0222
Dataset Card for structured_data_merged_v2v5_0222
Dataset Details
Dataset Description
structured_data_merged_v2v5_0222 is a dataset for Supervised Fine-Tuning (SFT) focused on structured data format conversion tasks — specifically, interconversion among JSON, XML, YAML, TOML, and CSV.
It was created by deduplicating and merging the following two existing datasets:
u-10bei/structured_data_with_cot_dataset_512_v2 (train split only)… See the full description on the dataset page: https://huggingface.co/datasets/takami2022/structured_data_merged_v2v5_0222.structured_data_with_cot_dataset_v2structured_data_with_cot_dataset_512structured_data_with_cot_dataset_512_v2_r0.6
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。このデータセットの目的は、自然言語プロンプトに基づいて構造化データを生成できるモデルに対し、CoTを活用してより良い推論と出力品質を達成するための高品質な学習例を提供することです。
データセットの概要
データセットの各エントリには、以下のものが含まれます。
messages: OpenAIチャット形式に従ったメッセージ辞書のリストで、以下で構成されます。
systemメッセージ: 特定のデータ形式の専門家としてのコンテキストを設定します。
userメッセージ: 特定のスキーマと形式で構造化データの生成を要求するプロンプトです。多様なプロンプトテンプレートが使用されています。
assistantメッセージ: 生成された構造化データに続く思考連鎖(CoT)推論が含まれます。… See the full description on the dataset page: https://huggingface.co/datasets/naoyasss/structured_data_with_cot_dataset_512_v2_r0.6.humanoid-domestic-task-structured-dataset-v2
Humanoid Domestic Task Structured Dataset
Overview
This dataset contains structured human instructions for basic
household assistance scenarios. It is designed to help humanoid
agents interpret natural language commands and convert them into
clear executable task representations.
The dataset focuses on simple real-world domestic tasks that reduce
human workload and improve everyday living environments.
Key Features
Natural human-written instructions
Structured… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/humanoid-domestic-task-structured-dataset-v2.combined-structured-dataset
Combined Structured Output Dataset
このデータセットは、構造化出力生成タスクのための統合データセットです。
Dataset Details
Total Samples: 3,425
Source Datasets:
u-10bei/structured_data_with_cot_dataset_v2 (2,500 samples)
daichira/structured-3k-mix-sft (3,000 samples)
Preprocessing:
Format normalization to 3-turn (system/user/assistant)
Deduplication by user content
Quality filtering
Format Distribution
JSON: 685 (20.0%)
YAML: 485 (14.2%)
TOML: 685 (20.0%)
XML: 885 (25.8%)
CSV: 685… See the full description on the dataset page: https://huggingface.co/datasets/kurota0612/combined-structured-dataset.structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetstructured_data_short_512structured_data_with_cot_dataset_512_v2_clean_v1
MSakae/structured_data_with_cot_dataset_512_v2_clean_v1
This dataset is a cleaned version of u-10bei/structured_data_with_cot_dataset_512_v2.
Cleaning policy
Read metadata.format and validate the assistant output by parsing:
json → json.loads
yaml → yaml.safe_load
xml → xml.etree.ElementTree.fromstring
toml → tomllib.loads
csv → csv.reader (light sanity checks)
Extract output from the final assistant message after the earliest output marker:
Output:, OUTPUT:, Final:… See the full description on the dataset page: https://huggingface.co/datasets/MSakae/structured_data_with_cot_dataset_512_v2_clean_v1.structured_data_with_cot_dataset_512_v2_filtered_1structured_data_with_cot_dataset_512_v2_filtered_1
This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2.
Usage
from datasets import load_dataset
# From local data
dataset = load_dataset(
'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_1", split='train'
)
print(dataset[0])
How to generate this dataset from the base one
from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_1.structured_data_with_cot_dataset_512_v2_filtered_2structured_data_with_cot_dataset_512_v2_filtered_2
This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2.
Usage
from datasets import load_dataset
# From local data
dataset = load_dataset(
'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_2", split='train'
)
print(dataset[0])
How to generate this dataset from the base one
from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_2.structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetstructured_data_with_cot_dataset_512_v5_cleanedstructured_data_with_cot_dataset_512_v2_filtered_4structured_data_with_cot_dataset_512_v2_filtered_4
This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2.
Usage
from datasets import load_dataset
# From local data
dataset = load_dataset(
'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_4", split='train'
)
print(dataset[0])
How to generate this dataset from the base one
from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_4.structured_data_with_cot_dataset_512_concatenate_v2_v4_v5This is the concatenated datasets of
u-10bei/structured_data_with_cot_dataset_512_v2 (train 3933 records)
u-10bei/structured_data_with_cot_dataset_512_v4 (train 4608 records)
u-10bei/structured_data_with_cot_dataset_512_v5 (train 4547 records)
without sampling.
cv-structured-dataset
