datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.Agent-IPI-Structured-Interaction-Datasets-v2
Adversarial Dataset for LLM Instruction Hijacking / Tool-Calling Attacks
This directory contains the processed training and test datasets for evaluating and training defenses against prompt injection / instruction hijacking attacks in LLM tool-calling scenarios.
The dataset includes both JSON and XML formatted inputs, with three difficulty buckets:
no_attack: clean (benign) examples
easy: value-level or structure-level single attacks
hard: structure-destroying attacks or combined… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets-v2.Turkish_5N1K_Tabanli_Sentetik_Veri_Seti_TR-Structured_Synthetic_Dataset
🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance).
🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish_5N1K_Tabanli_Sentetik_Veri_Seti_TR-Structured_Synthetic_Dataset.synthetic-structured-output-dataset
Synthetic Structured Output Dataset (SFT + DPO)
Synthetic training corpus for structured-output model tuning. This package contains SFT and DPO data focused on JSON schema compliance, structured extraction, and function calling.
Included files
sft_synthetic_json.jsonl — SFT samples for schema-conditioned JSON generation
sft_synthetic_extraction.jsonl — SFT samples for text-to-structured extraction
dpo_structured_output.jsonl — DPO chosen/rejected pairs for structured… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/synthetic-structured-output-dataset.structured_data_merged_v2v5_0222
Dataset Card for structured_data_merged_v2v5_0222
Dataset Details
Dataset Description
structured_data_merged_v2v5_0222 is a dataset for Supervised Fine-Tuning (SFT) focused on structured data format conversion tasks — specifically, interconversion among JSON, XML, YAML, TOML, and CSV.
It was created by deduplicating and merging the following two existing datasets:
u-10bei/structured_data_with_cot_dataset_512_v2 (train split only)… See the full description on the dataset page: https://huggingface.co/datasets/takami2022/structured_data_merged_v2v5_0222.combined-structured-dataset
Combined Structured Output Dataset
このデータセットは、構造化出力生成タスクのための統合データセットです。
Dataset Details
Total Samples: 3,425
Source Datasets:
u-10bei/structured_data_with_cot_dataset_v2 (2,500 samples)
daichira/structured-3k-mix-sft (3,000 samples)
Preprocessing:
Format normalization to 3-turn (system/user/assistant)
Deduplication by user content
Quality filtering
Format Distribution
JSON: 685 (20.0%)
YAML: 485 (14.2%)
TOML: 685 (20.0%)
XML: 885 (25.8%)
CSV: 685… See the full description on the dataset page: https://huggingface.co/datasets/kurota0612/combined-structured-dataset.
