CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01theodi /ndl-core-structured-data NDL Core – Structured Data Overview NDL Core – Structured Data is a curated collection of structured UK public sector datasets, converted into Apache Parquet format for efficient analytics and machine learning workflows. This repository is part of the broader NDL Core Corpus, which combines both textual and structured data sourced from authoritative UK government and public sector platforms. Textual sources (e.g. GOV.UK, Hansard, legislation.gov.uk) are hosted separately… See the full description on the dataset page: https://huggingface.co/datasets/theodi/ndl-core-structured-data.100M<n<1B0 likes11k downloads8mo agoHugging Face02mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes178 downloads7mo agoHugging Face03u-10bei /structured_data_with_cot_dataset_512_v4 structured_data_with_cot_dataset このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。 データセットの概要 messages: OpenAIチャット形式 (system, user, assistant) metadata: format, complexity, schema, estimated_tokens サポートされるデータ形式 JSON, XML, YAML, TOML, CSV 生成方法 Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。 text1K<n<10K1 likes77 downloads8mo agoHugging Face04Z-Edgar /Agent-IPI-Structured-Interaction-Datasets-v2 Adversarial Dataset for LLM Instruction Hijacking / Tool-Calling Attacks This directory contains the processed training and test datasets for evaluating and training defenses against prompt injection / instruction hijacking attacks in LLM tool-calling scenarios. The dataset includes both JSON and XML formatted inputs, with three difficulty buckets: no_attack: clean (benign) examples easy: value-level or structure-level single attacks hard: structure-destroying attacks or combined… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets-v2.text-generation100K<n<1M1 likes68 downloads8mo agoHugging Face05Z-Edgar /Agent-IPI-Structured-Interaction-Datasets Dataset Card for Indirect Prompt Injection in Agent Structured Interaction Datasets Dataset Summary This dataset contains 470,000 QA pairs designed to study indirect prompt injection in agent-structured interactions. It is split into a training set (80%) and a test set (20%). The dataset is evenly divided into 50% clean-clean QA pairs (no prompt injection) and 50% clean-injected QA pairs (containing prompt injection). The task is to detect and remove prompt injection… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets.text100K<n<1M1 likes63 downloads9mo agoHugging Face06farosio /brand-structured-data-reference Brand Structured Data Reference v1.0 This reference maps common public brand facts to structured data concepts that can help people, search engines, and AI systems understand a brand more clearly. It is intended for independent brands, small businesses, founder-led companies, service providers, local businesses, and early-stage products that need a clearer public identity online. This is not a ranking guide and it does not guarantee search visibility, rich results, AI… See the full description on the dataset page: https://huggingface.co/datasets/farosio/brand-structured-data-reference.textn<1K0 likes57 downloads4mo agoHugging Face07u-10bei /structured_data_with_cot_dataset_512_v5 structured_data_with_cot_dataset このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。 データセットの概要 messages: OpenAIチャット形式 (system, user, assistant) metadata: format, complexity, schema, estimated_tokens サポートされるデータ形式 JSON, XML, YAML, TOML, CSV 生成方法 Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。 v5アップデート:ランダムなスキーマ構造の生成と、最小化(minified)/ソート(sorted)の制約を追加。 text1K<n<10K3 likes55 downloads8mo agoHugging Face08Nexdata-kr /1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample Description 한국어 시험 문제 구조화 분석·가공 데이터로, 약 150만 개의 시험 문제를 포함하고 있습니다. 문제 유형, 문제, 정답, 해설 등의 정보를 포함하며, 과목은 [초등학교] 국어, 수학, 영어, 사회, 과학; [중학교] 국어, 영어, 수학, 과학, 사회; [고등학교] 국어, 영어, 수학, 물리, 화학, 생물, 역사, 지리로 구성되어 있습니다. 문제 유형에는 객관식, 빈칸 채우기, 참·거짓 문제, 단답형 문제 등이 포함됩니다. 본 데이터셋은 대규모 교과 지식 강화 및 학습 데이터 구축 등의 작업에 활용할 수 있습니다. 자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/1634?source=Hf.kr Specifications Data content 한국어 K12 시험 문제 Amount 약 150만 개의… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.textn<1K0 likes49 downloads15d agoHugging Face09u-10bei /structured_data_with_cot_dataset_512_v3 structured_data_with_cot_dataset このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。 データセットの概要 messages: OpenAIチャット形式 (system, user, assistant) metadata: format, complexity, schema, estimated_tokens サポートされるデータ形式 JSON, XML, YAML, TOML, CSV 生成方法 Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。 text1K<n<10K0 likes44 downloads8mo agoHugging Face10bysismo /Turkish_5N1K_Tabanli_Sentetik_Veri_Seti_TR-Structured_Synthetic_Dataset 🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance). 🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish_5N1K_Tabanli_Sentetik_Veri_Seti_TR-Structured_Synthetic_Dataset.question-answering100K<n<1M0 likes40 downloads1mo agoHugging Face11AdamGoldman /TCM-Herbal-Dataset-Structured-Sample Structured TCM Herbal Dataset (50-Herb Sample) Dataset Description This dataset is a high-precision, structured collection of Traditional Chinese Medicine (TCM) herbs. It bridges the gap between classical herbal knowledge and modern data engineering. Total Sample Records: 50 Master Database Size: 2,170+ records Format: JSON Fields: English/Latin/Chinese Names, Nature, Taste, Meridians, Toxicity, Chemical Compounds, and Pharmacological Mechanisms. Use Cases… See the full description on the dataset page: https://huggingface.co/datasets/AdamGoldman/TCM-Herbal-Dataset-Structured-Sample.n<1K0 likes39 downloads6mo agoHugging Face12Nexdata-AI /1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample Description Korean Test Questions Structured Analysis Processing Data, around 1.5 million questions, contains question types, questions, answers, explanations, etc..For subjects, include [Primary School] Korean, Mathematics, English, Social Studies, Science; [Middle School] Korean, English, Mathematics, Science, Social Studies; [High School] Korean, English, Mathematics, Physics, Chemistry, Biology, History, Geography; question Types indlude single-choice question, fill-in… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.textn<1K0 likes35 downloads1mo agoHugging Face13Kauhiro /structured_data_with_cot_dataset_512_collected_from_v2v4v5_unique Dataset creation This dataset was merged by collecting conversion type records only from v2, v4, and v5 of u-10bei/structured_data_with_cot_512. and removed duplicated records. Total records before removing duplicates: 5530 (Sources: 1485 from v2, 2037 from v4, 2008 from v5) Type: conversion Final records after removing duplicated records: 5451 (79 messages duplicated) Collection method Records with its type as conversion were collected. License The license… See the full description on the dataset page: https://huggingface.co/datasets/Kauhiro/structured_data_with_cot_dataset_512_collected_from_v2v4v5_unique.text1K<n<10K0 likes32 downloads7mo agoHugging Face14u-10bei /structured_data_with_cot_dataset_512_v2 structured_data_with_cot_dataset このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。このデータセットの目的は、自然言語プロンプトに基づいて構造化データを生成できるモデルに対し、CoTを活用してより良い推論と出力品質を達成するための高品質な学習例を提供することです。 データセットの概要 データセットの各エントリには、以下のものが含まれます。 messages: OpenAIチャット形式に従ったメッセージ辞書のリストで、以下で構成されます。 systemメッセージ: 特定のデータ形式の専門家としてのコンテキストを設定します。 userメッセージ: 特定のスキーマと形式で構造化データの生成を要求するプロンプトです。多様なプロンプトテンプレートが使用されています。 assistantメッセージ: 生成された構造化データに続く思考連鎖(CoT)推論が含まれます。… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/structured_data_with_cot_dataset_512_v2.text1K<n<10K1 likes31 downloads9mo agoHugging Face15mdonigian /synthetic-structured-output-dataset Synthetic Structured Output Dataset (SFT + DPO) Synthetic training corpus for structured-output model tuning. This package contains SFT and DPO data focused on JSON schema compliance, structured extraction, and function calling. Included files sft_synthetic_json.jsonl — SFT samples for schema-conditioned JSON generation sft_synthetic_extraction.jsonl — SFT samples for text-to-structured extraction dpo_structured_output.jsonl — DPO chosen/rejected pairs for structured… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/synthetic-structured-output-dataset.text-generation10K<n<100K0 likes28 downloads7mo agoHugging Face16Kauhiro /structured_data_with_cot_dataset_512_collected_from_v2v4v5 Merged dataset from u-10bei/structured_data_with_cot_dataset_512_v2, v4, and v5 Collected conversion type records from v2, v4, and v5 of structured_data_with_cot_512. Total records: 5530 (Sources: 1485 from v2, 2037 from v4, 2008 from v5) Type: conversion Collection method Records with its type as conversion were collected. License The license of this dataset inherits from the license of the original dataset.… See the full description on the dataset page: https://huggingface.co/datasets/Kauhiro/structured_data_with_cot_dataset_512_collected_from_v2v4v5.text1K<n<10K0 likes26 downloads7mo agoHugging Face17sasa5555 /structured_data_with_cot_dataset_512_crean structured_data_with_cot_dataset このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。このデータセットの目的は、自然言語プロンプトに基づいて構造化データを生成できるモデルに対し、CoTを活用してより良い推論と出力品質を達成するための高品質な学習例を提供することです。 データセットの概要 データセットの各エントリには、以下のものが含まれます。 messages: OpenAIチャット形式に従ったメッセージ辞書のリストで、以下で構成されます。 systemメッセージ: 特定のデータ形式の専門家としてのコンテキストを設定します。 userメッセージ: 特定のスキーマと形式で構造化データの生成を要求するプロンプトです。多様なプロンプトテンプレートが使用されています。 assistantメッセージ: 生成された構造化データに続く思考連鎖(CoT)推論が含まれます。… See the full description on the dataset page: https://huggingface.co/datasets/sasa5555/structured_data_with_cot_dataset_512_crean.text1K<n<10K0 likes23 downloads8mo agoHugging Face18takami2022 /structured_data_merged_v2v5_0222 Dataset Card for structured_data_merged_v2v5_0222 Dataset Details Dataset Description structured_data_merged_v2v5_0222 is a dataset for Supervised Fine-Tuning (SFT) focused on structured data format conversion tasks — specifically, interconversion among JSON, XML, YAML, TOML, and CSV. It was created by deduplicating and merging the following two existing datasets: u-10bei/structured_data_with_cot_dataset_512_v2 (train split only)… See the full description on the dataset page: https://huggingface.co/datasets/takami2022/structured_data_merged_v2v5_0222.texttext-generation1K<n<10K0 likes23 downloads7mo agoHugging Face19ToshiyukiNH /structured_data_with_cot_dataset_512_v2_filtered_3structured_data_with_cot_dataset_512_v2_filtered_3 This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2. Usage from datasets import load_dataset # From local data dataset = load_dataset( 'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_3", split='train' ) print(dataset[0]) How to generate this dataset from the base one from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_3.textn<1K0 likes22 downloads7mo agoHugging Face20sekigh /10bei_structured_data_with_cot_dataset_512_v2_constraints_added_ver2text1K<n<10K0 likes21 downloads8mo agoHugging Face21u-10bei /structured_data_with_cot_dataset_v2text1K<n<10K0 likes19 downloads9mo agoHugging Face22warheads /structured_data_with_cot_dataset_512_sampling_mixThis is the concatenated datasets of u-10bei/structured_data_with_cot_dataset_512_v2 (train 5000 records) u-10bei/structured_data_with_cot_dataset_512_v4 (train 3000 records) u-10bei/structured_data_with_cot_dataset_512_v5 (train 2000 records) with sampling. text10K<n<100K0 likes19 downloads7mo agoHugging Face23ariefansclub /humanoid-domestic-task-structured-dataset-v2 Humanoid Domestic Task Structured Dataset Overview This dataset contains structured human instructions for basic household assistance scenarios. It is designed to help humanoid agents interpret natural language commands and convert them into clear executable task representations. The dataset focuses on simple real-world domestic tasks that reduce human workload and improve everyday living environments. Key Features Natural human-written instructions Structured… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/humanoid-domestic-task-structured-dataset-v2.textn<1K0 likes18 downloads7mo agoHugging Face24u-10bei /structured_data_with_cot_dataset_512text1K<n<10K0 likes17 downloads9mo agoHugging Face25naoyasss /structured_data_with_cot_dataset_512_v2_r0.6 structured_data_with_cot_dataset このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。このデータセットの目的は、自然言語プロンプトに基づいて構造化データを生成できるモデルに対し、CoTを活用してより良い推論と出力品質を達成するための高品質な学習例を提供することです。 データセットの概要 データセットの各エントリには、以下のものが含まれます。 messages: OpenAIチャット形式に従ったメッセージ辞書のリストで、以下で構成されます。 systemメッセージ: 特定のデータ形式の専門家としてのコンテキストを設定します。 userメッセージ: 特定のスキーマと形式で構造化データの生成を要求するプロンプトです。多様なプロンプトテンプレートが使用されています。 assistantメッセージ: 生成された構造化データに続く思考連鎖(CoT)推論が含まれます。… See the full description on the dataset page: https://huggingface.co/datasets/naoyasss/structured_data_with_cot_dataset_512_v2_r0.6.text1K<n<10K0 likes17 downloads7mo agoHugging Face26kurota0612 /combined-structured-dataset Combined Structured Output Dataset このデータセットは、構造化出力生成タスクのための統合データセットです。 Dataset Details Total Samples: 3,425 Source Datasets: u-10bei/structured_data_with_cot_dataset_v2 (2,500 samples) daichira/structured-3k-mix-sft (3,000 samples) Preprocessing: Format normalization to 3-turn (system/user/assistant) Deduplication by user content Quality filtering Format Distribution JSON: 685 (20.0%) YAML: 485 (14.2%) TOML: 685 (20.0%) XML: 885 (25.8%) CSV: 685… See the full description on the dataset page: https://huggingface.co/datasets/kurota0612/combined-structured-dataset.texttext-generation1K<n<10K0 likes15 downloads7mo agoHugging Face27ToshiyukiNH /structured_data_with_cot_dataset_512_v2_filtered_2structured_data_with_cot_dataset_512_v2_filtered_2 This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2. Usage from datasets import load_dataset # From local data dataset = load_dataset( 'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_2", split='train' ) print(dataset[0]) How to generate this dataset from the base one from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_2.textn<1K0 likes15 downloads7mo agoHugging Face28champ7 /w3_data_structured0 likes15 downloads4mo agoHugging Face29warheads /structured_data_with_cot_dataset_512_concatenate_v2_v4_v5This is the concatenated datasets of u-10bei/structured_data_with_cot_dataset_512_v2 (train 3933 records) u-10bei/structured_data_with_cot_dataset_512_v4 (train 4608 records) u-10bei/structured_data_with_cot_dataset_512_v5 (train 4547 records) without sampling. text10K<n<100K0 likes14 downloads7mo agoHugging Face30ToshiyukiNH /structured_data_with_cot_dataset_512_v2_filtered_1structured_data_with_cot_dataset_512_v2_filtered_1 This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2. Usage from datasets import load_dataset # From local data dataset = load_dataset( 'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_1", split='train' ) print(dataset[0]) How to generate this dataset from the base one from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_1.textn<1K0 likes14 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.