structured_data
structured-data-classificationstructured-data-classification-grn-vsnyour-lora-repo-structured_data_with_cot_dataset_512_v2-0214qwen3-4b-structured-output-lora_data06qwen3-4b-structured-lora-u-10bei_structured_data_with_cot_dataset_512_v4-sig3mb12kr090-20260302qwen3-4b-structured-output-lora_data02your-lora-repo-structured_data_with_cot_dataset_512_v2-0218qwen3-4b-structured-output-lora_data_merged_v2v5_0222_Ver1-merged
ndl-core-structured-data
NDL Core – Structured Data
Overview
NDL Core – Structured Data is a curated collection of structured UK public sector datasets, converted into Apache Parquet format for efficient analytics and machine learning workflows.
This repository is part of the broader NDL Core Corpus, which combines both textual and structured data sourced from authoritative UK government and public sector platforms.
Textual sources (e.g. GOV.UK, Hansard, legislation.gov.uk) are hosted separately… See the full description on the dataset page: https://huggingface.co/datasets/theodi/ndl-core-structured-data.full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.structured_data_with_cot_dataset_512_v4
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。
データセットの概要
messages: OpenAIチャット形式 (system, user, assistant)
metadata: format, complexity, schema, estimated_tokens
サポートされるデータ形式
JSON, XML, YAML, TOML, CSV
生成方法
Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。
Agent-IPI-Structured-Interaction-Datasets-v2
Adversarial Dataset for LLM Instruction Hijacking / Tool-Calling Attacks
This directory contains the processed training and test datasets for evaluating and training defenses against prompt injection / instruction hijacking attacks in LLM tool-calling scenarios.
The dataset includes both JSON and XML formatted inputs, with three difficulty buckets:
no_attack: clean (benign) examples
easy: value-level or structure-level single attacks
hard: structure-destroying attacks or combined… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets-v2.Agent-IPI-Structured-Interaction-Datasets
Dataset Card for Indirect Prompt Injection in Agent Structured Interaction Datasets
Dataset Summary
This dataset contains 470,000 QA pairs designed to study indirect prompt injection in agent-structured interactions. It is split into a training set (80%) and a test set (20%). The dataset is evenly divided into 50% clean-clean QA pairs (no prompt injection) and 50% clean-injected QA pairs (containing prompt injection). The task is to detect and remove prompt injection… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets.brand-structured-data-reference
Brand Structured Data Reference v1.0
This reference maps common public brand facts to structured data concepts that can help people, search engines, and AI systems understand a brand more clearly.
It is intended for independent brands, small businesses, founder-led companies, service providers, local businesses, and early-stage products that need a clearer public identity online.
This is not a ranking guide and it does not guarantee search visibility, rich results, AI… See the full description on the dataset page: https://huggingface.co/datasets/farosio/brand-structured-data-reference.
