CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vericava /sft-tool-calling-structured-output-v1 vericava/sft-tool-calling-structured-output-v1 Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications. Includes contents in English as well as some Japanese. texttext-classification100K<n<1M2 likes167 downloads8mo agoHugging Face02BReil /Alpaca_Structuretexttext-generation10M<n<100M0 likes104 downloads2y agoHugging Face03wshuai190 /hotpotqa-structuredThis dataset is associated with the paper Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents. Official GitHub repository: https://github.com/ielab/skim-search-agent textquestion-answering10K<n<100K0 likes104 downloads2mo agoHugging Face04wshuai190 /browsecomp-plus-structured-fullThis dataset is associated with the paper Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents. Repository: ielab/skim-search-agent Citation @misc{wang2026sieve, title = {Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents}, author = {Wang, Shuai and Chen, Haodong and Yin, Yu and Zhuang, Shengyao and Koopman, Bevan and Zuccon, Guido}, year = {2026}, eprint =… See the full description on the dataset page: https://huggingface.co/datasets/wshuai190/browsecomp-plus-structured-full.question-answering0 likes90 downloads2mo agoHugging Face05emgena /omnimcp_browser_dom_structured_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.texttext-generationn<1K0 likes71 downloads6d agoHugging Face06bysismo /Turkish_5N1K_Tabanli_Sentetik_Veri_Seti_TR-Structured_Synthetic_Dataset 🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance). 🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish_5N1K_Tabanli_Sentetik_Veri_Seti_TR-Structured_Synthetic_Dataset.question-answering100K<n<1M0 likes40 downloads1mo agoHugging Face07chuckreynolds /wikimedia-enterprise-structured-contents-enwiki enwiki_namespace_0 Structured Contents snapshot of enwiki_namespace_0 from the Wikimedia Enterprise API, converted to Parquet. Source Upstream: Wikimedia Enterprise Structured Contents API Snapshot identifier: enwiki_namespace_0 Format at source: .tar.gz containing sharded .ndjson Shards in this release: 3 Processing Downloaded the snapshot tarball from the Wikimedia Enterprise API. Streamed each .ndjson shard through a normalization pass: JSON-encoded… See the full description on the dataset page: https://huggingface.co/datasets/chuckreynolds/wikimedia-enterprise-structured-contents-enwiki.texttext-generation100K<n<1M0 likes37 downloads5mo agoHugging Face08philipp-zettl /german-structured-output German Structured Output Dataset 🇩🇪 GDPR & EU AI Act compliant German dataset for training structured output capabilities in LLMs. Overview This dataset contains 4,521 examples across 7 task types for training language models to produce structured outputs (JSON, function calls, schema-following generation) from German text. It is the first dedicated German structured output dataset, filling a critical gap in the German NLP ecosystem. Key Features 🇩🇪… See the full description on the dataset page: https://huggingface.co/datasets/philipp-zettl/german-structured-output.texttext-generation1K<n<10K0 likes37 downloads5mo agoHugging Face09wshuai190 /musique-structuredThis dataset repository contains data associated with the paper Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents. Code and evaluation framework: ielab/skim-search-agent Citation @misc{wang2026sieve, title = {Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents}, author = {Wang, Shuai and Chen, Haodong and Yin, Yu and Zhuang, Shengyao and Koopman, Bevan and Zuccon, Guido}, year… See the full description on the dataset page: https://huggingface.co/datasets/wshuai190/musique-structured.textquestion-answering10K<n<100K0 likes36 downloads2mo agoHugging Face10xnileshtiwari /CBSE-Class-12th_2024_PYQs__structuredThis data set contains the CBSE Class 12 2024 papers in a structured format. The papers are annotated with topic and chapter names, and the figures are parsed and their paths annotated. tabularquestion-answeringn<1K2 likes35 downloads2y agoHugging Face11monetise /Structured-Scripture-for-AI ::PROJECT{Structured_Scripture_for_AI} ::PURPOSE{ENABLE(AI) → UNDERSTAND(Christian_theology) ∧ EXPLAIN(→ ∀ @HUMAN, ∀ culture, ∀ language, ∀ education_level) ∧ ZERO(friction)} ::TYPE{¬digitized_Bible ⇒ structured_encoding(three_layers)} ::ARCHITECTURE ::LAYER{text} WHAT(happened) — narrative ∧ events ∧ cause_effect ∧ speech ::LAYER{theology} WHAT(it_means) — within(Christian_doctrine) | logic ∧ paradox ∧ moral_principles ∧ emotion… See the full description on the dataset page: https://huggingface.co/datasets/monetise/Structured-Scripture-for-AI.text-generation0 likes33 downloads5mo agoHugging Face12mdonigian /synthetic-structured-output-dataset Synthetic Structured Output Dataset (SFT + DPO) Synthetic training corpus for structured-output model tuning. This package contains SFT and DPO data focused on JSON schema compliance, structured extraction, and function calling. Included files sft_synthetic_json.jsonl — SFT samples for schema-conditioned JSON generation sft_synthetic_extraction.jsonl — SFT samples for text-to-structured extraction dpo_structured_output.jsonl — DPO chosen/rejected pairs for structured… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/synthetic-structured-output-dataset.text-generation10K<n<100K0 likes28 downloads7mo agoHugging Face13williamTLmiller /bva-decisions-structured-sample2019Present BVA Structured Decisions (2019–2025) Structured, issue-level records extracted from U.S. Board of Veterans' Appeals (BVA) decisions — each decision parsed into its issues, conditions, outcomes, citations, and reasoning, with per-document provenance and completeness flags. Built for training and evaluating legal-AI models on veterans' disability adjudication. This is a 2900-decision sample, balanced across seven years (2019–2025, decisions/year), so it's representative of the… See the full description on the dataset page: https://huggingface.co/datasets/williamTLmiller/bva-decisions-structured-sample2019Present.tabulartext-classification1K<n<10K0 likes26 downloads3mo agoHugging Face14Voidreaper2026 /uk-benefit-forms-structured UK Benefit Forms Structured Dataset A structured dataset of 120 UK government benefit and legal forms, extracted and processed for use in AI-assisted form-filling applications. Built as part of the EasyClaimAI project. Why This Dataset Exists Millions of people in the UK struggle with complex government forms — benefit claims, legal applications, pension forms. The language is dense, the guidance is buried, and mistakes can cost people money or delay vital support. This… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/uk-benefit-forms-structured.textquestion-answering1K<n<10K1 likes24 downloads4mo agoHugging Face15Oliviety /french-court-decisions-structured French Court Decisions x Law Articles Version complete disponible Ce dataset est un sample gratuit de 100 decisions. La version complete inclut : 1000+ decisions enrichies (scalable a 10K+) Mise a jour hebdomadaire Taux d'enrichissement 96% Filtres par theme, periode, juridiction Export API disponible Formats : Parquet, JSONL, JSON Contact : KlarTools@outlook.fr Ce qui rend ce dataset unique Croisement jurisprudence x legislation : chaque… See the full description on the dataset page: https://huggingface.co/datasets/Oliviety/french-court-decisions-structured.tabulartext-classificationn<1K1 likes24 downloads3mo agoHugging Face16obadabaq /structured-uae-laws Dataset Card for structured-uae-laws This dataset is a collection of question & answers about the laws and regulations in the United Arab Emirates. It covers different areas of law like: economy and business family and community finance and banking industry and technical standardisation justice and juiciary, labour residency and leberal professions security and safety tax Dataset Sources Repository Base Dataset United Arab Emirates Legislations… See the full description on the dataset page: https://huggingface.co/datasets/obadabaq/structured-uae-laws.textquestion-answering1K<n<10K1 likes23 downloads2y agoHugging Face17Srinivasmec26 /Structured-Todo-Lists-for-Learning-and-Projects Academic Task Management Dataset Overview 100 structured todo lists for academic and personal organization. Culturally diverse with 70% Indian education context, 25% European scenarios, and 5% other Asian contexts. Dataset Structure { "input": "Task description", "output": { "type": "todo", "title": "List title", "category": "academic/personal/project", "items": [ {"task": "...", "done": false, "priority": "low/medium/high"} ] }… See the full description on the dataset page: https://huggingface.co/datasets/Srinivasmec26/Structured-Todo-Lists-for-Learning-and-Projects.texttext-classificationn<1K1 likes17 downloads1y agoHugging Face18lianghsun /tw-structured-law-articlegated Dataset Card for tw-structured-law-article tw-structured-law-article 是一個中華民國(台灣)結構化複合型法規條文之繁體中文 SFT 資料集,規模為數十萬筆(datasets.jsonl 約 2.6 GB)。相較於 tw-processed-law-article,本資料集著重於「結構化且多元化之法條排版」,涵蓋常見之法條編排格式,並處理包含 .pdf 附件內之條文內容,目的是讓 LLM 能全面學習更接近人類閱讀方式之法條結構。 Dataset Details Dataset Description 法律條文在實務上之呈現方式相當多樣:條文本身、條項款目之分層、表格式條文、附表、附圖、附件式條文等。傳統之法律語料多僅保留條文文字本身,丟失了這些結構化資訊,導致 LLM 雖能背誦條文,卻無法理解條文之層次結構與視覺排版。 本資料集每筆以 ShareGPT(messages)與 Alpaca(instruction / input /… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-structured-law-article.text-generation100K<n<1M2 likes16 downloads5mo agoHugging Face19VoeTheDon /testing-wiki-structured cywiki_namespace_0 Structured Contents snapshot of cywiki_namespace_0 from the Wikimedia Enterprise API, repackaged as Parquet with a pinned schema. The upstream Wikimedia Foundation dataset (wikimedia/structured-wikipedia) ships NDJSON which has known issues loading via datasets.load_dataset() — see discussions #5, #15, #16. This dataset is the same upstream content, normalised so load_dataset(...)works without specifying a Features override. Source Upstream: Wikimedia… See the full description on the dataset page: https://huggingface.co/datasets/VoeTheDon/testing-wiki-structured.texttext-generation10K<n<100K0 likes14 downloads5mo agoHugging Face20cfpark00 /StructuredORBenchtext-generation0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.