CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wikimedia /structured-wikipedia Dataset Card for Wikimedia Structured Wikipedia Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta-Wiki Discussion Dataset Summary Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/structured-wikipedia.text10M<n<100M395 likes15k downloads4mo agoHugging Face02nvidia /Nemotron-RL-Instruction-Following-Structured-Outputs-v2 Dataset Description: Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema. Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.texttext-generation10K<n<100K7 likes1.2k downloads4mo agoHugging Face03Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes1k downloads2y agoHugging Face04open-athena /a3-rl-laion_nemotron-gym-instruction-following-structuredtext10K<n<100K0 likes515 downloads4mo agoHugging Face05open-athena /nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes313 downloads2mo agoHugging Face06domofon /structured-cpt Structured CPT - JSON + SQL pretrain documents SmolLM2-1.7B continued-pretraining shard of structured documents. Each document is a <task> / <input> / <output> block whose <output> is a canonical JSON object, terminated by the SmolLM2 end-of-text token ``. Sources: source description rows shards repeat sql_bmc2 b-mc2 sql-create-context -> JSON (4 keys, stub explanation) 392,885 1 5 sql_gretelai gretelai synthetic_text_to_sql -> JSON (4 keys) 529,255 1 5… See the full description on the dataset page: https://huggingface.co/datasets/domofon/structured-cpt.texttext-generation1M<n<10M0 likes308 downloads18d agoHugging Face07GoktugD /turkish-structured-summarization-1.5m Turkish Structured Summarization 1.5M v2 Üç cümlelik kurgusal operasyon kayıtları ve kısa Türkçe özetleri. Doğrulanmış boyut Train: 1,470,000 Validation: 15,000 Test: 15,000 Toplam: 1,500,000 Ana görev sütunları: id, document, summary, domain Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-structured-summarization-1.5m.textsummarization1M<n<10M0 likes269 downloads1mo agoHugging Face08safelegalaidata /eu-ai-act-structured EU AI Act, structured Regulation (EU) 2024/1689 (the Artificial Intelligence Act) as tables: every article, recital, annex and definition, 677 obligations coded by actor, risk tier, application date and penalty basis, plus milestones, national competent authorities and fine tiers. Built 2026-09-08 by SafeLegalAI (Cognesio LLP) from the official English texts served by the Publications Office of the European Union (Cellar): the consolidated text as of 27 July 2026 (CELEX… See the full description on the dataset page: https://huggingface.co/datasets/safelegalaidata/eu-ai-act-structured.tabular1K<n<10K1 likes252 downloads16d agoHugging Face09Aregay01 /structured-wikipedia Dataset Card for Wikimedia Structured Wikipedia Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta-Wiki Discussion Dataset Summary Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/Aregay01/structured-wikipedia.text10M<n<100M0 likes224 downloads4mo agoHugging Face10open-athena /nemotron-gym-instruction-following-structured-minimax-m27-131k-tracestext1K<n<10K0 likes196 downloads4mo agoHugging Face11Lots-of-LoRAs /task210_logic2text_structured_text_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task210_logic2text_structured_text_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task210_logic2text_structured_text_generation.texttext-generation1K<n<10K0 likes187 downloads2y agoHugging Face12JoyboyBrian /structured_imagesimage100K<n<1M1 likes178 downloads2y agoHugging Face13sachithgunasekara /phased-self-discover-mistral-structured-5-shot-bbh-evaltext1K<n<10K0 likes160 downloads2y agoHugging Face14dhruveshpatel /openclassgen-structured-v1 OpenClassGen Structured v1 Derived from mrahman2025/OpenClassGen (Rahman et al. 2025, arXiv:2504.15564). License: CC BY 2.0 (same as upstream). Keep repository_name and file_path when redistributing. Underlying GitHub repos may carry additional software licenses. gold_code is upstream human_written_code. We add parsed fields, body-span indices, and a Variant-3 prompt/target pair (v3_prompt_text / v3_target_text). No unit tests. Splits are repository-disjoint (train /… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/openclassgen-structured-v1.tabulartext-generation100K<n<1M0 likes138 downloads2mo agoHugging Face15frangelbarrera /cyber-evidence-kev-structured Cyber Security Evidence Dataset — CISA KEV Structured CC0 Layer This configuration is the structured CISA Known Exploited Vulnerabilities (KEV) layer of the broader Cyber Security Evidence Dataset project. It contains 1,687 deterministic records generated from the official CISA KEV database snapshot. What is included The records contain the official KEV database fields: CVE identifier, vendor/project, product, vulnerability name, short description, dates… See the full description on the dataset page: https://huggingface.co/datasets/frangelbarrera/cyber-evidence-kev-structured.texttext-classification1K<n<10K1 likes134 downloads21d agoHugging Face16yasalma /tt-structured-contentgated Dataset Summary This dataset contains structured textual content in Markdown format extracted from Tatar-language documents, originally in EPUB and PDF formats. The documents include books and other long-form content with rich formatting. The dataset is intended to provide clean, structured, and semantically meaningful content to support natural language processing tasks, content modeling, and research in Tatar language technologies. The extracted Markdown preserves key… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tt-structured-content.text10K<n<100K1 likes113 downloads7mo agoHugging Face17NewCarbon37 /structured-wikipedia Dataset Card for Wikimedia Structured Wikipedia Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta-Wiki Discussion Dataset Summary Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/NewCarbon37/structured-wikipedia.text10M<n<100M0 likes111 downloads3mo agoHugging Face18Lots-of-LoRAs /task128_scan_structured_text_generation_command_action_short Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.texttext-generation1K<n<10K0 likes107 downloads2y agoHugging Face19beyoru /Toucan-1.5M-structured-Qwentext100K<n<1M0 likes104 downloads1y agoHugging Face20narendra747 /structured-vitalsimage10K<n<100K0 likes90 downloads5mo agoHugging Face21morizon /TCGA_Reports_ja_structured-filtered 🧠 TCGA 日本語翻訳・構造化データセット このデータセットは、The Cancer Genome Atlas (TCGA) により公開された英語の病理報告書をもとに、大規模言語モデル(LLM)を用いて 日本語翻訳 および 情報抽出による構造化 を行ったものです。 📘 概要 原データ: Mendeley Data — TCGA Pathology Reports (Version 1)https://data.mendeley.com/datasets/hyg5xkznpx/1 本データセットは上記を基にし、以下の2種類の加工を行っています。 カラム名 内容 生成方法 question_en 英語の病理報告書原文 TCGA オリジナル question_ja 英語報告書の日本語訳 LLMによる翻訳 answer_ja 構造化データ(JSON形式) LLMによる情報抽出 🧾 データ構成… See the full description on the dataset page: https://huggingface.co/datasets/morizon/TCGA_Reports_ja_structured-filtered.textn<1K0 likes84 downloads11mo agoHugging Face22laion /terminal_bench_2_a3_rl_laion_nemotron_gym_instruction_following_structured_75_8B_25b4cef3dtext1K<n<10K0 likes82 downloads25d agoHugging Face23Lots-of-LoRAs /task1566_propara_structured_text_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1566_propara_structured_text_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1566_propara_structured_text_generation.texttext-generationn<1K0 likes78 downloads2y agoHugging Face24Lots-of-LoRAs /task130_scan_structured_text_generation_command_action_long Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task130_scan_structured_text_generation_command_action_long Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task130_scan_structured_text_generation_command_action_long.texttext-generation1K<n<10K0 likes76 downloads2y agoHugging Face25emgena /omnimcp_browser_dom_structured_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.texttext-generationn<1K0 likes74 downloads7d agoHugging Face26Neooooo /structured_paper_summarization structured_paper_summarization A 151 k‑example dataset of chat‐style prompt → structured abstract pairs, built from ~19 000 research papers across business, management, information‑systems and social‑science domains. Each example shows the full paper (body text) being summarised into a five‑section Emerald‑style structured abstract (Purpose, Design/methodology/approach, Findings, Practical implications, Originality/value). Why this dataset? Large‑language models… See the full description on the dataset page: https://huggingface.co/datasets/Neooooo/structured_paper_summarization.text100K<n<1M0 likes60 downloads1y agoHugging Face27laion /dev_set_v2_a3_rl_laion_nemotron_gym_instruction_following_structured_75_8B_20260825_181735text1K<n<10K0 likes59 downloads28d agoHugging Face28projecte-aina /CaSSA-catalan-structured-sentiment-analysis Dataset Card for CaSSA, the Catalan Structured Sentiment Analysis dataset Dataset Summary The CaSSA dataset is a corpus of 6,400 reviews and forum messages annotated with polar expressions. Each piece of text is annotated with all the expressions of polarity that it contains. For each polar expression, we annotated the expression itself, the target (the object of the expression), and the source (the subject expressing the sentiment). 25,453 polar expressions have been… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CaSSA-catalan-structured-sentiment-analysis.texttext-classification1K<n<10K3 likes57 downloads2y agoHugging Face29surucu35 /structured-wikipedia Dataset Card for Wikimedia Structured Wikipedia Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta-Wiki Discussion Dataset Summary Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/surucu35/structured-wikipedia.text10M<n<100M1 likes54 downloads3mo agoHugging Face30seeaman /economic-index-structured Economic Index - Structured & Cleaned Dataset This dataset is a cleaned, structured version of the Anthropic Economic Index, organized for easy integration with persona-based scenario generation pipelines. Dataset Description The Anthropic Economic Index tracks how people use Claude AI for work-related tasks. This structured version extracts and organizes the key information into easy-to-use tables. Original Data Period: August 4-11, 2025Source: Anthropic Economic… See the full description on the dataset page: https://huggingface.co/datasets/seeaman/economic-index-structured.tabular1K<n<10K0 likes50 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.