CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wikimedia /structured-wikipedia Dataset Card for Wikimedia Structured Wikipedia Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta-Wiki Discussion Dataset Summary Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/structured-wikipedia.text10M<n<100M394 likes14k downloads4mo agoHugging Face02theodi /ndl-core-structured-data NDL Core – Structured Data Overview NDL Core – Structured Data is a curated collection of structured UK public sector datasets, converted into Apache Parquet format for efficient analytics and machine learning workflows. This repository is part of the broader NDL Core Corpus, which combines both textual and structured data sourced from authoritative UK government and public sector platforms. Textual sources (e.g. GOV.UK, Hansard, legislation.gov.uk) are hosted separately… See the full description on the dataset page: https://huggingface.co/datasets/theodi/ndl-core-structured-data.100M<n<1B0 likes11k downloads8mo agoHugging Face03Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes1.4k downloads2y agoHugging Face04nvidia /Nemotron-RL-Instruction-Following-Structured-Outputs-v2 Dataset Description: Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema. Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.texttext-generation10K<n<100K7 likes1.2k downloads4mo agoHugging Face05open-athena /a3-rl-laion_nemotron-gym-instruction-following-structuredtext10K<n<100K0 likes1.1k downloads4mo agoHugging Face06anonymous-structured-agent /structured-file-audit-benchmark Paper Data Release This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them. Contents datasets/ Benchmark data and per-task manifests for the three paper-facing splits. datasets/sc_flat/data SC-Flat is derived from DaBench, augmented with a replayable perturbation injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.texttable-question-answering1 likes658 downloads2mo agoHugging Face07nvidia /Nemotron-RL-instruction_following-structured_outputs Dataset Description: The Nemotron-RL-instruction_following-structured_outputs dataset tests the ability of the model to follow output formatting instructions under schema constraints under the JSON format. Each problem consists of three components: The document, output formatting Instruction (Schema), and question. The dataset varies the difficulty of each problem by varying the location of instructions, the comprehensiveness of instructions, the complexity of the schema, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-instruction_following-structured_outputs.text1K<n<10K40 likes563 downloads8mo agoHugging Face08helioom /3dvlm-structured3d_subset Structured3D Subset (3DVLM) A small, fast-to-download slice of the Structured3D synthetic indoor dataset, converted to a uniform posed-RGB-D format for quick model test-runs. This is a subset: 100 scenes (randomly sampled, seed 0) from collection 00, using the pre-rendered full (furnished) perspective views. Across the 100 scenes there are 2,198 frames (3–49 per scene). These are photorealistic synthetic renders with perfect dense ground-truth depth and exact camera poses — no… See the full description on the dataset page: https://huggingface.co/datasets/helioom/3dvlm-structured3d_subset.3ddepth-estimation1K<n<10K0 likes515 downloads3mo agoHugging Face09terminusresearch /pseudo-camera-10k-structured-json pseudo-camera-10k, structured JSON captions The 9,997 training images from bghira/pseudo-camera-10k, recaptioned into the structured JSON caption schema that Ideogram 4 consumes. The images are unchanged: free photographs from world class photographers, Lanczos-resized so the shorter edge is 1024px, nothing upsampled. The original dataset carries short CogVLM prose captions. This one replaces them with one JSON object per image describing the scene at three levels: an overall… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/pseudo-camera-10k-structured-json.imagetext-to-image10K<n<100K0 likes332 downloads12d agoHugging Face10open-athena /nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes302 downloads2mo agoHugging Face11Aregay01 /structured-wikipedia Dataset Card for Wikimedia Structured Wikipedia Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta-Wiki Discussion Dataset Summary Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/Aregay01/structured-wikipedia.text10M<n<100M0 likes298 downloads4mo agoHugging Face12domofon /structured-cpt Structured CPT - JSON + SQL pretrain documents SmolLM2-1.7B continued-pretraining shard of structured documents. Each document is a <task> / <input> / <output> block whose <output> is a canonical JSON object, terminated by the SmolLM2 end-of-text token ``. Sources: source description rows shards repeat sql_bmc2 b-mc2 sql-create-context -> JSON (4 keys, stub explanation) 392,885 1 5 sql_gretelai gretelai synthetic_text_to_sql -> JSON (4 keys) 529,255 1 5… See the full description on the dataset page: https://huggingface.co/datasets/domofon/structured-cpt.texttext-generation1M<n<10M0 likes298 downloads16d agoHugging Face13open-athena /nemotron-gym-instruction-following-structured-minimax-m27-131k-tracestext1K<n<10K0 likes245 downloads4mo agoHugging Face14GoktugD /turkish-structured-summarization-1.5m Turkish Structured Summarization 1.5M v2 Üç cümlelik kurgusal operasyon kayıtları ve kısa Türkçe özetleri. Doğrulanmış boyut Train: 1,470,000 Validation: 15,000 Test: 15,000 Toplam: 1,500,000 Ana görev sütunları: id, document, summary, domain Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-structured-summarization-1.5m.textsummarization1M<n<10M0 likes236 downloads1mo agoHugging Face15RayYoh /structured3d_gaussiancross0 likes234 downloads7mo agoHugging Face16safelegalaidata /eu-ai-act-structured EU AI Act, structured Regulation (EU) 2024/1689 (the Artificial Intelligence Act) as tables: every article, recital, annex and definition, 677 obligations coded by actor, risk tier, application date and penalty basis, plus milestones, national competent authorities and fine tiers. Built 2026-09-08 by SafeLegalAI (Cognesio LLP) from the official English texts served by the Publications Office of the European Union (Cellar): the consolidated text as of 27 July 2026 (CELEX… See the full description on the dataset page: https://huggingface.co/datasets/safelegalaidata/eu-ai-act-structured.tabular1K<n<10K1 likes232 downloads14d agoHugging Face17LLM4OR /StructuredOR0 likes228 downloads1y agoHugging Face18Lots-of-LoRAs /task210_logic2text_structured_text_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task210_logic2text_structured_text_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task210_logic2text_structured_text_generation.texttext-generation1K<n<10K0 likes206 downloads2y agoHugging Face19sillamainproject /augmented-skin-images-50k-structuredimage10K<n<100K0 likes187 downloads1y agoHugging Face20Rain729 /Structured3D0 likes185 downloads3mo agoHugging Face21JoyboyBrian /structured_imagesimage100K<n<1M1 likes182 downloads2y agoHugging Face22Pointcept /concerto_structured3d_compressed0 likes177 downloads1y agoHugging Face23KevinHuang /Structured3DStructured3D panorama dataset with BLIP3 text caption Missing data: Structured3D/scene_00212/2D_rendering/494/panorama/full/text_rawlight_blip3.txt Structured3D/scene_00411/2D_rendering/918447/panorama/full/text_rawlight_blip3.txt Structured3D/scene_00609/2D_rendering/159/panorama/full/text_rawlight_blip3.txt Structured3D/scene_01209/2D_rendering/4566/panorama/full/text_rawlight_blip3.txt Structured3D/scene_01210/2D_rendering/204/panorama/full/text_rawlight_blip3.txt… See the full description on the dataset page: https://huggingface.co/datasets/KevinHuang/Structured3D.1 likes175 downloads8mo agoHugging Face24sxiong /MLR_structured_trajectory Reasoning Trajectories with Step-Level Annotations This dataset contains structured reasoning trajectories introduced in (ICLR 2026) Enhancing Language Model Reasoning with Structured Multi-Level Modeling. Compared with the full-trajectory release, this dataset is a cleaned and segmented version designed for research on hierarchical reasoning, trajectory supervision, and multi-step policy training. Each example contains the original prompt and response fields together with a… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/MLR_structured_trajectory.tabular10K<n<100K1 likes169 downloads3mo agoHugging Face25gmahia /philosophy-classics-structured Classical Decision Frameworks — Philosophy Dataset Structured public domain philosophical texts focused on decision-making, leadership, and organizational ethics. All content is in the public domain. Content Works from classical philosophy structured for AI analysis: Stoic decision principles (Marcus Aurelius, Epictetus, Seneca) Political philosophy (Machiavelli, Aristotle) Virtue ethics (Aristotle, Plato) Sources All works published before 1928… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/philosophy-classics-structured.texttext-classificationn<1K0 likes164 downloads2mo agoHugging Face26sachithgunasekara /phased-self-discover-mistral-structured-5-shot-bbh-evaltext1K<n<10K0 likes158 downloads2y agoHugging Face27vericava /sft-tool-calling-structured-output-v1 vericava/sft-tool-calling-structured-output-v1 Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications. Includes contents in English as well as some Japanese. texttext-classification100K<n<1M2 likes147 downloads8mo agoHugging Face28Pointcept /structured3d-compressedgated100B<n<1T2 likes146 downloads2y agoHugging Face29strongminsu /ko-en-structured-translations Korean–English Multistyle Parallel Corpus 한국어 사용자에게 익숙한 표현 기반의 다도메인·다문체 한–영 병렬 코퍼스 소개(Introduction) 저는 머신러닝, 인공지능 수업을 진행하는 강사입니다.Seq2Seq, Attention, Transformer 등 자연어처리(NLP) 수업을 진행하며한국 학습자에게 자연스럽고 익숙한 한–영 번역 데이터셋의 부족을 경험했습니다. 기존 공개 데이터셋은 도메인 다양성이 부족하거나 문체가 한국 사용자에게 자연스럽지 않거나 전반적으로 문장의 퀄리티가 매우 부족하여 학습한 번역 모델의 실제 성능이 기대만큼 나오지 않는 문제가 있었습니다. 이 문제를 해결하기 위해, 딥러닝 강사로서 langchain을 사용하여 직접 고품질 병렬 데이터를 자동으로 생성·정제하여 구성한 데이터셋입니다. 한국어 사용자에게 익숙한 표현을 중심으로 다양한 문체, 문장 구조를… See the full description on the dataset page: https://huggingface.co/datasets/strongminsu/ko-en-structured-translations.texttranslation1K<n<10K17 likes136 downloads10mo agoHugging Face30yasalma /tt-structured-contentgated Dataset Summary This dataset contains structured textual content in Markdown format extracted from Tatar-language documents, originally in EPUB and PDF formats. The documents include books and other long-form content with rich formatting. The dataset is intended to provide clean, structured, and semantically meaningful content to support natural language processing tasks, content modeling, and research in Tatar language technologies. The extracted Markdown preserves key… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tt-structured-content.text10K<n<100K1 likes133 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.