CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01arjhinety /OpenGrad-ToolPolicy-Canonical-v2-M0-snapshot This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture. What this release is OpenGrad ToolPolicy Canonical v2 is a provenance-preserving, model-independent normalization of public tool-use and function-calling datasets. It was built to test one hypothesis with a measurement attached: that the tool-call collapse observed in the M0 SFT experiments on v1 was caused by the… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2-M0-snapshot.texttext-generation100K<n<1M0 likes635 downloads13d agoHugging Face02arjhinety /OpenGrad-ToolPolicy-Canonical-v1 This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture. What this release is OpenGrad ToolPolicy Canonical v1 is a provenance-preserving, model-independent normalization of several public tool-use and function-calling datasets. It is released as a pre-training candidate corpus for controlled research into tool-use policy in small open-weight language models. See OpenGrad… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v1.texttext-generation100K<n<1M0 likes563 downloads13d agoHugging Face03dongbobo /unified-toolcalls-canonical Unified Tool-Calling Corpus — Canonicalized Output Publish-ready conversion of two pinned Hugging Face dataset revisions into the single schema defined in docs/unified_format.md, with repeated records normalized by an explicit canonicalization rule and every surviving record kept faithful to its source row. Records in (source rows) 65,000 Records published (canonical survivors) 64,622 Duplicates collapsed 378 (343 duplicate groups) Records mutated during… See the full description on the dataset page: https://huggingface.co/datasets/dongbobo/unified-toolcalls-canonical.text-generation10K<n<100K0 likes521 downloads1mo agoHugging Face04laion /llama-nemotron-science-reasoning-on-canonical-think-full Llama-Nemotron science reasoning — Delphi canonical-think (COMPLETE, no length filter) The complete reasoning:on science split of nvidia/Llama-Nemotron-Post-Training-Dataset, converted once into the canonical Delphi chat-template thinking format. 708,920 rows. Unlike the cold-start warmup slice open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k (and its -canonical-think variant), this build applies no length cap and no subsample — every long-CoT science example is… See the full description on the dataset page: https://huggingface.co/datasets/laion/llama-nemotron-science-reasoning-on-canonical-think-full.texttext-generation100K<n<1M0 likes381 downloads21d agoHugging Face05arjhinety /OpenGrad-ToolPolicy-Canonical-v2 This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture. What this release is OpenGrad ToolPolicy Canonical v2 is a provenance-preserving, model-independent normalization of public tool-use datasets in which every record declares what it supervises. It exists because not every legitimate post-training corpus has the same conversational trajectory shape, and discarding a… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2.texttext-generation100K<n<1M0 likes297 downloads12d agoHugging Face06arjhinety /OpenGrad-ToolPolicy-Canonical-v2-minus-xlam This is the byte-identical training view for a joint xLAM-plus-CALL_PREDICTION removal experiment. xLAM is currently the corpus's only source of that supervision contract, so this is not a pure source-content ablation. It carries no result of its own and is not a recommended mixture. It is part of OpenGrad Study 001. What this is OpenGrad-ToolPolicy-Canonical-v2 with one source removed: xLAM/APIGen. Three sources remain, 115,895 canonical records, 118 shards. It is the exact… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2-minus-xlam.texttext-generation100K<n<1M0 likes297 downloads13d agoHugging Face07Nine1Eight /vil-canonical-glyph-system VIL Canonical Glyph System Canonical tri-layer glyph dataset for the GlyphMatics / SigilAGI / VIL stack. Tri-layer identity glyph = (visible, braille, hanzi) digest = SHA256(visible + braille + hanzi) Layers α-layer: visible canonical symbol / glyph role β-layer: Braille-inspired structural state γ-layer: Hanzi temporal-semantic context Canonical role system ID Name Role G0 Origin Root state G1 Split Branch G2 Bind Merge G3 Flow Transition G4 Gate Conditional G5… See the full description on the dataset page: https://huggingface.co/datasets/Nine1Eight/vil-canonical-glyph-system.textfeature-extraction1K<n<10K0 likes187 downloads2mo agoHugging Face08Ikkoikko /thermodivergence-canonical-papers Thermikron & Thermodivergence — Canonical Open-Access Papers Author: Asim Patel · Bangalore, India · ORCID: 0009-0006-2732-8323 Organisations: Thermikron · The Thermodivergence Foundation Dataset Description This dataset contains the full text of two canonical open-access preprints that establish the foundational terminology for two interconnected disciplines: Biothermal microconditioning — integrating biological thermal actors with mechanical HVAC for personalised… See the full description on the dataset page: https://huggingface.co/datasets/Ikkoikko/thermodivergence-canonical-papers.documenttext-generationn<1K0 likes59 downloads7mo agoHugging Face09ArabicNLPWorld /canonical-islamic-corpusgated 🕌 Canonical Islamic Corpus (Quran + Hadith) Description Comprehensive corpus of authentic Islamic texts: 6,236 verses of the Holy Quran from Tanzil (Simple Clean) 315,913 unique hadith matns from four curated collectionsTotal: 322,149 texts with rich metadata. Prepared by Mullosharaf Arabov for IslamicEval 2026 Shared Task (Subtask 2: Hallucination Detection). 📊 Corpus Statistics Metric Value Total entries 322,149 Quran verses 6… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/canonical-islamic-corpus.tabulartext-generation100K<n<1M0 likes21 downloads3mo agoHugging Face10mohamed-ahmed-58059 /text2sql-canonical-v3.1 text2sql-canonical-v3.1 Training data for a SQLite text-to-SQL model: 51,976 training rows and a 527-row validation split. Each row holds a database schema rendered as text, a natural-language question, an optional evidence hint, and the reference SQL. The mix caps synthetic data at 40 percent and gives real benchmark rows double weight. Split Rows BIRD Spider SynSQL train 51,976 33% 27% 40% val 527 34% 27% 40% Columns db_id, question, gold_sql… See the full description on the dataset page: https://huggingface.co/datasets/mohamed-ahmed-58059/text2sql-canonical-v3.1.texttext-generation10K<n<100K0 likes16 downloads1mo agoHugging Face11open-athena /llama-nemotron-science-reasoning-on-le3000tok-100k-canonical-think Llama-Nemotron science reasoning — Delphi cold-start CoT warmup (canonical think tokens) Regenerated (Fix C) variant of laion/llama-nemotron-science-reasoning-on-le3000tok-100k. The original repo's assistant turns carry inline <think>...</think>. LLaMA-Factory's ReasoningTemplate.encode_oneturn checks for the literal canonical string <|start_think|> in the assistant content; inline <think> does NOT satisfy that check, so LF injects an EMPTY canonical block… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k-canonical-think.texttext-generation100K<n<1M0 likes3 downloads21d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.