datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenGrad-ToolPolicy-Canonical-v2-M0-snapshot
This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture.
What this release is
OpenGrad ToolPolicy Canonical v2 is a provenance-preserving, model-independent normalization of public tool-use and function-calling datasets. It was built to test one hypothesis with a measurement attached: that the tool-call collapse observed in the M0 SFT experiments on v1 was caused by the… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2-M0-snapshot.OpenGrad-ToolPolicy-Canonical-v1
This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture.
What this release is
OpenGrad ToolPolicy Canonical v1 is a provenance-preserving, model-independent normalization of several public tool-use and function-calling datasets. It is released as a pre-training candidate corpus for controlled research into tool-use policy in small open-weight language models. See OpenGrad… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v1.unified-toolcalls-canonical
Unified Tool-Calling Corpus — Canonicalized Output
Publish-ready conversion of two pinned Hugging Face dataset revisions into the single
schema defined in docs/unified_format.md, with repeated
records normalized by an explicit canonicalization rule and every surviving record
kept faithful to its source row.
Records in (source rows)
65,000
Records published (canonical survivors)
64,622
Duplicates collapsed
378 (343 duplicate groups)
Records mutated during… See the full description on the dataset page: https://huggingface.co/datasets/dongbobo/unified-toolcalls-canonical.llama-nemotron-science-reasoning-on-canonical-think-full
Llama-Nemotron science reasoning — Delphi canonical-think (COMPLETE, no length filter)
The complete reasoning:on science split of
nvidia/Llama-Nemotron-Post-Training-Dataset, converted once into the canonical
Delphi chat-template thinking format. 708,920 rows.
Unlike the cold-start warmup slice
open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k
(and its -canonical-think variant), this build applies no length cap and no subsample — every
long-CoT science example is… See the full description on the dataset page: https://huggingface.co/datasets/laion/llama-nemotron-science-reasoning-on-canonical-think-full.OpenGrad-ToolPolicy-Canonical-v2
This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture.
What this release is
OpenGrad ToolPolicy Canonical v2 is a provenance-preserving, model-independent normalization of
public tool-use datasets in which every record declares what it supervises. It exists because
not every legitimate post-training corpus has the same conversational trajectory shape, and
discarding a… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2.OpenGrad-ToolPolicy-Canonical-v2-minus-xlam
This is the byte-identical training view for a joint xLAM-plus-CALL_PREDICTION removal
experiment. xLAM is currently the corpus's only source of that supervision contract, so this is
not a pure source-content ablation. It carries no result of its own and is not a recommended
mixture. It is part of OpenGrad Study 001.
What this is
OpenGrad-ToolPolicy-Canonical-v2
with one source removed: xLAM/APIGen. Three sources remain, 115,895 canonical records, 118
shards.
It is the exact… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2-minus-xlam.vil-canonical-glyph-system
VIL Canonical Glyph System
Canonical tri-layer glyph dataset for the GlyphMatics / SigilAGI / VIL stack.
Tri-layer identity
glyph = (visible, braille, hanzi)
digest = SHA256(visible + braille + hanzi)
Layers
α-layer: visible canonical symbol / glyph role
β-layer: Braille-inspired structural state
γ-layer: Hanzi temporal-semantic context
Canonical role system
ID
Name
Role
G0
Origin
Root state
G1
Split
Branch
G2
Bind
Merge
G3
Flow
Transition
G4
Gate
Conditional
G5… See the full description on the dataset page: https://huggingface.co/datasets/Nine1Eight/vil-canonical-glyph-system.thermodivergence-canonical-papers
Thermikron & Thermodivergence — Canonical Open-Access Papers
Author: Asim Patel · Bangalore, India · ORCID: 0009-0006-2732-8323
Organisations: Thermikron · The Thermodivergence Foundation
Dataset Description
This dataset contains the full text of two canonical open-access preprints that establish the foundational terminology for two interconnected disciplines:
Biothermal microconditioning — integrating biological thermal actors with mechanical HVAC for personalised… See the full description on the dataset page: https://huggingface.co/datasets/Ikkoikko/thermodivergence-canonical-papers.canonical-islamic-corpus
🕌 Canonical Islamic Corpus (Quran + Hadith)
Description
Comprehensive corpus of authentic Islamic texts:
6,236 verses of the Holy Quran from Tanzil (Simple Clean)
315,913 unique hadith matns from four curated collectionsTotal: 322,149 texts with rich metadata.
Prepared by Mullosharaf Arabov for IslamicEval 2026 Shared Task (Subtask 2: Hallucination Detection).
📊 Corpus Statistics
Metric
Value
Total entries
322,149
Quran verses
6… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/canonical-islamic-corpus.text2sql-canonical-v3.1
text2sql-canonical-v3.1
Training data for a SQLite text-to-SQL model: 51,976 training rows and
a 527-row validation split. Each row holds a database schema rendered
as text, a natural-language question, an optional evidence hint, and
the reference SQL. The mix caps synthetic data at 40 percent and gives
real benchmark rows double weight.
Split
Rows
BIRD
Spider
SynSQL
train
51,976
33%
27%
40%
val
527
34%
27%
40%
Columns
db_id, question, gold_sql… See the full description on the dataset page: https://huggingface.co/datasets/mohamed-ahmed-58059/text2sql-canonical-v3.1.llama-nemotron-science-reasoning-on-le3000tok-100k-canonical-think
Llama-Nemotron science reasoning — Delphi cold-start CoT warmup (canonical think tokens)
Regenerated (Fix C) variant of
laion/llama-nemotron-science-reasoning-on-le3000tok-100k.
The original repo's assistant turns carry inline <think>...</think>. LLaMA-Factory's
ReasoningTemplate.encode_oneturn checks for the literal canonical string <|start_think|>
in the assistant content; inline <think> does NOT satisfy that check, so LF injects an EMPTY
canonical block… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k-canonical-think.
