datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
betty-dota2-canonical-v1canonical_Xperience
canonical_Xperience
Xperience hand-depth data at 256-pixel resolution.
Access to this dataset is manually reviewed by the repository owner.
Repository layout
Hugging Face limits each directory to 10,000 files. The first 9,990 files retain
their original paths under stereo/; the remaining 4,992 stereo files are stored
under stereo/overflow/. Filenames are unchanged. The original source README is
preserved as SOURCE_README.md.
betty-dota2-canonical-v1
Betty Dota 2 Canonical Dataset
Enriched version of the Dota 2 match data.
Created during backfill process.
OpenGrad-ToolPolicy-Canonical-v2-M0-snapshot
This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture.
What this release is
OpenGrad ToolPolicy Canonical v2 is a provenance-preserving, model-independent normalization of public tool-use and function-calling datasets. It was built to test one hypothesis with a measurement attached: that the tool-call collapse observed in the M0 SFT experiments on v1 was caused by the… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2-M0-snapshot.OpenGrad-ToolPolicy-Canonical-v1
This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture.
What this release is
OpenGrad ToolPolicy Canonical v1 is a provenance-preserving, model-independent normalization of several public tool-use and function-calling datasets. It is released as a pre-training candidate corpus for controlled research into tool-use policy in small open-weight language models. See OpenGrad… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v1.canonical-pores
SubstrateCommons/canonical-pores
Replay packs (.rpk) built with dmipy_sim: one converged Monte-Carlo walk each, stored so any acquisition can be replayed on it. Load one with ReplayPack.load("hf://SubstrateCommons/canonical-pores/<path>"); the manifest (manifest.json) holds the sha256 every load is checked against. This file is rendered from the manifest by dmipy_sim.replay.publish.
Substrate
analytic/sphere
box: 0.4 × 0.4 × 0.4 µm
boundary: open, open, open… See the full description on the dataset page: https://huggingface.co/datasets/SubstrateCommons/canonical-pores.ZINC-canonicalized
dataset description
We downloaded ZINC dataset from here and canonicalized it.
We used the following function to canonicalize the data and removed some SMILES that cannot be read by RDKit.
from rdkit import Chem
def canonicalize(mol):
mol = Chem.MolToSmiles(Chem.MolFromSmiles(mol),True)
return mol
We randomly split the preprocessed data into train and validation. The ratio is 9 : 1.
unified-toolcalls-canonical
Unified Tool-Calling Corpus — Canonicalized Output
Publish-ready conversion of two pinned Hugging Face dataset revisions into the single
schema defined in docs/unified_format.md, with repeated
records normalized by an explicit canonicalization rule and every surviving record
kept faithful to its source row.
Records in (source rows)
65,000
Records published (canonical survivors)
64,622
Duplicates collapsed
378 (343 duplicate groups)
Records mutated during… See the full description on the dataset page: https://huggingface.co/datasets/dongbobo/unified-toolcalls-canonical.nmr-canonical-cleaned
Canonical NMR Dataset Collection — Data Card
Dataset release: v4
Canonical schema: v2
Spectral modalities: 1H and 13C resonance-level peak lists
Collection overview
This release brings several of the largest openly available processed NMR
corpora used by current deep-learning methods into one model-independent
schema. It combines simulated and literature-derived spectra while preserving
the provenance and annotation coverage of every source.
The collection has… See the full description on the dataset page: https://huggingface.co/datasets/niccogreek/nmr-canonical-cleaned.aquatype-canonical-ctx-20260706canonical-formulas-v1
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZLHOLDINGS/canonical-formulas-v1
The canonical SZL formula registry — 21 pure, typed, no-IO Python formulas, the
matching Lean 4 obligation theorems, and the Codex-Kernel governed-loop composer.
Contents
File
What
code/python/formulas.py
21 canonical formulas, each… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/canonical-formulas-v1.llama-nemotron-science-reasoning-on-canonical-think-full
Llama-Nemotron science reasoning — Delphi canonical-think (COMPLETE, no length filter)
The complete reasoning:on science split of
nvidia/Llama-Nemotron-Post-Training-Dataset, converted once into the canonical
Delphi chat-template thinking format. 708,920 rows.
Unlike the cold-start warmup slice
open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k
(and its -canonical-think variant), this build applies no length cap and no subsample — every
long-CoT science example is… See the full description on the dataset page: https://huggingface.co/datasets/laion/llama-nemotron-science-reasoning-on-canonical-think-full.canonical_prompt_collectionasolaria-record-231-canonical
asolaria-record-231-canonical
The photographic record of Jesse Daniel Brown, in his own numbering.
What is here
path
what it is
photos/
231 photographs, numbered 001–231, each keeping its original camera filename after the number
MAPPING.tsv
the canonical index: number, path, original filename, byte size, SHA-256
CHECKSUMS-231.sha256
machine-checkable form of the same, for sha256sum -c
OBSERVATION.md
the written observation document, 8,194 lines… See the full description on the dataset page: https://huggingface.co/datasets/Jessedbrown/asolaria-record-231-canonical.OpenGrad-ToolPolicy-Canonical-v2
This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture.
What this release is
OpenGrad ToolPolicy Canonical v2 is a provenance-preserving, model-independent normalization of
public tool-use datasets in which every record declares what it supervises. It exists because
not every legitimate post-training corpus has the same conversational trajectory shape, and
discarding a… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2.OpenGrad-ToolPolicy-Canonical-v2-minus-xlam
This is the byte-identical training view for a joint xLAM-plus-CALL_PREDICTION removal
experiment. xLAM is currently the corpus's only source of that supervision contract, so this is
not a pure source-content ablation. It carries no result of its own and is not a recommended
mixture. It is part of OpenGrad Study 001.
What this is
OpenGrad-ToolPolicy-Canonical-v2
with one source removed: xLAM/APIGen. Three sources remain, 115,895 canonical records, 118
shards.
It is the exact… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2-minus-xlam.gurbani-sehajpath-yt-captions-canonical
Gurbani Sehajpath — Canonical-aligned ASR corpus
Stage-1 + Stage-2 canonical pipeline output for sehaj-path (calm recitation of the Guru Granth Sahib). Built from publicly available audio recordings with aligned transcripts, chunked by caption timing and aligned against the canonical Guru Granth Sahib Ji text (SGGS).
Columns
Schema is auto-inferred from the parquet shards. Primary columns:
audio — 16 kHz mono waveform
final_text — canonical Gurmukhi transcription (post… See the full description on the dataset page: https://huggingface.co/datasets/surindersinghssj/gurbani-sehajpath-yt-captions-canonical.pubchem-10m-canonicalized
dataset description
We downloaded PubChem-10m dataset from here and canonicalized it.
We used the following function to canonicalize the data and removed some SMILES that cannot be read by RDKit.
from rdkit import Chem
def canonicalize(mol):
mol = Chem.MolToSmiles(Chem.MolFromSmiles(mol),True)
return mol
We randomly split the preprocessed data into train and validation. The ratio is 9 : 1.
vil-canonical-glyph-system
VIL Canonical Glyph System
Canonical tri-layer glyph dataset for the GlyphMatics / SigilAGI / VIL stack.
Tri-layer identity
glyph = (visible, braille, hanzi)
digest = SHA256(visible + braille + hanzi)
Layers
α-layer: visible canonical symbol / glyph role
β-layer: Braille-inspired structural state
γ-layer: Hanzi temporal-semantic context
Canonical role system
ID
Name
Role
G0
Origin
Root state
G1
Split
Branch
G2
Bind
Merge
G3
Flow
Transition
G4
Gate
Conditional
G5… See the full description on the dataset page: https://huggingface.co/datasets/Nine1Eight/vil-canonical-glyph-system.canonical_obligation_datasettinyperson-copy-paste-canonical-matrix-runscanonical-dataset
canonical-dataset
A parallel, multilingual dataset on rule-following
Languages
en — English
am — Amharic
de — German
hi — Hindi
ig — Igbo
it — Italian
ko — Korean
ru — Russian
sw — Swahili
ta — Tamil
tr — Turkish
ur — Urdu
yo — Yoruba
Loading
from datasets import load_dataset
en = load_dataset("canonical-dataset", "en", split="test")
yo = load_dataset("canonical-dataset", "yo", split="test")
lichess-stockfish-canonicalDelighting-Benchmark-v1-Canonical-Outputs
Delighting Benchmark v1 — Canonical Outputs
Canonical Paint3D and Hunyuan3D-2.0 texture-generation outputs for 30 test
cases containing visible illumination and specular effects:
10 ABO matte cases
10 ABO glossy cases
10 Poly Haven PBR cases
Layout
Results are grouped by method:
paint3d/<case_id>/
hunyuan3d/<case_id>/
Every case contains exactly seven files:
input_image.png
input_mesh.obj
textured.glb
albedo_render.mp4
albedo_texture.png
gt_lit_render.mp4… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/Delighting-Benchmark-v1-Canonical-Outputs.CIC-IoT-2023-canonical-neto
CIC-IoT-2023 — Canonical (Neto et al.) Variant
This is the canonical CIC-IoT-2023 dataset, sourced from
bencorn/CIC-IoT-2023's
CSV/MERGED_CSV/ folder, which contains Neto et al.'s authentic merged CSVs
WITH embedded labels (vs. bencorn's other CSV/CSV/<attack>/ re-organization
which lost ~6.5M rows during the folder-restructure).
Why this exists: prior lacg030175/CIC-IoT-2023-full and -full-raw were
built from CSV/CSV/ and contained only 38.5M rows. This one contains
~45,019,243… See the full description on the dataset page: https://huggingface.co/datasets/lacg030175/CIC-IoT-2023-canonical-neto.the-cohesive-tetrad-canonical
The Cohesive Tetrad — Canonical Dataset (v1.0.0)
Dataset ini adalah korpus kanonis untuk melatih dan mengevaluasi model instruksiThe Cohesive Tetrad (TCT), khususnya:
suratkiade/the-cohesive-tetrad-instruct-base (sebagai model dasar / mirror TinyLlama), dan
suratkiade/the-cohesive-tetrad-instruct-v1 (sebagai model instruksi kanonis hasil fine-tuning).
Seluruh isi dataset dan berkas sumber dinyatakan di bawah lisensi CC0 1.0 (Public Domain Dedication).Secara epistemik, dataset… See the full description on the dataset page: https://huggingface.co/datasets/suratkiade/the-cohesive-tetrad-canonical.christianity-canonical-corpus
Christianity Canonical Corpus (Theological and Scriptural Corpus)
A comprehensive, curated, and machine-readable JSON dataset encompassing the canonical biblical scriptures in original languages and historical translations, Thomas Aquinas's Summa Theologiae, the Early Church Fathers (Ante-Nicene and Nicene series), historic ecumenical creeds, Protestant confessions, and Matthew Henry's commentaries.
📖 Corpus Overview & Structure
christianity-canonical-corpus/
├──… See the full description on the dataset page: https://huggingface.co/datasets/Yehuda-Rubin/christianity-canonical-corpus.canonical-order-problem
List Extraction Dataset
A dataset for studying how large language models (LLMs) represent and generate multi-valued relations — relations in which a single subject is associated with a set of entities (e.g., "all countries in South America", "all elements of the periodic table").
It accompanies the paper:
The Canonical Order Problem: When Large Language Models Are Unreliable Knowledge Bases for Multi-Valued Relations
Timo Pierre Schrader, Annemarie Friedrich, Simon Razniewski… See the full description on the dataset page: https://huggingface.co/datasets/timo-pierre-schrader/canonical-order-problem.augmented_canonical_pubchem_13m
PubChem 10M - Augmented SMILES Dataset
This dataset is derived from the original PubChem 10M and has been canonicalized using RDKit (2024.9.4) to ensure structural consistency.
To enhance molecular diversity, 33% of the dataset was randomly sampled and augmented using RDKit’s Chem.MolToRandomSmilesVect function, following an approach similar to NVIDIA’s molmim method for SMILES augmentation.
Dataset Overview:
Source: PubChem 10M
Canonicalization: RDKit (2024.9.4)… See the full description on the dataset page: https://huggingface.co/datasets/Derify/augmented_canonical_pubchem_13m.augmented_canonical_druglike_QED_43m
Druglike QED 43M - Augmented SMILES Dataset
This dataset is derived from the Druglike molecule datasets for drug discovery dataset and has been canonicalized using RDKit (2024.9.4) to ensure structural consistency.
To enhance molecular diversity, 33% of the dataset was randomly sampled and augmented using RDKit’s Chem.MolToRandomSmilesVect function, following an approach similar to NVIDIA's molmim method for SMILES augmentation.
Dataset Overview:
Source:… See the full description on the dataset page: https://huggingface.co/datasets/Derify/augmented_canonical_druglike_QED_43m.
