CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wolframko /betty-dota2-canonical-v10 likes3.7k downloads6mo agoHugging Face02mojimoji61 /canonical_Xperiencegated canonical_Xperience Xperience hand-depth data at 256-pixel resolution. Access to this dataset is manually reviewed by the repository owner. Repository layout Hugging Face limits each directory to 10,000 files. The first 9,990 files retain their original paths under stereo/; the remaining 4,992 stereo files are stored under stereo/overflow/. Filenames are unchanged. The original source README is preserved as SOURCE_README.md. videodepth-estimation10K<n<100K0 likes1.3k downloads24d agoHugging Face03apararti /betty-dota2-canonical-v1 Betty Dota 2 Canonical Dataset Enriched version of the Dota 2 match data. Created during backfill process. 0 likes1.1k downloads6mo agoHugging Face04arjhinety /OpenGrad-ToolPolicy-Canonical-v2-M0-snapshot This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture. What this release is OpenGrad ToolPolicy Canonical v2 is a provenance-preserving, model-independent normalization of public tool-use and function-calling datasets. It was built to test one hypothesis with a measurement attached: that the tool-call collapse observed in the M0 SFT experiments on v1 was caused by the… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2-M0-snapshot.texttext-generation100K<n<1M0 likes635 downloads13d agoHugging Face05arjhinety /OpenGrad-ToolPolicy-Canonical-v1 This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture. What this release is OpenGrad ToolPolicy Canonical v1 is a provenance-preserving, model-independent normalization of several public tool-use and function-calling datasets. It is released as a pre-training candidate corpus for controlled research into tool-use policy in small open-weight language models. See OpenGrad… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v1.texttext-generation100K<n<1M0 likes563 downloads13d agoHugging Face06SubstrateCommons /canonical-pores SubstrateCommons/canonical-pores Replay packs (.rpk) built with dmipy_sim: one converged Monte-Carlo walk each, stored so any acquisition can be replayed on it. Load one with ReplayPack.load("hf://SubstrateCommons/canonical-pores/<path>"); the manifest (manifest.json) holds the sha256 every load is checked against. This file is rendered from the manifest by dmipy_sim.replay.publish. Substrate analytic/sphere box: 0.4 × 0.4 × 0.4 µm boundary: open, open, open… See the full description on the dataset page: https://huggingface.co/datasets/SubstrateCommons/canonical-pores.1 likes554 downloads0m agoHugging Face07sagawa /ZINC-canonicalized dataset description We downloaded ZINC dataset from here and canonicalized it. We used the following function to canonicalize the data and removed some SMILES that cannot be read by RDKit. from rdkit import Chem def canonicalize(mol): mol = Chem.MolToSmiles(Chem.MolFromSmiles(mol),True) return mol We randomly split the preprocessed data into train and validation. The ratio is 9 : 1. text10M<n<100M0 likes542 downloads4y agoHugging Face08dongbobo /unified-toolcalls-canonical Unified Tool-Calling Corpus — Canonicalized Output Publish-ready conversion of two pinned Hugging Face dataset revisions into the single schema defined in docs/unified_format.md, with repeated records normalized by an explicit canonicalization rule and every surviving record kept faithful to its source row. Records in (source rows) 65,000 Records published (canonical survivors) 64,622 Duplicates collapsed 378 (343 duplicate groups) Records mutated during… See the full description on the dataset page: https://huggingface.co/datasets/dongbobo/unified-toolcalls-canonical.text-generation10K<n<100K0 likes521 downloads1mo agoHugging Face09niccogreek /nmr-canonical-cleaned Canonical NMR Dataset Collection — Data Card Dataset release: v4 Canonical schema: v2 Spectral modalities: 1H and 13C resonance-level peak lists Collection overview This release brings several of the largest openly available processed NMR corpora used by current deep-learning methods into one model-independent schema. It combines simulated and literature-derived spectra while preserving the provenance and annotation coverage of every source. The collection has… See the full description on the dataset page: https://huggingface.co/datasets/niccogreek/nmr-canonical-cleaned.feature-extraction100M<n<1B0 likes429 downloads3d agoHugging Face10dio7153 /aquatype-canonical-ctx-20260706tabular100M<n<1B1 likes405 downloads3mo agoHugging Face11SZLHOLDINGS /canonical-formulas-v1 Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance. SZLHOLDINGS/canonical-formulas-v1 The canonical SZL formula registry — 21 pure, typed, no-IO Python formulas, the matching Lean 4 obligation theorems, and the Codex-Kernel governed-loop composer. Contents File What code/python/formulas.py 21 canonical formulas, each… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/canonical-formulas-v1.othern<1K0 likes401 downloads25d agoHugging Face12laion /llama-nemotron-science-reasoning-on-canonical-think-full Llama-Nemotron science reasoning — Delphi canonical-think (COMPLETE, no length filter) The complete reasoning:on science split of nvidia/Llama-Nemotron-Post-Training-Dataset, converted once into the canonical Delphi chat-template thinking format. 708,920 rows. Unlike the cold-start warmup slice open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k (and its -canonical-think variant), this build applies no length cap and no subsample — every long-CoT science example is… See the full description on the dataset page: https://huggingface.co/datasets/laion/llama-nemotron-science-reasoning-on-canonical-think-full.texttext-generation100K<n<1M0 likes381 downloads21d agoHugging Face13SkillFactory /canonical_prompt_collectiontext100K<n<1M0 likes352 downloads10mo agoHugging Face14Jessedbrown /asolaria-record-231-canonical asolaria-record-231-canonical The photographic record of Jesse Daniel Brown, in his own numbering. What is here path what it is photos/ 231 photographs, numbered 001–231, each keeping its original camera filename after the number MAPPING.tsv the canonical index: number, path, original filename, byte size, SHA-256 CHECKSUMS-231.sha256 machine-checkable form of the same, for sha256sum -c OBSERVATION.md the written observation document, 8,194 lines… See the full description on the dataset page: https://huggingface.co/datasets/Jessedbrown/asolaria-record-231-canonical.imagen<1K0 likes340 downloads2mo agoHugging Face15arjhinety /OpenGrad-ToolPolicy-Canonical-v2 This is a provenance-preserving canonical candidate corpus. It is a pre-training canonical release, not an empirically selected or recommended training mixture. What this release is OpenGrad ToolPolicy Canonical v2 is a provenance-preserving, model-independent normalization of public tool-use datasets in which every record declares what it supervises. It exists because not every legitimate post-training corpus has the same conversational trajectory shape, and discarding a… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2.texttext-generation100K<n<1M0 likes297 downloads12d agoHugging Face16arjhinety /OpenGrad-ToolPolicy-Canonical-v2-minus-xlam This is the byte-identical training view for a joint xLAM-plus-CALL_PREDICTION removal experiment. xLAM is currently the corpus's only source of that supervision contract, so this is not a pure source-content ablation. It carries no result of its own and is not a recommended mixture. It is part of OpenGrad Study 001. What this is OpenGrad-ToolPolicy-Canonical-v2 with one source removed: xLAM/APIGen. Three sources remain, 115,895 canonical records, 118 shards. It is the exact… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2-minus-xlam.texttext-generation100K<n<1M0 likes297 downloads13d agoHugging Face17surindersinghssj /gurbani-sehajpath-yt-captions-canonical Gurbani Sehajpath — Canonical-aligned ASR corpus Stage-1 + Stage-2 canonical pipeline output for sehaj-path (calm recitation of the Guru Granth Sahib). Built from publicly available audio recordings with aligned transcripts, chunked by caption timing and aligned against the canonical Guru Granth Sahib Ji text (SGGS). Columns Schema is auto-inferred from the parquet shards. Primary columns: audio — 16 kHz mono waveform final_text — canonical Gurmukhi transcription (post… See the full description on the dataset page: https://huggingface.co/datasets/surindersinghssj/gurbani-sehajpath-yt-captions-canonical.audioautomatic-speech-recognition10K<n<100K0 likes240 downloads5mo agoHugging Face18sagawa /pubchem-10m-canonicalized dataset description We downloaded PubChem-10m dataset from here and canonicalized it. We used the following function to canonicalize the data and removed some SMILES that cannot be read by RDKit. from rdkit import Chem def canonicalize(mol): mol = Chem.MolToSmiles(Chem.MolFromSmiles(mol),True) return mol We randomly split the preprocessed data into train and validation. The ratio is 9 : 1. text1M<n<10M7 likes197 downloads4y agoHugging Face19Nine1Eight /vil-canonical-glyph-system VIL Canonical Glyph System Canonical tri-layer glyph dataset for the GlyphMatics / SigilAGI / VIL stack. Tri-layer identity glyph = (visible, braille, hanzi) digest = SHA256(visible + braille + hanzi) Layers α-layer: visible canonical symbol / glyph role β-layer: Braille-inspired structural state γ-layer: Hanzi temporal-semantic context Canonical role system ID Name Role G0 Origin Root state G1 Split Branch G2 Bind Merge G3 Flow Transition G4 Gate Conditional G5… See the full description on the dataset page: https://huggingface.co/datasets/Nine1Eight/vil-canonical-glyph-system.textfeature-extraction1K<n<10K0 likes187 downloads2mo agoHugging Face20nunaa /canonical_obligation_datasettext1K<n<10K0 likes180 downloads20d agoHugging Face21duyle2408 /tinyperson-copy-paste-canonical-matrix-runstabular1M<n<10M0 likes179 downloads8d agoHugging Face22crosslingual-rule-following /canonical-dataset canonical-dataset A parallel, multilingual dataset on rule-following Languages en — English am — Amharic de — German hi — Hindi ig — Igbo it — Italian ko — Korean ru — Russian sw — Swahili ta — Tamil tr — Turkish ur — Urdu yo — Yoruba Loading from datasets import load_dataset en = load_dataset("canonical-dataset", "en", split="test") yo = load_dataset("canonical-dataset", "yo", split="test") text10K<n<100K0 likes161 downloads1mo agoHugging Face23SaakethS /lichess-stockfish-canonicaltext100M<n<1B0 likes110 downloads11h agoHugging Face24ZeyuLing /Delighting-Benchmark-v1-Canonical-Outputs Delighting Benchmark v1 — Canonical Outputs Canonical Paint3D and Hunyuan3D-2.0 texture-generation outputs for 30 test cases containing visible illumination and specular effects: 10 ABO matte cases 10 ABO glossy cases 10 Poly Haven PBR cases Layout Results are grouped by method: paint3d/<case_id>/ hunyuan3d/<case_id>/ Every case contains exactly seven files: input_image.png input_mesh.obj textured.glb albedo_render.mp4 albedo_texture.png gt_lit_render.mp4… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/Delighting-Benchmark-v1-Canonical-Outputs.3d0 likes104 downloads2mo agoHugging Face25lacg030175 /CIC-IoT-2023-canonical-neto CIC-IoT-2023 — Canonical (Neto et al.) Variant This is the canonical CIC-IoT-2023 dataset, sourced from bencorn/CIC-IoT-2023's CSV/MERGED_CSV/ folder, which contains Neto et al.'s authentic merged CSVs WITH embedded labels (vs. bencorn's other CSV/CSV/<attack>/ re-organization which lost ~6.5M rows during the folder-restructure). Why this exists: prior lacg030175/CIC-IoT-2023-full and -full-raw were built from CSV/CSV/ and contained only 38.5M rows. This one contains ~45,019,243… See the full description on the dataset page: https://huggingface.co/datasets/lacg030175/CIC-IoT-2023-canonical-neto.tabular10M<n<100M0 likes83 downloads5mo agoHugging Face26suratkiade /the-cohesive-tetrad-canonical The Cohesive Tetrad — Canonical Dataset (v1.0.0) Dataset ini adalah korpus kanonis untuk melatih dan mengevaluasi model instruksiThe Cohesive Tetrad (TCT), khususnya: suratkiade/the-cohesive-tetrad-instruct-base (sebagai model dasar / mirror TinyLlama), dan suratkiade/the-cohesive-tetrad-instruct-v1 (sebagai model instruksi kanonis hasil fine-tuning). Seluruh isi dataset dan berkas sumber dinyatakan di bawah lisensi CC0 1.0 (Public Domain Dedication).Secara epistemik, dataset… See the full description on the dataset page: https://huggingface.co/datasets/suratkiade/the-cohesive-tetrad-canonical.1K<n<10K0 likes72 downloads10mo agoHugging Face27Yehuda-Rubin /christianity-canonical-corpus Christianity Canonical Corpus (Theological and Scriptural Corpus) A comprehensive, curated, and machine-readable JSON dataset encompassing the canonical biblical scriptures in original languages and historical translations, Thomas Aquinas's Summa Theologiae, the Early Church Fathers (Ante-Nicene and Nicene series), historic ecumenical creeds, Protestant confessions, and Matthew Henry's commentaries. 📖 Corpus Overview & Structure christianity-canonical-corpus/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Yehuda-Rubin/christianity-canonical-corpus.100K<n<1M0 likes68 downloads1mo agoHugging Face28timo-pierre-schrader /canonical-order-problem List Extraction Dataset A dataset for studying how large language models (LLMs) represent and generate multi-valued relations — relations in which a single subject is associated with a set of entities (e.g., "all countries in South America", "all elements of the periodic table"). It accompanies the paper: The Canonical Order Problem: When Large Language Models Are Unreliable Knowledge Bases for Multi-Valued Relations Timo Pierre Schrader, Annemarie Friedrich, Simon Razniewski… See the full description on the dataset page: https://huggingface.co/datasets/timo-pierre-schrader/canonical-order-problem.tabularother10K<n<100K0 likes67 downloads13d agoHugging Face29Derify /augmented_canonical_pubchem_13m PubChem 10M - Augmented SMILES Dataset This dataset is derived from the original PubChem 10M and has been canonicalized using RDKit (2024.9.4) to ensure structural consistency. To enhance molecular diversity, 33% of the dataset was randomly sampled and augmented using RDKit’s Chem.MolToRandomSmilesVect function, following an approach similar to NVIDIA’s molmim method for SMILES augmentation. Dataset Overview: Source: PubChem 10M Canonicalization: RDKit (2024.9.4)… See the full description on the dataset page: https://huggingface.co/datasets/Derify/augmented_canonical_pubchem_13m.textfeature-extraction10M<n<100M0 likes63 downloads1y agoHugging Face30Derify /augmented_canonical_druglike_QED_43m Druglike QED 43M - Augmented SMILES Dataset This dataset is derived from the Druglike molecule datasets for drug discovery dataset and has been canonicalized using RDKit (2024.9.4) to ensure structural consistency. To enhance molecular diversity, 33% of the dataset was randomly sampled and augmented using RDKit’s Chem.MolToRandomSmilesVect function, following an approach similar to NVIDIA's molmim method for SMILES augmentation. Dataset Overview: Source:… See the full description on the dataset page: https://huggingface.co/datasets/Derify/augmented_canonical_druglike_QED_43m.textfeature-extraction10M<n<100M1 likes63 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.