datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glyph-datasets
glyph-datasets
Data artifacts for Glyph — chemical-structure transcription with two models:
OCSRGlyph (molecule image → SMILES) and MarkushGlyph (patent Markush
image → CXSMILES + R-group table). This repository holds the training indices,
training metadata, and the frozen self-contained evaluation benchmarks that the
glyph package auto-downloads.
Code: EdisonScientific/glyph (GitHub)
Weights: EdisonScientific/OCSRGlyph, EdisonScientific/MarkushGlyph
License: Apache-2.0. See… See the full description on the dataset page: https://huggingface.co/datasets/EdisonScientific/glyph-datasets.vil-canonical-glyph-system
VIL Canonical Glyph System
Canonical tri-layer glyph dataset for the GlyphMatics / SigilAGI / VIL stack.
Tri-layer identity
glyph = (visible, braille, hanzi)
digest = SHA256(visible + braille + hanzi)
Layers
α-layer: visible canonical symbol / glyph role
β-layer: Braille-inspired structural state
γ-layer: Hanzi temporal-semantic context
Canonical role system
ID
Name
Role
G0
Origin
Root state
G1
Split
Branch
G2
Bind
Merge
G3
Flow
Transition
G4
Gate
Conditional
G5… See the full description on the dataset page: https://huggingface.co/datasets/Nine1Eight/vil-canonical-glyph-system.glyph_machinatwkm-maya-glyphs
TWKM Maya Glyphs
An ML-ready, relational snapshot of the University of Bonn's Text Database and Dictionary of Classic Mayan (TWKM) digital sign catalogue. It packages the catalogue as seven Parquet configurations and embeds the explicitly CC BY 4.0 standardized graph drawings in the graphs configuration.
This is an independent preservation and interoperability package, not an official TWKM publication. The source of authority for sign classification and interpretation remains… See the full description on the dataset page: https://huggingface.co/datasets/tadad/twkm-maya-glyphs.openr1_glyph_1223_4kopenr1_glyph_1223_4k_qwenopus-4.6-frontend-development
CoT Code Debugging Dataset
Synthetic code debugging examples with chain-of-thought (CoT) reasoning and solutions, built with a three-stage pipeline: seed problem → evolved problem → detailed solve. Topics emphasize frontend / UI engineering (CSS, React, accessibility, layout, design systems, SSR/hydration, and related product UI issues).
Each line in dataset.jsonl is one JSON object (JSONL format).
Data fields
Field
Description
id
16-character hex id:… See the full description on the dataset page: https://huggingface.co/datasets/glyphsoftware/opus-4.6-frontend-development.openr1_glyph_1127_oriGlyphnet
GlyphNet: Homoglyph Domains Dataset
Data for detecting homoglyph
phishing domains (e.g. facebook.com spoofed with visually-similar Unicode
characters). Every genuine domain is paired with a synthetically generated
homoglyph variant, and each domain is also rendered to a 256x256 grayscale image
so the task can be tackled as text or image classification.
Paper: arXiv:2306.10392 ·
Code: github.com/Akshat4112/Glyphnet
Configs & splits
Both configs share the same train… See the full description on the dataset page: https://huggingface.co/datasets/Akshat4112/Glyphnet.m-digital-calligraphy-mini-set
M.Digital Calligraphy Project – Mini Set
ver.1.2 / 2026-01-15
🖋 Overview
M.Digital Calligraphy Project reconstructs the beauty of traditional Japanese calligraphy through digital brush techniques and AI-ready datasets.
The Mini Set features 10 representative kanji characters inspired by nature and seasons. This set is provided as a free trial and promotional version for research and personal evaluation.
📜 Included Characters (10)
Kanji… See the full description on the dataset page: https://huggingface.co/datasets/M-Glyph/m-digital-calligraphy-mini-set.GlyphByT5Pretraining
GlyphByT5 Pretraining
This is the pretraining data for Glyph-ByT5.
Dataset Details
Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text RenderingZeyu Liu, Weicong Liang, Zhanhao Liang, Chong Luo, Ji Li, Gao Huang, Yuhui Yuan
Microsoft Research Asia; Tsinghua University; Peking University; The Australian National UniversityPreprint
Dataset Structure
This dataset contains a json file containing the annotation file needed for Glyph-ByT5 pretraining.… See the full description on the dataset page: https://huggingface.co/datasets/GlyphByT5/GlyphByT5Pretraining.glyphmatics-complete-training-dataset
GlyphMatics Complete Training Dataset
Canonical synthetic training data for GlyphMatics / SigilAGI.
Covers
glyph encoding
glyph decoding
semantic compression
reconstruction
Alpha/Beta/Gamma mapping
SigilAGI routing
VIL normalization
GIIBL lattice blocks
RC3 cube encoding
Quantum Glyph states
mobile deployment planning
safety-aware symbolic transformation
Dataset Viewer
The public dataset viewer is configured only for:
data/train.jsonl
data/validation.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Nine1Eight/glyphmatics-complete-training-dataset.myanmar-synthetic-syllable-glyphs
🇲🇲 Myanmar Synthetic Syllable Glyphs (MSSG)
The Myanmar Synthetic Syllable Glyphs (MSSG) is a massive-scale, high-fidelity synthetic image dataset containing 14,295,552 heavily augmented glyph images (128x64 pixels, grayscale) representing the structural combinatorial matrix of the Burmese script.
Developed and engineered by Khant Sint Heinn (Kalix Louis), this core foundational dataset is officially published and maintained under DatarrX (Myanmar Open Source Organization… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-synthetic-syllable-glyphs.VLM-GlyphSlimOrca-DeDup-Glyphe-20kmyanmar-word-glyphs
🇲🇲 Myanmar Word Glyphs (MWG)
The Myanmar Word Glyphs (MWG) is a curated vocabulary-based synthetic image dataset containing 49,800 high-quality word/phrase glyph images (256x64 pixels, grayscale).
Developed and engineered by Khant Sint Heinn, this dataset is officially published and distributed under DatarrX (Myanmar Open Source Organization, NPO). While our sibling project—MSSG—explores the absolute mathematical grid of theoretical syllables, MWG is designed to map out authentic… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-word-glyphs.myanmar-numeral-glyphs
🇲🇲 Myanmar Numeral Glyphs (MNG)
The Myanmar Numeral Glyphs (MNG) dataset is a curated, high-quality hybrid image dataset designed for optical character recognition (OCR) and image classification tasks targeting native Burmese digits (၀ to ၉). Released under DatarrX, this dataset bridges the gap in low-resource language resources by combining clean, human-annotated handwritten data with robust computer-generated font variations.
📌 Dataset Overview
Total Images: 1… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-numeral-glyphs.multiscript_glyph_images_v1
Multiscript Glyph Images (v1)
Balanced 128×128-pixel grayscale glyph set for multilingual diffusion and OCR
experiments.
Script
Unicode
Char
Caption
# images
Latin
U+0041 A
A
english character A
100
Latin
U+0042 B
B
english character B
100
Latin
U+0047 G
G
english character G
100
Devanagari
U+0915 क
ka
hindi character ka
100
Devanagari
U+0932 ल
la
hindi character la
100
Devanagari
U+0917 ग
ga
hindi character ga
100
Arabic
U+0641 ف
faa
arabic character faa
400… See the full description on the dataset page: https://huggingface.co/datasets/adishri/multiscript_glyph_images_v1.Glyph-SVG-LLaVA-Datasetglyph-word
