CoolFace
Datasetpublic

mjbommar/opengloss-v2.0-pretrain

Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Pretrain The release rendered as continuous prose for language-model pretraining or continued pretraining: four document templates per entry — a dictionary entry, a thesaurus entry, an encyclopedia article and a usage note — written as plain text… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-pretrain.

sourceHugging Facecc-by-4.0updated 20d agoView on Hugging Face
0likes123downloads

mjbommar/opengloss-v2.0-pretrain · main · files are served by the source, never re-hosted here