CoolFace
Datasetpublic

mjbommar/opengloss-v2.0-pretrain

Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Pretrain The release rendered as continuous prose for language-model pretraining or continued pretraining: four document templates per entry — a dictionary entry, a thesaurus entry, an encyclopedia article and a usage note — written as plain text… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-pretrain.

sourceHugging Facecc-by-4.0updated 20d agoView on Hugging Face
0likes123downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
mjbommar/opengloss-v2.0-pretrain · CoolFace