mjbommar/opengloss-v2.0-pretrain
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Pretrain The release rendered as continuous prose for language-model pretraining or continued pretraining: four document templates per entry — a dictionary entry, a thesaurus entry, an encyclopedia article and a usage note — written as plain text… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-pretrain.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face