datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opengloss-v2.0-encyclopedia
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Encyclopedia
The long-form entry-level prose of OpenGloss v2.0, one row per rendition. The encyclopedia config holds the 300–500-word article about each headword, written at up to five reading levels; the explanation config holds the shorter "why… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-encyclopedia.opengloss-v2.0-definitions
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Definitions
The flat definition view: one row for every stored rendition of every live sense's definition, the canonical (neutral, plain) gloss included. This is the reading-level and register grading of OpenGloss v2.0 laid out one row at a time… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-definitions.opengloss-v2.0-lexicon
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Lexicon
The entry-level view of OpenGloss v2.0: one row per lexeme, with everything that belongs to the entry rather than to one of its meanings — the kind discriminator, per-POS morphology, structured etymology, the lexical explanation, the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-lexicon.opengloss-v2.0-senses
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Senses
The sense-level view of OpenGloss v2.0 and the repo most consumers want: one row per live sense, with its canonical gloss, its eight reading-level and register renditions, its sense-tagged example sentences with headword character spans, its… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-senses.AdParaphrase-v2.0
AdParaphrase v2.0
This repository contains data for our paper "AdParaphrase v2.0: Generating Attractive Ad Texts Using a Preference-Annotated Paraphrase Dataset" (ACL2025 Findings).
Overview
AdParaphrase v2.0 is a dataset for ad text paraphrasing, containing human preference data, to enable the analysis of the linguistic factors and to support the development of methods for generating attractive ad texts. Compared with AdParaphrase v1.0, this dataset is 20 times larger… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/AdParaphrase-v2.0.
