datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hebrew-lexical-references
Hebrew Lexical Reference Indices
Four structured, Strong's-linked transcriptions of external Hebrew (and one Hebrew↔Greek) lexical
reference sources. These are not our own synonymy judgments — each config faithfully represents
what an established outside source, or an actual historical translation record, already asserts (an
etymological dictionary's own root groupings, a WordNet's own synset membership, five named scholars'
own verified structural analysis, the Septuagint's own… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/hebrew-lexical-references.chew_lexical
Dataset Card for Dataset Name
This is the lexical/no-overlapping split of the CHEW dataset(CHEW: A Dataset of CHanging Events in Wikipedia).
Dataset Details
Dataset Description
This dataset is the Lexical/No-overlapping split of the CHEW Dataset,where CHEW stands for CHanging Events in Wikipedia. It contains Wikipedia titles, text in two timestamped versions and Binary Label showing Change(1) or No change(0). Change here means there has been informationm… See the full description on the dataset page: https://huggingface.co/datasets/hsuvaskakoty/chew_lexical.lexical-decisionThis dataset contains words/sentences for lexical decision tests, which we created with wuggy.
If you use this dataset, please cite the following preprint:
@misc{bunzeck2025subwordmodelsstruggleword,
title={Subword models struggle with word learning, but surprisal hides it},
author={Bastian Bunzeck and Sina Zarrieß},
year={2025},
eprint={2502.12835},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2502.12835},
}
chinese-lexical-normalization
chinese-lexical-normalization
This dataset contains informal-formal-explanation triples from the chinese-lexical-normalization dataset. Note that there are duplicate informal-formal pairs due to multiple explanations.
Example usage:
from datasets import load_dataset
dataset = load_dataset("larrylawl/chinese-lexical-normalization")
lemi_lexical_lists
Romanian Grade Word Lists
Dataset Description
This dataset contains word lists grouped by grade level (1–4), derived from model-based difficulty predictions across multiple Romanian texts. Each word is associated with a mean grade score indicating its predicted difficulty level.
Dataset Structure
Each file corresponds to a grade level and contains the following columns:
cuvant: the Romanian word (token)
grad_mediu: mean predicted grade level for that word… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/lemi_lexical_lists.lexical-fieldlexical-decision-germanThis dataset contains German words/sentences for lexical decision tests, which we created with wuggy.
If you use this dataset, please cite the following preprint:
If you use this dataset, please cite the following publication:
@inproceedings{bunzeck-etal-2025-construction,
title = "Do Construction Distributions Shape Formal Language Learning In {G}erman {B}aby{LM}s?",
author = "Bunzeck, Bastian and
Duran, Daniel and
Zarrie{\ss}, Sina",
editor = "Boleda, Gemma and… See the full description on the dataset page: https://huggingface.co/datasets/bbunzeck/lexical-decision-german.lexical-dataset-50-lexical-fieldslexical-fields-with-tokenslexical-frames
language:
en
---+
This repository contains the lexical frames used to prompt the models in the frame completion task we introduced in Child-directed speech facilitates production, not comprehension, in BabyLMs (Bunzeck & Zarrieß @ CoNLL 2026).
If you use this data, please cite:
@inproceedings{bunzeck-zarriess-2026-child,
title = "Child-directed speech facilitates production, not comprehension, in {B}aby{LM}s",
author = "Bunzeck, Bastian and
Zarrie{\ss}, Sina",
editor =… See the full description on the dataset page: https://huggingface.co/datasets/bbunzeck/lexical-frames.lexical_field_normalizedcustomerservice_qa_semantic_lexical_results
