datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MATH_qCoT_LLMquery_questionasquery_lexicalqueryDatasets from Paper: https://huggingface.co/papers/2505.18405
lexica-stable-diffusion-v1-5
Stable Diffusion Dataset
This is a set of about 80,000 Image-Prompt pairs generated by stable-diffusion-v1-5.
The Prompts come from dataset Stable-Diffusion-Prompts which filtered and extracted from the image finder for Stable Diffusion: "Lexica.art".
lexical_relation_classification[Lexical Relation Classification](https://aclanthology.org/P19-1169/)lexica_dataset
LexicaDataset
LexicaDataset is a large-scale text-to-image prompt dataset shared in [USENIX'24] Prompt Stealing Attacks Against Text-to-Image Generation Models.
It contains 61,467 prompt-image pairs collected from Lexica.
All prompts are curated by real users and images are generated by Stable Diffusion.
Data collection details can be found in the paper.
Data Splits
We randomly sample 80% of a dataset as the training dataset and the rest 20% as the testing dataset.… See the full description on the dataset page: https://huggingface.co/datasets/vera365/lexica_dataset.wordnet-lexical-topology
WordNet Lexical Topology Dataset
Dataset Summary
The WordNet Lexical Topology Dataset provides comprehensive n-gram frequency analysis from multiple sources:
NLTK WordNet: Original Princeton WordNet with 117,659 synsets
HF WordNet: Frequency-weighted definitions from 864,894 entries with cardinality data
Unicode: Character names from 143,041 Unicode codepoints
This dataset preserves sequential information crucial for language modeling and text generation, with over 12… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/wordnet-lexical-topology.anno-lexicallexical_substitutionThe Lexical Substitution Task Test Set comprehends the test set and the gold labels used in the Lexical Substitution Task (https://www.evalita.it/2009/tasks/lexical), organised as part of the EVALITA 2009 evaluation campaign (http://www.evalita.it/2009). The task challenged participants to build systems that could automatically find synonyms for a set of 231 words appearing in different contexts.
The data set contains 1710 sentences extracted from the Italian Syntactic Semantic Treebank (ISST)… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/lexical_substitution.Plosives_and_Non_Lexical_Consonant_Bursts_Preview
Harmonic Frontier Audio -- Plosives and Non-Lexical Consonant Bursts (Preview, v0.95)
A high-fidelity human vocal dataset designed for AI training, speech
research, and articulation-aware voice modeling.
Plosives and Non-Lexical Consonant Bursts (Preview), created by
Harmonic Frontier Audio, provides a compact reference set
demonstrating the quality, formatting, and metadata conventions used in
the Harmonic Frontier Audio Human Vocality Primitives series.
🔎 Summary… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Plosives_and_Non_Lexical_Consonant_Bursts_Preview.LexicalTripletshebrew-lexical-references
Hebrew Lexical Reference Indices
Four structured, Strong's-linked transcriptions of external Hebrew (and one Hebrew↔Greek) lexical
reference sources. These are not our own synonymy judgments — each config faithfully represents
what an established outside source, or an actual historical translation record, already asserts (an
etymological dictionary's own root groupings, a WordNet's own synset membership, five named scholars'
own verified structural analysis, the Septuagint's own… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/hebrew-lexical-references.asia-owid-age-of-electoral-democracy-lexical
Age Of Electoral Democracy Lexical | Asia (Our World in Data)
🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 8,453 observations of Age Of Electoral Democracy Lexical data across 49 Asia countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Age Of Electoral Democracy Lexical
Geographic… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-age-of-electoral-democracy-lexical.asia-owid-political-opposition-lexical
Political Opposition Lexical | Asia (Our World in Data)
🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 8,453 observations of Political Opposition Lexical data across 49 Asia countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Political Opposition Lexical
Geographic coverage
49 Asia… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-political-opposition-lexical.cefr-lexical-balance-dataset-50-50-50
Dataset Card for "cefr-lexical-balance-dataset-50-50-50"
More Information needed
lexic-ai-tutorial-datasetWarhammer-Fantasy-Lexicanum-RAG_v1.12
Warhammer Fantasy Lexicanum - RAG-Optimized Dataset v1.12
Dataset Description
This dataset contains structured information scraped from the Warhammer Fantasy Lexicanum, meticulously cleaned, and processed for Retrieval-Augmented Generation (RAG) applications. It is designed to serve as a comprehensive knowledge base for private, lore-accurate Warhammer Fantasy Roleplay (WFRP) sessions powered by Large Language Models (LLMs).
The primary goal of this dataset is to… See the full description on the dataset page: https://huggingface.co/datasets/s1arsky/Warhammer-Fantasy-Lexicanum-RAG_v1.12.bookmia_lexical_unique_trio_ratio_1.50_adaptive_match_mink_random_7_p0.25_a0.25realec-lexical-alpacaeurope-owid-full-democracy-lexical
Full Democracy Lexical | Europe (Our World in Data)
🇪🇺 7,863 observations · 44 Europe countries · 1789–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 7,863 observations of Full Democracy Lexical data across 44 Europe countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Full Democracy Lexical
Geographic coverage
44 Europe… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-full-democracy-lexical.VALUE_wikitext103_lexical
Dataset Card for "VALUE_wikitext103_lexical"
More Information needed
anno-lexical-coresetasia-owid-full-democracy-lexical
Full Democracy Lexical | Asia (Our World in Data)
🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 8,453 observations of Full Democracy Lexical data across 49 Asia countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Full Democracy Lexical
Geographic coverage
49 Asia countries · top… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-full-democracy-lexical.olympiads_paraphrased_lexical_unique_trio_ratio_2.0_adaptive_match_minkplus_random_7_p0.25asia-owid-universal-suffrage-lexical
Universal Suffrage Lexical | Asia (Our World in Data)
🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 8,453 observations of Universal Suffrage Lexical data across 49 Asia countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Universal Suffrage Lexical
Geographic coverage
49 Asia… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-universal-suffrage-lexical.kazakh-lexical-complexity-classes
Kazakh Lexical Complexity Classes
A CEFR-graded lexical resource for the Kazakh language. The lexicon contains 4,561 lemma–POS entries graded across five CEFR proficiency levels.
Data Format
The dataset is provided as a single JSON file. Each entry has the following fields:
Field
Type
Description
lemma
string
Kazakh word (Cyrillic script)
pos
string
Part of speech (NOUN, VERB, ADJ, ADV, NUM, PRON, OTHER, etc.)
cefr
string
CEFR proficiency level (A1, A2, B1… See the full description on the dataset page: https://huggingface.co/datasets/Gulnur7/kazakh-lexical-complexity-classes.chew_lexical
Dataset Card for Dataset Name
This is the lexical/no-overlapping split of the CHEW dataset(CHEW: A Dataset of CHanging Events in Wikipedia).
Dataset Details
Dataset Description
This dataset is the Lexical/No-overlapping split of the CHEW Dataset,where CHEW stands for CHanging Events in Wikipedia. It contains Wikipedia titles, text in two timestamped versions and Binary Label showing Change(1) or No change(0). Change here means there has been informationm… See the full description on the dataset page: https://huggingface.co/datasets/hsuvaskakoty/chew_lexical.acereason_ge15_lexical_v2_100000_diverselexical-decisionThis dataset contains words/sentences for lexical decision tests, which we created with wuggy.
If you use this dataset, please cite the following preprint:
@misc{bunzeck2025subwordmodelsstruggleword,
title={Subword models struggle with word learning, but surprisal hides it},
author={Bastian Bunzeck and Sina Zarrieß},
year={2025},
eprint={2502.12835},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2502.12835},
}
asia-owid-democracy-lexical
Democracy Lexical | Asia (Our World in Data)
🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 8,453 observations of Democracy Lexical data across 49 Asia countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Democracy Lexical
Geographic coverage
49 Asia countries · top rows shown below… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-democracy-lexical.europe-owid-universal-suffrage-women-lexical
Universal Suffrage Women Lexical | Europe (Our World in Data)
🇪🇺 7,863 observations · 44 Europe countries · 1789–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 7,863 observations of Universal Suffrage Women Lexical data across 44 Europe countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Universal Suffrage Women Lexical
Geographic… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-universal-suffrage-women-lexical.chinese-lexical-normalization
chinese-lexical-normalization
This dataset contains informal-formal-explanation triples from the chinese-lexical-normalization dataset. Note that there are duplicate informal-formal pairs due to multiple explanations.
Example usage:
from datasets import load_dataset
dataset = load_dataset("larrylawl/chinese-lexical-normalization")
