datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MATH_qCoT_LLMquery_questionasquery_lexicalqueryDatasets from Paper: https://huggingface.co/papers/2505.18405
lexica-stable-diffusion-v1-5
Stable Diffusion Dataset
This is a set of about 80,000 Image-Prompt pairs generated by stable-diffusion-v1-5.
The Prompts come from dataset Stable-Diffusion-Prompts which filtered and extracted from the image finder for Stable Diffusion: "Lexica.art".
lexical_stress_dataset@article{allouche2026does,
title={How does a deep neural network look at lexical stress in English words?},
author={Allouche, Itai and Asael, Itay and Rousso, Rotem and Dassa, Vered and Bradlow, Ann and Kim, Seung-Eun and Goldrick, Matthew and Keshet, Joseph},
journal={The Journal of the Acoustical Society of America},
volume={159},
number={2},
pages={1348--1358},
year={2026},
publisher={AIP Publishing}
}
lexical_relation_classification[Lexical Relation Classification](https://aclanthology.org/P19-1169/)lexica_dataset
LexicaDataset
LexicaDataset is a large-scale text-to-image prompt dataset shared in [USENIX'24] Prompt Stealing Attacks Against Text-to-Image Generation Models.
It contains 61,467 prompt-image pairs collected from Lexica.
All prompts are curated by real users and images are generated by Stable Diffusion.
Data collection details can be found in the paper.
Data Splits
We randomly sample 80% of a dataset as the training dataset and the rest 20% as the testing dataset.… See the full description on the dataset page: https://huggingface.co/datasets/vera365/lexica_dataset.wordnet-lexical-topology
WordNet Lexical Topology Dataset
Dataset Summary
The WordNet Lexical Topology Dataset provides comprehensive n-gram frequency analysis from multiple sources:
NLTK WordNet: Original Princeton WordNet with 117,659 synsets
HF WordNet: Frequency-weighted definitions from 864,894 entries with cardinality data
Unicode: Character names from 143,041 Unicode codepoints
This dataset preserves sequential information crucial for language modeling and text generation, with over 12… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/wordnet-lexical-topology.anno-lexicallexical_substitutionThe Lexical Substitution Task Test Set comprehends the test set and the gold labels used in the Lexical Substitution Task (https://www.evalita.it/2009/tasks/lexical), organised as part of the EVALITA 2009 evaluation campaign (http://www.evalita.it/2009). The task challenged participants to build systems that could automatically find synonyms for a set of 231 words appearing in different contexts.
The data set contains 1710 sentences extracted from the Italian Syntactic Semantic Treebank (ISST)… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/lexical_substitution.Plosives_and_Non_Lexical_Consonant_Bursts_Preview
Harmonic Frontier Audio -- Plosives and Non-Lexical Consonant Bursts (Preview, v0.95)
A high-fidelity human vocal dataset designed for AI training, speech
research, and articulation-aware voice modeling.
Plosives and Non-Lexical Consonant Bursts (Preview), created by
Harmonic Frontier Audio, provides a compact reference set
demonstrating the quality, formatting, and metadata conventions used in
the Harmonic Frontier Audio Human Vocality Primitives series.
🔎 Summary… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Plosives_and_Non_Lexical_Consonant_Bursts_Preview.dfm10-danish-lexical-sentiment-sft
dfm10-danish-lexical-sentiment-sft
Gold lexical-polarity supervision derived from the Danish Sentiment Lexicon.
Contents
Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz
Schema: chat messages, optional condition and tools, plus provenance
Shards: 1
Rows: 13,698
Category: Danish lexical sentiment
Upstream material
dsldk/danish-sentiment-lexicon
Processing
Gold batched mappings are supplemented by separately generated… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-danish-lexical-sentiment-sft.LexicalTripletslexicapLexicap contains the captions for every Lex Friedman Podcast episode. It it created by [Dr. Andrej Karpathy](https://twitter.com/karpathy).
There are 430 caption files available. There are 2 types of files:
- large
- small
Each file name follows the format `episode_{episode_number}_{file_type}.vtt`.hebrew-lexical-references
Hebrew Lexical Reference Indices
Four structured, Strong's-linked transcriptions of external Hebrew (and one Hebrew↔Greek) lexical
reference sources. These are not our own synonymy judgments — each config faithfully represents
what an established outside source, or an actual historical translation record, already asserts (an
etymological dictionary's own root groupings, a WordNet's own synset membership, five named scholars'
own verified structural analysis, the Septuagint's own… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/hebrew-lexical-references.asia-owid-age-of-electoral-democracy-lexical
Age Of Electoral Democracy Lexical | Asia (Our World in Data)
🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 8,453 observations of Age Of Electoral Democracy Lexical data across 49 Asia countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Age Of Electoral Democracy Lexical
Geographic… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-age-of-electoral-democracy-lexical.asia-owid-political-opposition-lexical
Political Opposition Lexical | Asia (Our World in Data)
🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 8,453 observations of Political Opposition Lexical data across 49 Asia countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Political Opposition Lexical
Geographic coverage
49 Asia… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-political-opposition-lexical.cefr-lexical-balance-dataset-50-50-50
Dataset Card for "cefr-lexical-balance-dataset-50-50-50"
More Information needed
lexic-ai-tutorial-datasetWarhammer-Fantasy-Lexicanum-RAG_v1.12
Warhammer Fantasy Lexicanum - RAG-Optimized Dataset v1.12
Dataset Description
This dataset contains structured information scraped from the Warhammer Fantasy Lexicanum, meticulously cleaned, and processed for Retrieval-Augmented Generation (RAG) applications. It is designed to serve as a comprehensive knowledge base for private, lore-accurate Warhammer Fantasy Roleplay (WFRP) sessions powered by Large Language Models (LLMs).
The primary goal of this dataset is to… See the full description on the dataset page: https://huggingface.co/datasets/s1arsky/Warhammer-Fantasy-Lexicanum-RAG_v1.12.bookmia_lexical_unique_trio_ratio_1.50_adaptive_match_mink_random_7_p0.25_a0.25realec-lexical-alpacaeurope-owid-full-democracy-lexical
Full Democracy Lexical | Europe (Our World in Data)
🇪🇺 7,863 observations · 44 Europe countries · 1789–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 7,863 observations of Full Democracy Lexical data across 44 Europe countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Full Democracy Lexical
Geographic coverage
44 Europe… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-full-democracy-lexical.VALUE_wikitext103_lexical
Dataset Card for "VALUE_wikitext103_lexical"
More Information needed
anno-lexical-coresetasia-owid-full-democracy-lexical
Full Democracy Lexical | Asia (Our World in Data)
🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 8,453 observations of Full Democracy Lexical data across 49 Asia countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Full Democracy Lexical
Geographic coverage
49 Asia countries · top… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-full-democracy-lexical.lexical-system-fr
[!NOTE]
Dataset origin: https://www.ortolang.fr/market/lexicons/lexical-system-fr
Description
Caractérisation du Réseau Lexical du Français (RL-fr)
Le Réseau Lexical du Français (RL-fr) est un modèle formel du lexique du français contemporain, en cours de construction au laboratoire ATILF du CNRS. Il possède les trois particularités suivantes :
Le RL-fr est, formellement, un [Système Lexical(https://lexical-systems.atilf.fr/) [Polguère 2014], c'est-à-dire une modélisation sous… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/lexical-system-fr.olympiads_paraphrased_lexical_unique_trio_ratio_2.0_adaptive_match_minkplus_random_7_p0.25asia-owid-universal-suffrage-lexical
Universal Suffrage Lexical | Asia (Our World in Data)
🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 8,453 observations of Universal Suffrage Lexical data across 49 Asia countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Universal Suffrage Lexical
Geographic coverage
49 Asia… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-universal-suffrage-lexical.kazakh-lexical-complexity-classes
Kazakh Lexical Complexity Classes
A CEFR-graded lexical resource for the Kazakh language. The lexicon contains 4,561 lemma–POS entries graded across five CEFR proficiency levels.
Data Format
The dataset is provided as a single JSON file. Each entry has the following fields:
Field
Type
Description
lemma
string
Kazakh word (Cyrillic script)
pos
string
Part of speech (NOUN, VERB, ADJ, ADV, NUM, PRON, OTHER, etc.)
cefr
string
CEFR proficiency level (A1, A2, B1… See the full description on the dataset page: https://huggingface.co/datasets/Gulnur7/kazakh-lexical-complexity-classes.chew_lexical
Dataset Card for Dataset Name
This is the lexical/no-overlapping split of the CHEW dataset(CHEW: A Dataset of CHanging Events in Wikipedia).
Dataset Details
Dataset Description
This dataset is the Lexical/No-overlapping split of the CHEW Dataset,where CHEW stands for CHanging Events in Wikipedia. It contains Wikipedia titles, text in two timestamped versions and Binary Label showing Change(1) or No change(0). Change here means there has been informationm… See the full description on the dataset page: https://huggingface.co/datasets/hsuvaskakoty/chew_lexical.acereason_ge15_lexical_v2_100000_diverse
