lexica
Datasets
All datasets matching “lexica”MATH_qCoT_LLMquery_questionasquery_lexicalqueryDatasets from Paper: https://huggingface.co/papers/2505.18405
lexica-stable-diffusion-v1-5
Stable Diffusion Dataset
This is a set of about 80,000 Image-Prompt pairs generated by stable-diffusion-v1-5.
The Prompts come from dataset Stable-Diffusion-Prompts which filtered and extracted from the image finder for Stable Diffusion: "Lexica.art".
lexical_stress_dataset@article{allouche2026does,
title={How does a deep neural network look at lexical stress in English words?},
author={Allouche, Itai and Asael, Itay and Rousso, Rotem and Dassa, Vered and Bradlow, Ann and Kim, Seung-Eun and Goldrick, Matthew and Keshet, Joseph},
journal={The Journal of the Acoustical Society of America},
volume={159},
number={2},
pages={1348--1358},
year={2026},
publisher={AIP Publishing}
}
lexical_relation_classification[Lexical Relation Classification](https://aclanthology.org/P19-1169/)lexica_dataset
LexicaDataset
LexicaDataset is a large-scale text-to-image prompt dataset shared in [USENIX'24] Prompt Stealing Attacks Against Text-to-Image Generation Models.
It contains 61,467 prompt-image pairs collected from Lexica.
All prompts are curated by real users and images are generated by Stable Diffusion.
Data collection details can be found in the paper.
Data Splits
We randomly sample 80% of a dataset as the training dataset and the rest 20% as the testing dataset.… See the full description on the dataset page: https://huggingface.co/datasets/vera365/lexica_dataset.wordnet-lexical-topology
WordNet Lexical Topology Dataset
Dataset Summary
The WordNet Lexical Topology Dataset provides comprehensive n-gram frequency analysis from multiple sources:
NLTK WordNet: Original Princeton WordNet with 117,659 synsets
HF WordNet: Frequency-weighted definitions from 864,894 entries with cardinality data
Unicode: Character names from 143,041 Unicode codepoints
This dataset preserves sequential information crucial for language modeling and text generation, with over 12… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/wordnet-lexical-topology.
