CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Raderspace /MATH_qCoT_LLMquery_questionasquery_lexicalqueryDatasets from Paper: https://huggingface.co/papers/2505.18405 text10K<n<100K2 likes986 downloads1y agoHugging Face02yuwan0 /lexica-stable-diffusion-v1-5 Stable Diffusion Dataset This is a set of about 80,000 Image-Prompt pairs generated by stable-diffusion-v1-5. The Prompts come from dataset Stable-Diffusion-Prompts which filtered and extracted from the image finder for Stable Diffusion: "Lexica.art". image10K<n<100K4 likes430 downloads3y agoHugging Face03MLSpeech /lexical_stress_dataset@article{allouche2026does, title={How does a deep neural network look at lexical stress in English words?}, author={Allouche, Itai and Asael, Itay and Rousso, Rotem and Dassa, Vered and Bradlow, Ann and Kim, Seung-Eun and Goldrick, Matthew and Keshet, Joseph}, journal={The Journal of the Acoustical Society of America}, volume={159}, number={2}, pages={1348--1358}, year={2026}, publisher={AIP Publishing} } 10K<n<100K0 likes340 downloads13d agoHugging Face04relbert /lexical_relation_classification[Lexical Relation Classification](https://aclanthology.org/P19-1169/)text100K<n<1M3 likes163 downloads4y agoHugging Face05vera365 /lexica_dataset LexicaDataset LexicaDataset is a large-scale text-to-image prompt dataset shared in [USENIX'24] Prompt Stealing Attacks Against Text-to-Image Generation Models. It contains 61,467 prompt-image pairs collected from Lexica. All prompts are curated by real users and images are generated by Stable Diffusion. Data collection details can be found in the paper. Data Splits We randomly sample 80% of a dataset as the training dataset and the rest 20% as the testing dataset.… See the full description on the dataset page: https://huggingface.co/datasets/vera365/lexica_dataset.imagetext-to-image10K<n<100K8 likes122 downloads2y agoHugging Face06AbstractPhil /wordnet-lexical-topology WordNet Lexical Topology Dataset Dataset Summary The WordNet Lexical Topology Dataset provides comprehensive n-gram frequency analysis from multiple sources: NLTK WordNet: Original Princeton WordNet with 117,659 synsets HF WordNet: Frequency-weighted definitions from 864,894 entries with cardinality data Unicode: Character names from 143,041 Unicode codepoints This dataset preserves sequential information crucial for language modeling and text generation, with over 12… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/wordnet-lexical-topology.tabulartext-generation10M<n<100M2 likes113 downloads1y agoHugging Face07mediabiasgroup /anno-lexicaltext10K<n<100K1 likes78 downloads7mo agoHugging Face08evalitahf /lexical_substitutionThe Lexical Substitution Task Test Set comprehends the test set and the gold labels used in the Lexical Substitution Task (https://www.evalita.it/2009/tasks/lexical), organised as part of the EVALITA 2009 evaluation campaign (http://www.evalita.it/2009). The task challenged participants to build systems that could automatically find synonyms for a set of 231 words appearing in different contexts. The data set contains 1710 sentences extracted from the Italian Syntactic Semantic Treebank (ISST)… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/lexical_substitution.texttext-generation1K<n<10K0 likes73 downloads2y agoHugging Face09Harmonic-Frontier-Audio /Plosives_and_Non_Lexical_Consonant_Bursts_Preview Harmonic Frontier Audio -- Plosives and Non-Lexical Consonant Bursts (Preview, v0.95) A high-fidelity human vocal dataset designed for AI training, speech research, and articulation-aware voice modeling. Plosives and Non-Lexical Consonant Bursts (Preview), created by Harmonic Frontier Audio, provides a compact reference set demonstrating the quality, formatting, and metadata conventions used in the Harmonic Frontier Audio Human Vocality Primitives series. 🔎 Summary… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Plosives_and_Non_Lexical_Consonant_Bursts_Preview.audioothern<1K2 likes73 downloads7mo agoHugging Face10schneiderkamplab /dfm10-danish-lexical-sentiment-sft dfm10-danish-lexical-sentiment-sft Gold lexical-polarity supervision derived from the Danish Sentiment Lexicon. Contents Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz Schema: chat messages, optional condition and tools, plus provenance Shards: 1 Rows: 13,698 Category: Danish lexical sentiment Upstream material dsldk/danish-sentiment-lexicon Processing Gold batched mappings are supplemented by separately generated… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-danish-lexical-sentiment-sft.0 likes67 downloads27d agoHugging Face11cnmoro /LexicalTripletstext10M<n<100M0 likes59 downloads10mo agoHugging Face12shubhamg2208 /lexicapLexicap contains the captions for every Lex Friedman Podcast episode. It it created by [Dr. Andrej Karpathy](https://twitter.com/karpathy). There are 430 caption files available. There are 2 types of files: - large - small Each file name follows the format `episode_{episode_number}_{file_type}.vtt`.text-classificationn<1K0 likes47 downloads4y agoHugging Face13bcv-commons /hebrew-lexical-references Hebrew Lexical Reference Indices Four structured, Strong's-linked transcriptions of external Hebrew (and one Hebrew↔Greek) lexical reference sources. These are not our own synonymy judgments — each config faithfully represents what an established outside source, or an actual historical translation record, already asserts (an etymological dictionary's own root groupings, a WordNet's own synset membership, five named scholars' own verified structural analysis, the Septuagint's own… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/hebrew-lexical-references.text1K<n<10K0 likes45 downloads2mo agoHugging Face14electricsheepasia /asia-owid-age-of-electoral-democracy-lexical Age Of Electoral Democracy Lexical | Asia (Our World in Data) 🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia TL;DR This dataset contains 8,453 observations of Age Of Electoral Democracy Lexical data across 49 Asia countries, spanning 1789–2025. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Age Of Electoral Democracy Lexical Geographic… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-age-of-electoral-democracy-lexical.texttabular-classification1K<n<10K0 likes44 downloads4mo agoHugging Face15electricsheepasia /asia-owid-political-opposition-lexical Political Opposition Lexical | Asia (Our World in Data) 🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia TL;DR This dataset contains 8,453 observations of Political Opposition Lexical data across 49 Asia countries, spanning 1789–2025. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Political Opposition Lexical Geographic coverage 49 Asia… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-political-opposition-lexical.tabulartabular-classification1K<n<10K0 likes44 downloads4mo agoHugging Face16hafidikhsan /cefr-lexical-balance-dataset-50-50-50 Dataset Card for "cefr-lexical-balance-dataset-50-50-50" More Information needed text1K<n<10K1 likes42 downloads3y agoHugging Face17smartloop-ai /lexic-ai-tutorial-datasettextquestion-answeringn<1K0 likes41 downloads2y agoHugging Face18s1arsky /Warhammer-Fantasy-Lexicanum-RAG_v1.12 Warhammer Fantasy Lexicanum - RAG-Optimized Dataset v1.12 Dataset Description This dataset contains structured information scraped from the Warhammer Fantasy Lexicanum, meticulously cleaned, and processed for Retrieval-Augmented Generation (RAG) applications. It is designed to serve as a comprehensive knowledge base for private, lore-accurate Warhammer Fantasy Roleplay (WFRP) sessions powered by Large Language Models (LLMs). The primary goal of this dataset is to… See the full description on the dataset page: https://huggingface.co/datasets/s1arsky/Warhammer-Fantasy-Lexicanum-RAG_v1.12.text10K<n<100K2 likes38 downloads7mo agoHugging Face19ADRA-RL /bookmia_lexical_unique_trio_ratio_1.50_adaptive_match_mink_random_7_p0.25_a0.25textn<1K0 likes34 downloads7mo agoHugging Face20rhythm00 /realec-lexical-alpacatext1K<n<10K0 likes32 downloads1y agoHugging Face21electricsheepeurope /europe-owid-full-democracy-lexical Full Democracy Lexical | Europe (Our World in Data) 🇪🇺 7,863 observations · 44 Europe countries · 1789–2025 · Repackaged by Electric Sheep Europe TL;DR This dataset contains 7,863 observations of Full Democracy Lexical data across 44 Europe countries, spanning 1789–2025. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Full Democracy Lexical Geographic coverage 44 Europe… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-full-democracy-lexical.tabulartabular-classification1K<n<10K0 likes32 downloads4mo agoHugging Face22liuyanchen1015 /VALUE_wikitext103_lexical Dataset Card for "VALUE_wikitext103_lexical" More Information needed text100K<n<1M0 likes27 downloads4y agoHugging Face23mediabiasgroup /anno-lexical-coresettexttext-classification1K<n<10K0 likes27 downloads2y agoHugging Face24electricsheepasia /asia-owid-full-democracy-lexical Full Democracy Lexical | Asia (Our World in Data) 🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia TL;DR This dataset contains 8,453 observations of Full Democracy Lexical data across 49 Asia countries, spanning 1789–2025. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Full Democracy Lexical Geographic coverage 49 Asia countries · top… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-full-democracy-lexical.tabulartabular-classification1K<n<10K0 likes27 downloads4mo agoHugging Face25datasets-CNRS /lexical-system-fr [!NOTE] Dataset origin: https://www.ortolang.fr/market/lexicons/lexical-system-fr Description Caractérisation du Réseau Lexical du Français (RL-fr) Le Réseau Lexical du Français (RL-fr) est un modèle formel du lexique du français contemporain, en cours de construction au laboratoire ATILF du CNRS. Il possède les trois particularités suivantes : Le RL-fr est, formellement, un [Système Lexical(https://lexical-systems.atilf.fr/) [Polguère 2014], c'est-à-dire une modélisation sous… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/lexical-system-fr.0 likes26 downloads1y agoHugging Face26ADRA-RL /olympiads_paraphrased_lexical_unique_trio_ratio_2.0_adaptive_match_minkplus_random_7_p0.25textn<1K0 likes26 downloads7mo agoHugging Face27electricsheepasia /asia-owid-universal-suffrage-lexical Universal Suffrage Lexical | Asia (Our World in Data) 🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia TL;DR This dataset contains 8,453 observations of Universal Suffrage Lexical data across 49 Asia countries, spanning 1789–2025. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Universal Suffrage Lexical Geographic coverage 49 Asia… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-universal-suffrage-lexical.tabulartabular-classification1K<n<10K0 likes25 downloads4mo agoHugging Face28Gulnur7 /kazakh-lexical-complexity-classes Kazakh Lexical Complexity Classes A CEFR-graded lexical resource for the Kazakh language. The lexicon contains 4,561 lemma–POS entries graded across five CEFR proficiency levels. Data Format The dataset is provided as a single JSON file. Each entry has the following fields: Field Type Description lemma string Kazakh word (Cyrillic script) pos string Part of speech (NOUN, VERB, ADJ, ADV, NUM, PRON, OTHER, etc.) cefr string CEFR proficiency level (A1, A2, B1… See the full description on the dataset page: https://huggingface.co/datasets/Gulnur7/kazakh-lexical-complexity-classes.texttext-classification1K<n<10K1 likes23 downloads6mo agoHugging Face29hsuvaskakoty /chew_lexical Dataset Card for Dataset Name This is the lexical/no-overlapping split of the CHEW dataset(CHEW: A Dataset of CHanging Events in Wikipedia). Dataset Details Dataset Description This dataset is the Lexical/No-overlapping split of the CHEW Dataset,where CHEW stands for CHanging Events in Wikipedia. It contains Wikipedia titles, text in two timestamped versions and Binary Label showing Change(1) or No change(0). Change here means there has been informationm… See the full description on the dataset page: https://huggingface.co/datasets/hsuvaskakoty/chew_lexical.texttext-classification1K<n<10K0 likes22 downloads2y agoHugging Face30shizhuo2 /acereason_ge15_lexical_v2_100000_diversetext10K<n<100K0 likes22 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.