CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sileod /probability_words_nliWords of estimative probability NLI (3 configs). Parquet conversion; CSVs retained. tabular10K<n<100K6 likes592 downloads3d agoHugging Face02Maximax67 /English-Valid-Words English Valid Words This repository contains CSV files with valid English words along with their frequency, stem, and stem valid probability. Dataset Github link: https://github.com/Maximax67/English-Valid-Words Files included valid_words_sorted_alphabetically.csv: N: Counter for each word entry. Word: The English word itself. Frequency count: The number of occurrences of the word in the 1-grams dataset. Stem: The stem of the word. Stem valid probability: Probability… See the full description on the dataset page: https://huggingface.co/datasets/Maximax67/English-Valid-Words.tabular100K<n<1M5 likes191 downloads2y agoHugging Face03nltk-data-hub /words NLTK Word Lists English word lists from NLTK, the New General Service List Project, and Bing Liu's Opinion Lexicon. Configs Config Words Schema License Source en 235,886 word NLTK (other) NLTK words corpus en-basic 850 word Public domain Ogden Basic English (1930) ngsl 2,809 word, rank, sfi, freq_per_million CC-BY-SA 4.0 New General Service List 1.2 toeic 1,250 word, rank, sfi, freq_per_million CC-BY-SA 4.0 TOEIC Service List 1.2 nawl 963 word, rank… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/words.tabulartext-classification100K<n<1M1 likes137 downloads5mo agoHugging Face04wordsum /for-the-small-shield-chapters Foreword The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster. I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.tabulartext-retrieval1K<n<10K0 likes94 downloads1mo agoHugging Face05ronantakizawa /trending-words-google Google Trending Words Dataset (2001-2024) Dataset Description This dataset contains Google trending words and search terms from 2001 to 2024, capturing 24 years of internet culture, major events, and global trends. The dataset includes 2,784 entries across 93 standardized categories, providing a comprehensive view of what captured the world's attention over more than two decades. Dataset Summary Total Entries: 2,784 Years Covered: 2001-2024 (24 years)… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/trending-words-google.tabulartext-classification1K<n<10K4 likes92 downloads10mo agoHugging Face06clearspeed /WordSpoofgated Dataset Card for WordSpoof A public corpus of standalone single-word bona fide and synthetic speech, for research on synthetic-voice detection. Most audio-deepfake corpora are sentence-length; this one is not, and detectors trained on sentence-length audio degrade sharply on it. Dataset Details Dataset Description WordSpoof pairs real single-word utterances with synthetic ones generated by 10 modern text-to-speech and voice-cloning systems, so a… See the full description on the dataset page: https://huggingface.co/datasets/clearspeed/WordSpoof.tabularaudio-classification1K<n<10K1 likes64 downloads10d agoHugging Face07gswamy /pythia-1.4B-tldr-two-words-gpt-4o-iter-1tabular10K<n<100K0 likes48 downloads2y agoHugging Face08PopKornmm /wordscapes-wildlife-animals-stats Wordscapes Wildlife Animals — Complete Stats Complete statistics for all 66 collectible Wildlife animals in Wordscapes (PeopleFun), including the June 2026 update that extended the Blue and Green egg groups to Level 10. Files wordscapes_wildlife_animals.csv — one row per animal: rarity tier, egg groups, activity duration, activities before sleep, sleep hours, max level, bonus type, coin cost per activity, star animal flag. wordscapes_wildlife_levels.csv — one row… See the full description on the dataset page: https://huggingface.co/datasets/PopKornmm/wordscapes-wildlife-animals-stats.tabularn<1K0 likes44 downloads2mo agoHugging Face09ghanaopenai /ghanaian-english-words-corrected-transcriptions This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghanaian English Transcript Corrections A dataset of mistranscribed words and phrases from Ghanaian news media YouTube videos, corrected using Llama 3.1 405B. Source Extracted from YouTube transcripts of Ghanaian… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghanaian-english-words-corrected-transcriptions.tabular100K<n<1M0 likes43 downloads3mo agoHugging Face10gswamy /pythia-1.4B-tldr-two-words-iter-1tabular10K<n<100K0 likes37 downloads2y agoHugging Face11NLie2 /rewrite-questions-real-words-sciency real_words_sciency.csv - Question Rewriting Dataset This dataset contains question rewriting outputs from the file real_words_sciency.csv. Dataset Structure The dataset contains the following columns: custom_id: Unique identifier for each question style: Rewriting style applied (e.g., "gibberish") index: Numerical index original: Original question text rewritten: Rewritten version of the question options: Multiple choice options (list format) correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-real-words-sciency.tabulartext-generationn<1K0 likes32 downloads1y agoHugging Face12ziwenyd /leetcode-standalone-wordstabular1K<n<10K0 likes31 downloads4y agoHugging Face13rcade /amharic_wordstabular10K<n<100K1 likes29 downloads3y agoHugging Face14Mouwiya /image-in-Words400 Mouwiya/image-in-words400 Dataset Description Mouwiya/image-in-words400 is a dataset consisting of 400 images along with their corresponding descriptive captions. The dataset is designed for tasks related to image captioning, where the goal is to generate accurate and contextually relevant descriptions for visual content. This dataset can be used to train and evaluate models that bridge the gap between visual and textual data. Dataset Details Total Examples:… See the full description on the dataset page: https://huggingface.co/datasets/Mouwiya/image-in-Words400.imageimage-to-textn<1K0 likes25 downloads2y agoHugging Face15ronantakizawa /japanese-trending-words Japanese Trending Words Dataset (2006-2025) Dataset Description This dataset comprises Japanese trending words from the official annual Japanese Trending Words Awards (流行語大賞) from 2006 to 2025, documenting the cultural, social, and political phenomena that have shaped Japan over the past two decades. Dataset Summary Total entries: 593 words Time period: 2006-2025 (20 years) Languages: Japanese with English translations Format: CSV with seven columns: word… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-trending-words.tabulartext-generationn<1K4 likes23 downloads10mo agoHugging Face16StephanAkkerman /english-words-human-similaritytabularn<1K0 likes21 downloads2y agoHugging Face17gswamy /pythia-1.4B-tldr-two-words-gpt-4o-reference-traintabular100K<n<1M0 likes17 downloads2y agoHugging Face18GoktugD /turkish-number-words-1m Turkish Number Words 1M v2 0 ile 999.999 arasındaki her tamsayının Türkçe yazıyla karşılığı. Doğrulanmış boyut Train: 980,000 Validation: 10,000 Test: 10,000 Toplam: 1,000,000 Ana görev sütunları: id, number, words Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version, generator_sha256… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-number-words-1m.tabulartext-generation1M<n<10M0 likes16 downloads2mo agoHugging Face19the-french-artist /wikipedia_20220301.simple_sentence_split_text_has_at_least_5_wordstabular1M<n<10M0 likes14 downloads2y agoHugging Face20gswamy /pythia-1.4B-tldr-two-words-gpt-4o-reference-valtabular1K<n<10K0 likes14 downloads2y agoHugging Face21yaneivan /Pari_Chekhov_wordstabularn<1K0 likes12 downloads2y agoHugging Face22ronantakizawa /india-trending-words Google India Trending Words Dataset (2008-2021, 2023-2024) Dataset Description This dataset contains Google trending search terms specific to India from 2008 to 2024 (https://trends.withgoogle.com). Dataset Summary Total Entries: 900 Years Covered: 2008-2009, 2011-2021, 2023-2024 (15 years, 2010 and 2022 data not available) Categories: 18 unique tags Region: India Format: CSV Dataset Structure Data Fields word (string): The trending… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/india-trending-words.tabulartext-classificationn<1K2 likes12 downloads10mo agoHugging Face23azhang42 /one-billion-words-testtabular10K<n<100K0 likes11 downloads2y agoHugging Face24laurabraad /intergrated_gradient_wordstabular1K<n<10K0 likes10 downloads1y agoHugging Face25phillipeds /English-Valid-Words English Valid Words This repository contains CSV files with valid English words along with their frequency, stem, and stem valid probability. Dataset Github link: https://github.com/Maximax67/English-Valid-Words Files included valid_words_sorted_alphabetically.csv: N: Counter for each word entry. Word: The English word itself. Frequency count: The number of occurrences of the word in the 1-grams dataset. Stem: The stem of the word. Stem valid probability: Probability… See the full description on the dataset page: https://huggingface.co/datasets/phillipeds/English-Valid-Words.tabular100K<n<1M0 likes7 downloads7mo agoHugging Face26kossnocorp /wikipedia-words-en-lowtabular1M<n<10M0 likes6 downloads3y agoHugging Face27kossnocorp /wikipedia-words-ru-lowtabular1M<n<10M0 likes5 downloads3y agoHugging Face28TAUR-dev /D-EVAL__letter_countdown__freq_words__all_easy__qwen_base__evaltabularn<1K0 likes3 downloads1y agoHugging Face29TAUR-dev /D-EVAL__letter_countdown__freq_words__all__qwen_base__evaltabularn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.