CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MongoDB /english-words-definitions English Words Definitions This dataset contains definitions and important facts about 467k words that appear in the context of English texts. It has been used to train our high-performance, compact text embedding models mdbr-leaf-ir and mdbr-leaf-mt. The original list of words stems from here. We have extended it with definitions and important facts about each word using Claude 3.7 Sonnet. textfeature-extraction100K<n<1M4 likes173 downloads1y agoHugging Face02hmar-heritage-org /numeral-wordsgated numeral-words A 3-way parallel digital dataset containing 999,999 spelled-out Hmar & English number words mapped in sequential numerical order (1 to 999,999). Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev). Overview Languages: Hmar (hmr, ISO 639-3, Glottolog: hmar1241), English (en) Family: Zo Languages Volume: 999,999 parallel rows (1 to 999,999) Format: Compressed JSONL (data/train-*.jsonl.gz) License: Apache-2.0 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/numeral-words.texttranslation100K<n<1M0 likes129 downloads10d agoHugging Face03ameer4wisam /iraqi_words_finetuning Iraqi Words A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a dependency-free BM25 retriever and a fine-tuning data generator built on top of it. Why this exists Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA) and higher-resource dialects such as Egyptian or Levantine. Lexical resources that map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.texttranslationn<1K0 likes111 downloads2mo agoHugging Face04meilisearch /stop_wordstext10K<n<100K0 likes95 downloads9mo agoHugging Face05wordsum /for-the-small-shield-chapters Foreword The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster. I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.tabulartext-retrieval1K<n<10K0 likes94 downloads1mo agoHugging Face06lysandre /anki-words anki:Kaishi 1.5k Vocabulary Dataset Japanese vocabulary dataset exported from Anki deck. Dataset Information Deck Name: anki:Kaishi 1.5k Word Count: 478 Format: JSONL (one word per line) Last Updated: 2026-02-21 Schema Each word entry contains: word (string): The Japanese word in kanji/kana reading (string): The reading in hiragana/katakana meaning (string): English translation/meaning jlpt: JLPT level (N5, N4, N3, N2, N1) N5: Basic level (easiest) N4:… See the full description on the dataset page: https://huggingface.co/datasets/lysandre/anki-words.textn<1K0 likes82 downloads7mo agoHugging Face07youngermax /digital-sat-words-in-context-llmtextn<1K0 likes55 downloads1y agoHugging Face08MartialTerran /Eval_Counting_Letters_in_WordsLetters in Words Evaluation Dataset "The strawberry question is pretty much the new Turing Test for future AI" BlakeSergin OP 3mo agohttps://www.reddit.com/r/singularity/comments/1enqk04/how_many_rs_in_strawberry_why_is_this_a_very/ This dataset .json provides a simple yet effective way to assess the basic letter-counting abilities of Large Language Models (LLMs). (Try it on the new SmolLM2 models.) It consists of a set of questions designed to evaluate an LLM's capacity for: Understanding… See the full description on the dataset page: https://huggingface.co/datasets/MartialTerran/Eval_Counting_Letters_in_Words.textn<1K0 likes53 downloads2y agoHugging Face09much1na /english-words-definitions English Words Definitions This dataset contains definitions and important facts about 467k words that appear in the context of English texts. It has been used to train our high-performance, compact text embedding models mdbr-leaf-ir and mdbr-leaf-mt. The original list of words stems from here. We have extended it with definitions and important facts about each word using Claude 3.7 Sonnet. textfeature-extraction100K<n<1M0 likes37 downloads6mo agoHugging Face10NaNoBotCo /thai-trade-words-chiang-mai คำบนป้าย — Thai trade words of Chiang Mai and Chiang Rai 1,782 Thai trade terms taken from the tags on business listings in Chiang Mai and Chiang Rai — the words the city uses for what a shop does — each with an RTGS reading and an English gloss. Plus 403 administrative place names (tambon, amphoe, city, province) with their readings. ซ่อมมอเตอร์ไซค์ Som Motoesai motorcycle repair 191 places ตู้น้ำดื่มหยอดเหรียญ Tunam Duem Yotrian coin-op drinking-water… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-trade-words-chiang-mai.text1K<n<10K0 likes36 downloads17d agoHugging Face11kth8 /english-words-wordnet English Words WordNet Dataset This dataset contains a comprehensive collection of English words paired with their detailed definitions from WordNet. Includes words between 2 and 20 characters in length. Each line represents a single record containing the following structure: { "word": "zwiebacks", "index_0": { "pos": "n", "definition":"slice of sweet raised bread baked again until it is brown and hard and crisp", "examples": [] } } text100K<n<1M0 likes32 downloads6mo agoHugging Face12Arabicc /arabizi_toxic_harassment_words license: other task_categories: - text-classification task_ids: - hate-speech-detection language: - ar tags: - arabizi - dialectal-arabic - toxicity - harassment - social-media-safety - red-teaming size_categories: - n<1K 🛡️ Pro-Grade Arabizi 250- Sentences - Safety & Moderation Dataset Purchase Full Access 👉 Buy the Complete Dataset on Gumroad Arabizi Toxicity & Content Moderation Dataset A dataset containing Arabizi (3araby) toxic, abusive, and… See the full description on the dataset page: https://huggingface.co/datasets/Arabicc/arabizi_toxic_harassment_words.texttext-classificationn<1K0 likes32 downloads1mo agoHugging Face13Rustem-Kaimolla /kazakh-swear-words Kazakh Swear Words Dataset 🇰🇿 Dataset of Kazakh obscene and profane expressions for NLP tasks including text classification, content moderation, toxicity detection, and LLM fine-tuning. Dataset Description This is a low-resource language dataset containing Kazakh profanity, swear words, and offensive expressions along with neutral examples for binary classification tasks. Languages Kazakh (kk) Dataset Structure Data Files data.jsonl -… See the full description on the dataset page: https://huggingface.co/datasets/Rustem-Kaimolla/kazakh-swear-words.texttext-classificationn<1K0 likes28 downloads9mo agoHugging Face14ahmetkaansever /turkish-rare-wordstextn<1K0 likes27 downloads2y agoHugging Face15proshady2 /english-words-definitions English Words Definitions This dataset contains definitions and important facts about 467k words that appear in the context of English texts. It has been used to train our high-performance, compact text embedding models mdbr-leaf-ir and mdbr-leaf-mt. The original list of words stems from here. We have extended it with definitions and important facts about each word using Claude 3.7 Sonnet. textfeature-extraction100K<n<1M0 likes27 downloads9mo agoHugging Face16kurtn718 /scrambled_wordstext10K<n<100K0 likes26 downloads3y agoHugging Face17ilyaxin /Russian-Swear-WordsSource: https://github.com/nickname76/russian-swears textn<1K1 likes26 downloads2y agoHugging Face18Harley-ml /es-en-words Words A dataset comprised of 753k words, 90k of them are Spanish, and 660k of them are English. Key Value Entries (words) 753,232 Tokens 3,225,398 Characters 7,022,310 Avg. Tokens Per Entry ~4.2 Avg. Words Per Entry 1 Avg. Chars Per Entry ~9.3 Longest Entry (Tokens) 36 Shortest Entry (Tokens) 1 English Words~660k Spanish Words ~90k Check out Tiny-Word: A Model Trained on 753k Words Have fun. ALotta Words for you to enjoy! texttext-generation100K<n<1M0 likes26 downloads9mo agoHugging Face19rustemgareev /russian-foreign-words Russian Foreign Words This dataset is based on the Dictionary of Foreign Words developed by the Institute for Linguistic Studies of the Russian Academy of Sciences. Usage The dataset can be loaded using the Hugging Face datasets library. from datasets import load_dataset dataset = load_dataset("rustemgareev/russian-foreign-words", split='train') Dataset Structure Each entry in the dataset represents a dictionary article and is stored as a JSON object with the… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-foreign-words.text10K<n<100K0 likes26 downloads7mo agoHugging Face20text2font /words_with_path_tags_version_2_splittedtext100K<n<1M0 likes24 downloads3y agoHugging Face21AlhitawiMohammed22 /words_hu_dictThis dataset was generated from a Hungarian dictionary, where 60345 sample given The command used to generate data : python3 run.py -i "dicts/hu.txt" -t 8 -f 64 -l hu -c 60345 -na 2 --output_dir "out/words/hu/" --font_dir fonts/hu/ -b 3 -al 0 TRDGHuMu is used for generating text: https://github.com/Mohammed20201991/TextRecognitionDataGeneratorHuMu23 imageimage-to-text10K<n<100K0 likes19 downloads3y agoHugging Face22kurtn718 /scrambled_words_multiple_choicetext10K<n<100K0 likes16 downloads3y agoHugging Face23freococo /pali-words-myanmar-script Pali Words in Myanmar Script (Master Index) This dataset is a master index of 220,252 unique Pali words written in Myanmar (Burmese) Unicode script, intended for reuse across linguistic, religious, and computational workflows. Data Fields Each record contains: word_id: A stable, sequential integer identifier. pali_word: A Pali lexical item rendered in Myanmar Unicode script. Data Processing Methodology The dataset was constructed using the following steps:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/pali-words-myanmar-script.text100K<n<1M0 likes16 downloads8mo agoHugging Face24OmerKuru /old_turkish_wordstext10K<n<100K0 likes14 downloads2mo agoHugging Face25text2font /words_with_path_tags_version_2_validtext10K<n<100K0 likes13 downloads3y agoHugging Face26Highgroundbkk /anki-words anki:Kaishi 1.5k Vocabulary Dataset Japanese vocabulary dataset exported from Anki deck. Dataset Information Deck Name: anki:Kaishi 1.5k Word Count: 478 Format: JSONL (one word per line) Last Updated: 2026-02-21 Schema Each word entry contains: word (string): The Japanese word in kanji/kana reading (string): The reading in hiragana/katakana meaning (string): English translation/meaning jlpt: JLPT level (N5, N4, N3, N2, N1) N5: Basic level (easiest) N4:… See the full description on the dataset page: https://huggingface.co/datasets/Highgroundbkk/anki-words.textn<1K0 likes13 downloads7mo agoHugging Face270x22almostEvil /words-operations-rewards-5k Dataset Card for words-operations-rewards-5k with 5K entries. Dataset Summary License: Apache-2.0. Contains JSONL. Use this for Reward Models. Solved tasks: Count Letters; Write Backwards; Write Character on a Position; Repeat Word; Write In Case; Change Case on a Position; Write Numbering; Connect Characters; Write a Word from Characters; Count Syllables; Example: { "message_tree_id": "00000000-0000-0000-0000-000000000004", "tree_state":… See the full description on the dataset page: https://huggingface.co/datasets/0x22almostEvil/words-operations-rewards-5k.texttext-classification1K<n<10K1 likes12 downloads3y agoHugging Face28text2font /words_with_path_tags_version_2text100K<n<1M0 likes12 downloads3y agoHugging Face29ByunByun /6key_wordstext1K<n<10K0 likes12 downloads3y agoHugging Face30text2font /words_with_path_tags_version_2_traintext100K<n<1M0 likes11 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.