CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ClusterlabAi /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.texttext-generation10M<n<100M73 likes1.8k downloads2y agoHugging Face02Lots-of-LoRAs /task044_essential_terms_identifying_essential_words Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task044_essential_terms_identifying_essential_words Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task044_essential_terms_identifying_essential_words.texttext-generation1K<n<10K0 likes332 downloads2y agoHugging Face03muhammadrizo5721 /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/muhammadrizo5721/101_billion_arabic_words_dataset.texttext-generation10M<n<100M0 likes278 downloads7mo agoHugging Face04MohamedRashad /arabic-billion-words Arabic Billion Words Dataset 🌕 The Abu El-Khair Arabic News Corpus (arabic-billion-words) is a comprehensive collection of Arabic text, encompassing over five million newspaper articles. The corpus is rich in linguistic diversity, containing more than a billion and a half words, with approximately three million unique words. The text is encoded in two formats: UTF-8 and Windows CP-1256, and marked up using two markup languages: SGML and XML. Data Example An example… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-billion-words.texttext-generation1M<n<10M12 likes233 downloads3y agoHugging Face05Lots-of-LoRAs /task039_qasc_find_overlapping_words Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task039_qasc_find_overlapping_words Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task039_qasc_find_overlapping_words.texttext-generation1K<n<10K1 likes220 downloads2y agoHugging Face06hmar-heritage-org /numeral-wordsgated numeral-words A 3-way parallel digital dataset containing 999,999 spelled-out Hmar & English number words mapped in sequential numerical order (1 to 999,999). Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev). Overview Languages: Hmar (hmr, ISO 639-3, Glottolog: hmar1241), English (en) Family: Zo Languages Volume: 999,999 parallel rows (1 to 999,999) Format: Compressed JSONL (data/train-*.jsonl.gz) License: Apache-2.0 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/numeral-words.texttranslation100K<n<1M0 likes129 downloads10d agoHugging Face07Lots-of-LoRAs /task089_swap_words_verification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task089_swap_words_verification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task089_swap_words_verification.texttext-generation1K<n<10K0 likes124 downloads2y agoHugging Face08ameer4wisam /iraqi_words_finetuning Iraqi Words A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a dependency-free BM25 retriever and a fine-tuning data generator built on top of it. Why this exists Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA) and higher-resource dialects such as Egyptian or Levantine. Lexical resources that map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.texttranslationn<1K0 likes111 downloads2mo agoHugging Face09Lots-of-LoRAs /task163_count_words_ending_with_letter Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task163_count_words_ending_with_letter Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task163_count_words_ending_with_letter.texttext-generation1K<n<10K0 likes98 downloads2y agoHugging Face10wordsum /for-the-small-shield-chapters Foreword The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster. I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.tabulartext-retrieval1K<n<10K0 likes94 downloads1mo agoHugging Face11Lots-of-LoRAs /task162_count_words_starting_with_letter Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task162_count_words_starting_with_letter Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task162_count_words_starting_with_letter.texttext-generation1K<n<10K0 likes89 downloads2y agoHugging Face12Lots-of-LoRAs /task378_reverse_words_of_given_length Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task378_reverse_words_of_given_length Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task378_reverse_words_of_given_length.texttext-generation1K<n<10K0 likes83 downloads2y agoHugging Face13MoryBinM /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/MoryBinM/101_billion_arabic_words_dataset.texttext-generation10M<n<100M0 likes79 downloads9mo agoHugging Face14Lots-of-LoRAs /task161_count_words_containing_letter Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task161_count_words_containing_letter Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task161_count_words_containing_letter.texttext-generation1K<n<10K0 likes78 downloads2y agoHugging Face15Lots-of-LoRAs /task377_remove_words_of_given_length Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task377_remove_words_of_given_length Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task377_remove_words_of_given_length.texttext-generation1K<n<10K0 likes74 downloads2y agoHugging Face16yo-lo-pregunto /spanish-billion-words-v2texttext-generation10M<n<100M3 likes66 downloads10mo agoHugging Face17NLie2 /rewrite-questions-real-words-sciency real_words_sciency.csv - Question Rewriting Dataset This dataset contains question rewriting outputs from the file real_words_sciency.csv. Dataset Structure The dataset contains the following columns: custom_id: Unique identifier for each question style: Rewriting style applied (e.g., "gibberish") index: Numerical index original: Original question text rewritten: Rewritten version of the question options: Multiple choice options (list format) correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-real-words-sciency.tabulartext-generationn<1K0 likes32 downloads1y agoHugging Face18Harley-ml /es-en-words Words A dataset comprised of 753k words, 90k of them are Spanish, and 660k of them are English. Key Value Entries (words) 753,232 Tokens 3,225,398 Characters 7,022,310 Avg. Tokens Per Entry ~4.2 Avg. Words Per Entry 1 Avg. Chars Per Entry ~9.3 Longest Entry (Tokens) 36 Shortest Entry (Tokens) 1 English Words~660k Spanish Words ~90k Check out Tiny-Word: A Model Trained on 753k Words Have fun. ALotta Words for you to enjoy! texttext-generation100K<n<1M0 likes26 downloads9mo agoHugging Face19ronantakizawa /japanese-trending-words Japanese Trending Words Dataset (2006-2025) Dataset Description This dataset comprises Japanese trending words from the official annual Japanese Trending Words Awards (流行語大賞) from 2006 to 2025, documenting the cultural, social, and political phenomena that have shaped Japan over the past two decades. Dataset Summary Total entries: 593 words Time period: 2006-2025 (20 years) Languages: Japanese with English translations Format: CSV with seven columns: word… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-trending-words.tabulartext-generationn<1K4 likes23 downloads10mo agoHugging Face20nhagar /101_billion_arabic_words_dataset_urls Dataset Card for 101_billion_arabic_words_dataset_urls This dataset provides the URLs and top-level domains associated with training records in ClusterlabAi/101_billion_arabic_words_dataset. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/101_billion_arabic_words_dataset_urls.texttext-generation10M<n<100M0 likes21 downloads1y agoHugging Face21GoktugD /turkish-number-words-1m Turkish Number Words 1M v2 0 ile 999.999 arasındaki her tamsayının Türkçe yazıyla karşılığı. Doğrulanmış boyut Train: 980,000 Validation: 10,000 Test: 10,000 Toplam: 1,000,000 Ana görev sütunları: id, number, words Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version, generator_sha256… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-number-words-1m.tabulartext-generation1M<n<10M0 likes16 downloads2mo agoHugging Face22VanessaSchenkel /pt-all-words Dataset Card for Dicionário Português It is a list of portuguese words with its inflections How to use it: from datasets import load_dataset remote_dataset = load_dataset("VanessaSchenkel/pt-all-words") remote_dataset textother10K<n<100K7 likes14 downloads4y agoHugging Face23wordsum /for-the-small-shield-instruct For The Small Shield — Instruction Data The training data to fine-tune an LLM is derived from a 1.2-million-word manuscript called For The Small Shield (https://github.com/wordsum/For_The_Small_Shield), which I open-sourced 9 years ago. For The Small Shield is grimdark, so the QA pairs may be grimdark. The system role in the training files contains the only words I wrote in the dataset and are intended to make the model just darkish. I've used this to fine-tune a Llama model… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-instruct.texttext-generation1K<n<10K0 likes14 downloads2mo agoHugging Face24ronantakizawa /india-trending-words Google India Trending Words Dataset (2008-2021, 2023-2024) Dataset Description This dataset contains Google trending search terms specific to India from 2008 to 2024 (https://trends.withgoogle.com). Dataset Summary Total Entries: 900 Years Covered: 2008-2009, 2011-2021, 2023-2024 (15 years, 2010 and 2022 data not available) Categories: 18 unique tags Region: India Format: CSV Dataset Structure Data Fields word (string): The trending… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/india-trending-words.tabulartext-classificationn<1K2 likes12 downloads10mo agoHugging Face25Lots-of-LoRAs /task159_check_frequency_of_words_in_sentence_pair Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task159_check_frequency_of_words_in_sentence_pair Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task159_check_frequency_of_words_in_sentence_pair.texttext-generation1K<n<10K0 likes10 downloads2y agoHugging Face26Lots-of-LoRAs /task158_count_frequency_of_words Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task158_count_frequency_of_words Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task158_count_frequency_of_words.texttext-generation1K<n<10K0 likes9 downloads2y agoHugging Face27Lots-of-LoRAs /task376_reverse_order_of_words Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task376_reverse_order_of_words Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task376_reverse_order_of_words.texttext-generation1K<n<10K0 likes8 downloads2y agoHugging Face28BigShort /bok_words_700gatedtexttext-generationn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.