CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ClusterlabAi /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.texttext-generation10M<n<100M73 likes1.8k downloads2y agoHugging Face02Lots-of-LoRAs /task044_essential_terms_identifying_essential_words Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task044_essential_terms_identifying_essential_words Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task044_essential_terms_identifying_essential_words.texttext-generation1K<n<10K0 likes332 downloads2y agoHugging Face03muhammadrizo5721 /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/muhammadrizo5721/101_billion_arabic_words_dataset.texttext-generation10M<n<100M0 likes279 downloads7mo agoHugging Face04MohamedRashad /arabic-billion-words Arabic Billion Words Dataset 🌕 The Abu El-Khair Arabic News Corpus (arabic-billion-words) is a comprehensive collection of Arabic text, encompassing over five million newspaper articles. The corpus is rich in linguistic diversity, containing more than a billion and a half words, with approximately three million unique words. The text is encoded in two formats: UTF-8 and Windows CP-1256, and marked up using two markup languages: SGML and XML. Data Example An example… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-billion-words.texttext-generation1M<n<10M12 likes233 downloads3y agoHugging Face05Lots-of-LoRAs /task039_qasc_find_overlapping_words Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task039_qasc_find_overlapping_words Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task039_qasc_find_overlapping_words.texttext-generation1K<n<10K1 likes220 downloads2y agoHugging Face06abuelkhair-corpus /arabic_billion_wordsAbu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles. It contains over a billion and a half words in total, out of which, there are about three million unique words. The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256. Also it was marked with two mark-up languages, namely: SGML, and XML.text-generation100K<n<1M35 likes204 downloads3y agoHugging Face07crscardellino /spanish_billion_wordsAn unannotated Spanish corpus of nearly 1.5 billion words, compiled from different resources from the web. This resources include the spanish portions of SenSem, the Ancora Corpus, some OPUS Project Corpora and the Europarl, the Tibidabo Treebank, the IULA Spanish LSP Treebank, and dumps from the Spanish Wikipedia, Wikisource and Wikibooks. This corpus is a compilation of 100 text files. Each line of these files represents one of the 50 million sentences from the corpus.other10M<n<100M13 likes173 downloads3y agoHugging Face08hmar-heritage-org /numeral-wordsgated numeral-words A 3-way parallel digital dataset containing 999,999 spelled-out Hmar & English number words mapped in sequential numerical order (1 to 999,999). Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev). Overview Languages: Hmar (hmr, ISO 639-3, Glottolog: hmar1241), English (en) Family: Zo Languages Volume: 999,999 parallel rows (1 to 999,999) Format: Compressed JSONL (data/train-*.jsonl.gz) License: Apache-2.0 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/numeral-words.texttranslation100K<n<1M0 likes129 downloads11d agoHugging Face09Lots-of-LoRAs /task089_swap_words_verification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task089_swap_words_verification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task089_swap_words_verification.texttext-generation1K<n<10K0 likes124 downloads2y agoHugging Face10ameer4wisam /iraqi_words_finetuning Iraqi Words A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a dependency-free BM25 retriever and a fine-tuning data generator built on top of it. Why this exists Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA) and higher-resource dialects such as Egyptian or Levantine. Lexical resources that map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.texttranslationn<1K0 likes108 downloads2mo agoHugging Face11Lots-of-LoRAs /task163_count_words_ending_with_letter Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task163_count_words_ending_with_letter Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task163_count_words_ending_with_letter.texttext-generation1K<n<10K0 likes98 downloads2y agoHugging Face12oserikov /arabic_billion_wordsTHIS IS A FORK FOR LOCAL USAGE. Abu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles. It contains over a billion and a half words in total, out of which, there are about three million unique words. The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256. Also it was marked with two mark-up languages, namely: SGML, and XML.text-generation100K<n<1M0 likes94 downloads3y agoHugging Face13wordsum /for-the-small-shield-chapters Foreword The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster. I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.tabulartext-retrieval1K<n<10K0 likes94 downloads1mo agoHugging Face14Lots-of-LoRAs /task162_count_words_starting_with_letter Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task162_count_words_starting_with_letter Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task162_count_words_starting_with_letter.texttext-generation1K<n<10K0 likes89 downloads2y agoHugging Face15Lots-of-LoRAs /task378_reverse_words_of_given_length Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task378_reverse_words_of_given_length Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task378_reverse_words_of_given_length.texttext-generation1K<n<10K0 likes83 downloads2y agoHugging Face16MoryBinM /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/MoryBinM/101_billion_arabic_words_dataset.texttext-generation10M<n<100M0 likes80 downloads9mo agoHugging Face17Lots-of-LoRAs /task161_count_words_containing_letter Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task161_count_words_containing_letter Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task161_count_words_containing_letter.texttext-generation1K<n<10K0 likes78 downloads2y agoHugging Face18Lots-of-LoRAs /task377_remove_words_of_given_length Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task377_remove_words_of_given_length Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task377_remove_words_of_given_length.texttext-generation1K<n<10K0 likes74 downloads2y agoHugging Face19yo-lo-pregunto /spanish-billion-words-v2texttext-generation10M<n<100M3 likes66 downloads10mo agoHugging Face20abdelhaqueidali /Amazigh-Numbers-To-Words-Dataset Amazigh Numbers Dataset Dataset Summary This dataset maps integers to their Amazigh textual representations across three different numeral counting systems. It is ideal for NLP tasks, localization for the Amazigh language. Dataset Structure Data Fields Number: The integer numerical value. Ten System Abrv: Base-10 abbreviated representation - The current standard (e.g., 20 is ⵙⵉⵎⵔⴰⵡ). Ten System Ext: Base-10 extended representation -… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amazigh-Numbers-To-Words-Dataset.translation0 likes48 downloads29d agoHugging Face21oserikov /arabic_billion_words_old Dataset Card for Arabic Billion Words Corpus Dataset Summary Abu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles. It contains over a billion and a half words in total, out of which, there are about three million unique words. The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256. Also it was marked with two mark-up languages, namely: SGML, and XML. NB: this dataset is based on the unofficial… See the full description on the dataset page: https://huggingface.co/datasets/oserikov/arabic_billion_words_old.text-generation100K<n<1M0 likes43 downloads3y agoHugging Face22NLie2 /rewrite-questions-real-words-sciency real_words_sciency.csv - Question Rewriting Dataset This dataset contains question rewriting outputs from the file real_words_sciency.csv. Dataset Structure The dataset contains the following columns: custom_id: Unique identifier for each question style: Rewriting style applied (e.g., "gibberish") index: Numerical index original: Original question text rewritten: Rewritten version of the question options: Multiple choice options (list format) correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-real-words-sciency.tabulartext-generationn<1K0 likes32 downloads1y agoHugging Face23ronantakizawa /japanese-trending-words Japanese Trending Words Dataset (2006-2025) Dataset Description This dataset comprises Japanese trending words from the official annual Japanese Trending Words Awards (流行語大賞) from 2006 to 2025, documenting the cultural, social, and political phenomena that have shaped Japan over the past two decades. Dataset Summary Total entries: 593 words Time period: 2006-2025 (20 years) Languages: Japanese with English translations Format: CSV with seven columns: word… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-trending-words.tabulartext-generationn<1K4 likes24 downloads10mo agoHugging Face24Harley-ml /es-en-words Words A dataset comprised of 753k words, 90k of them are Spanish, and 660k of them are English. Key Value Entries (words) 753,232 Tokens 3,225,398 Characters 7,022,310 Avg. Tokens Per Entry ~4.2 Avg. Words Per Entry 1 Avg. Chars Per Entry ~9.3 Longest Entry (Tokens) 36 Shortest Entry (Tokens) 1 English Words~660k Spanish Words ~90k Check out Tiny-Word: A Model Trained on 753k Words Have fun. ALotta Words for you to enjoy! texttext-generation100K<n<1M0 likes24 downloads9mo agoHugging Face25nhagar /101_billion_arabic_words_dataset_urls Dataset Card for 101_billion_arabic_words_dataset_urls This dataset provides the URLs and top-level domains associated with training records in ClusterlabAi/101_billion_arabic_words_dataset. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/101_billion_arabic_words_dataset_urls.texttext-generation10M<n<100M0 likes21 downloads1y agoHugging Face26GoktugD /turkish-number-words-1m Turkish Number Words 1M v2 0 ile 999.999 arasındaki her tamsayının Türkçe yazıyla karşılığı. Doğrulanmış boyut Train: 980,000 Validation: 10,000 Test: 10,000 Toplam: 1,000,000 Ana görev sütunları: id, number, words Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version, generator_sha256… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-number-words-1m.tabulartext-generation1M<n<10M0 likes19 downloads2mo agoHugging Face27VanessaSchenkel /pt-all-words Dataset Card for Dicionário Português It is a list of portuguese words with its inflections How to use it: from datasets import load_dataset remote_dataset = load_dataset("VanessaSchenkel/pt-all-words") remote_dataset textother10K<n<100K7 likes14 downloads4y agoHugging Face28wordsum /for-the-small-shield-instruct For The Small Shield — Instruction Data The training data to fine-tune an LLM is derived from a 1.2-million-word manuscript called For The Small Shield (https://github.com/wordsum/For_The_Small_Shield), which I open-sourced 9 years ago. For The Small Shield is grimdark, so the QA pairs may be grimdark. The system role in the training files contains the only words I wrote in the dataset and are intended to make the model just darkish. I've used this to fine-tune a Llama model… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-instruct.texttext-generation1K<n<10K0 likes14 downloads2mo agoHugging Face29erogluegemen /TDK_Turkish_WordsgatedThis dataset contains a collection of Turkish dictionary definitions extracted from the official website of the Turkish Language Association (TDK). It provides comprehensive definitions for a wide range of Turkish words and phrases. The dataset is intended to be a valuable resource for researchers, linguists, language enthusiasts, and anyone interested in the Turkish language. It can be used for various purposes, such as natural language processing tasks, language analysis, and educational… See the full description on the dataset page: https://huggingface.co/datasets/erogluegemen/TDK_Turkish_Words.text-classification10K<n<100K8 likes12 downloads3y agoHugging Face30ronantakizawa /india-trending-words Google India Trending Words Dataset (2008-2021, 2023-2024) Dataset Description This dataset contains Google trending search terms specific to India from 2008 to 2024 (https://trends.withgoogle.com). Dataset Summary Total Entries: 900 Years Covered: 2008-2009, 2011-2021, 2023-2024 (15 years, 2010 and 2022 data not available) Categories: 18 unique tags Region: India Format: CSV Dataset Structure Data Fields word (string): The trending… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/india-trending-words.tabulartext-classificationn<1K2 likes12 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.