datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.task044_essential_terms_identifying_essential_words
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task044_essential_terms_identifying_essential_words
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task044_essential_terms_identifying_essential_words.101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/muhammadrizo5721/101_billion_arabic_words_dataset.arabic-billion-words
Arabic Billion Words Dataset 🌕
The Abu El-Khair Arabic News Corpus (arabic-billion-words) is a comprehensive collection of Arabic text, encompassing over five million newspaper articles. The corpus is rich in linguistic diversity, containing more than a billion and a half words, with approximately three million unique words. The text is encoded in two formats: UTF-8 and Windows CP-1256, and marked up using two markup languages: SGML and XML.
Data Example
An example… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-billion-words.task039_qasc_find_overlapping_words
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task039_qasc_find_overlapping_words
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task039_qasc_find_overlapping_words.arabic_billion_wordsAbu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles.
It contains over a billion and a half words in total, out of which, there are about three million unique words.
The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256.
Also it was marked with two mark-up languages, namely: SGML, and XML.spanish_billion_wordsAn unannotated Spanish corpus of nearly 1.5 billion words, compiled from different resources from the web.
This resources include the spanish portions of SenSem, the Ancora Corpus, some OPUS Project Corpora and the Europarl,
the Tibidabo Treebank, the IULA Spanish LSP Treebank, and dumps from the Spanish Wikipedia, Wikisource and Wikibooks.
This corpus is a compilation of 100 text files. Each line of these files represents one of the 50 million sentences from the corpus.numeral-words
numeral-words
A 3-way parallel digital dataset containing 999,999 spelled-out Hmar & English number words mapped in sequential numerical order (1 to 999,999).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Languages: Hmar (hmr, ISO 639-3, Glottolog: hmar1241), English (en)
Family: Zo Languages
Volume: 999,999 parallel rows (1 to 999,999)
Format: Compressed JSONL (data/train-*.jsonl.gz)
License: Apache-2.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/numeral-words.task089_swap_words_verification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task089_swap_words_verification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task089_swap_words_verification.iraqi_words_finetuning
Iraqi Words
A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a
dependency-free BM25 retriever and a fine-tuning data generator built on top of it.
Why this exists
Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA)
and higher-resource dialects such as Egyptian or Levantine. Lexical resources that
map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or
instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.task163_count_words_ending_with_letter
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task163_count_words_ending_with_letter
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task163_count_words_ending_with_letter.arabic_billion_wordsTHIS IS A FORK FOR LOCAL USAGE.
Abu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles.
It contains over a billion and a half words in total, out of which, there are about three million unique words.
The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256.
Also it was marked with two mark-up languages, namely: SGML, and XML.for-the-small-shield-chapters
Foreword
The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster.
I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct
I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.task162_count_words_starting_with_letter
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task162_count_words_starting_with_letter
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task162_count_words_starting_with_letter.task378_reverse_words_of_given_length
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task378_reverse_words_of_given_length
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task378_reverse_words_of_given_length.101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/MoryBinM/101_billion_arabic_words_dataset.task161_count_words_containing_letter
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task161_count_words_containing_letter
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task161_count_words_containing_letter.task377_remove_words_of_given_length
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task377_remove_words_of_given_length
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task377_remove_words_of_given_length.spanish-billion-words-v2Amazigh-Numbers-To-Words-Dataset
Amazigh Numbers Dataset
Dataset Summary
This dataset maps integers to their Amazigh textual representations across three different numeral counting systems. It is ideal for NLP tasks, localization for the Amazigh language.
Dataset Structure
Data Fields
Number: The integer numerical value.
Ten System Abrv: Base-10 abbreviated representation - The current standard (e.g., 20 is ⵙⵉⵎⵔⴰⵡ).
Ten System Ext: Base-10 extended representation -… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amazigh-Numbers-To-Words-Dataset.arabic_billion_words_old
Dataset Card for Arabic Billion Words Corpus
Dataset Summary
Abu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles.
It contains over a billion and a half words in total, out of which, there are about three million unique words.
The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256.
Also it was marked with two mark-up languages, namely: SGML, and XML.
NB: this dataset is based on the unofficial… See the full description on the dataset page: https://huggingface.co/datasets/oserikov/arabic_billion_words_old.rewrite-questions-real-words-sciency
real_words_sciency.csv - Question Rewriting Dataset
This dataset contains question rewriting outputs from the file real_words_sciency.csv.
Dataset Structure
The dataset contains the following columns:
custom_id: Unique identifier for each question
style: Rewriting style applied (e.g., "gibberish")
index: Numerical index
original: Original question text
rewritten: Rewritten version of the question
options: Multiple choice options (list format)
correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-real-words-sciency.japanese-trending-words
Japanese Trending Words Dataset (2006-2025)
Dataset Description
This dataset comprises Japanese trending words from the official annual Japanese Trending Words Awards (流行語大賞) from 2006 to 2025, documenting the cultural, social, and political phenomena that have shaped Japan over the past two decades.
Dataset Summary
Total entries: 593 words
Time period: 2006-2025 (20 years)
Languages: Japanese with English translations
Format: CSV with seven columns: word… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-trending-words.es-en-words
Words
A dataset comprised of 753k words, 90k of them are Spanish, and 660k of them are English.
Key
Value
Entries (words)
753,232
Tokens
3,225,398
Characters
7,022,310
Avg. Tokens Per Entry
~4.2
Avg. Words Per Entry
1
Avg. Chars Per Entry
~9.3
Longest Entry (Tokens)
36
Shortest Entry (Tokens)
1
English Words~660k
Spanish Words
~90k
Check out Tiny-Word: A Model Trained on 753k Words
Have fun. ALotta Words for you to enjoy!
101_billion_arabic_words_dataset_urls
Dataset Card for 101_billion_arabic_words_dataset_urls
This dataset provides the URLs and top-level domains associated with training records in ClusterlabAi/101_billion_arabic_words_dataset. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/101_billion_arabic_words_dataset_urls.turkish-number-words-1m
Turkish Number Words 1M v2
0 ile 999.999 arasındaki her tamsayının Türkçe yazıyla karşılığı.
Doğrulanmış boyut
Train: 980,000
Validation: 10,000
Test: 10,000
Toplam: 1,000,000
Ana görev sütunları: id, number, words
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version, generator_sha256… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-number-words-1m.pt-all-words
Dataset Card for Dicionário Português
It is a list of portuguese words with its inflections
How to use it:
from datasets import load_dataset
remote_dataset = load_dataset("VanessaSchenkel/pt-all-words")
remote_dataset
for-the-small-shield-instruct
For The Small Shield — Instruction Data
The training data to fine-tune an LLM is derived from a 1.2-million-word
manuscript called For The Small Shield (https://github.com/wordsum/For_The_Small_Shield),
which I open-sourced 9 years ago.
For The Small Shield is grimdark, so the QA pairs may be grimdark.
The system role in the training files contains the only words
I wrote in the dataset and are intended to make the model just darkish.
I've used this to fine-tune a Llama model… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-instruct.TDK_Turkish_WordsThis dataset contains a collection of Turkish dictionary definitions extracted from the official website of the Turkish Language Association (TDK). It provides comprehensive definitions for a wide range of Turkish words and phrases.
The dataset is intended to be a valuable resource for researchers, linguists, language enthusiasts, and anyone interested in the Turkish language. It can be used for various purposes, such as natural language processing tasks, language analysis, and educational… See the full description on the dataset page: https://huggingface.co/datasets/erogluegemen/TDK_Turkish_Words.india-trending-words
Google India Trending Words Dataset (2008-2021, 2023-2024)
Dataset Description
This dataset contains Google trending search terms specific to India from 2008 to 2024 (https://trends.withgoogle.com).
Dataset Summary
Total Entries: 900
Years Covered: 2008-2009, 2011-2021, 2023-2024 (15 years, 2010 and 2022 data not available)
Categories: 18 unique tags
Region: India
Format: CSV
Dataset Structure
Data Fields
word (string): The trending… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/india-trending-words.
