datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.task044_essential_terms_identifying_essential_words
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task044_essential_terms_identifying_essential_words
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task044_essential_terms_identifying_essential_words.101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/muhammadrizo5721/101_billion_arabic_words_dataset.arabic-billion-words
Arabic Billion Words Dataset 🌕
The Abu El-Khair Arabic News Corpus (arabic-billion-words) is a comprehensive collection of Arabic text, encompassing over five million newspaper articles. The corpus is rich in linguistic diversity, containing more than a billion and a half words, with approximately three million unique words. The text is encoded in two formats: UTF-8 and Windows CP-1256, and marked up using two markup languages: SGML and XML.
Data Example
An example… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-billion-words.task039_qasc_find_overlapping_words
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task039_qasc_find_overlapping_words
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task039_qasc_find_overlapping_words.numeral-words
numeral-words
A 3-way parallel digital dataset containing 999,999 spelled-out Hmar & English number words mapped in sequential numerical order (1 to 999,999).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Languages: Hmar (hmr, ISO 639-3, Glottolog: hmar1241), English (en)
Family: Zo Languages
Volume: 999,999 parallel rows (1 to 999,999)
Format: Compressed JSONL (data/train-*.jsonl.gz)
License: Apache-2.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/numeral-words.task089_swap_words_verification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task089_swap_words_verification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task089_swap_words_verification.iraqi_words_finetuning
Iraqi Words
A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a
dependency-free BM25 retriever and a fine-tuning data generator built on top of it.
Why this exists
Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA)
and higher-resource dialects such as Egyptian or Levantine. Lexical resources that
map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or
instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.task163_count_words_ending_with_letter
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task163_count_words_ending_with_letter
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task163_count_words_ending_with_letter.for-the-small-shield-chapters
Foreword
The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster.
I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct
I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.task162_count_words_starting_with_letter
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task162_count_words_starting_with_letter
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task162_count_words_starting_with_letter.task378_reverse_words_of_given_length
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task378_reverse_words_of_given_length
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task378_reverse_words_of_given_length.101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/MoryBinM/101_billion_arabic_words_dataset.task161_count_words_containing_letter
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task161_count_words_containing_letter
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task161_count_words_containing_letter.task377_remove_words_of_given_length
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task377_remove_words_of_given_length
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task377_remove_words_of_given_length.spanish-billion-words-v2rewrite-questions-real-words-sciency
real_words_sciency.csv - Question Rewriting Dataset
This dataset contains question rewriting outputs from the file real_words_sciency.csv.
Dataset Structure
The dataset contains the following columns:
custom_id: Unique identifier for each question
style: Rewriting style applied (e.g., "gibberish")
index: Numerical index
original: Original question text
rewritten: Rewritten version of the question
options: Multiple choice options (list format)
correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-real-words-sciency.es-en-words
Words
A dataset comprised of 753k words, 90k of them are Spanish, and 660k of them are English.
Key
Value
Entries (words)
753,232
Tokens
3,225,398
Characters
7,022,310
Avg. Tokens Per Entry
~4.2
Avg. Words Per Entry
1
Avg. Chars Per Entry
~9.3
Longest Entry (Tokens)
36
Shortest Entry (Tokens)
1
English Words~660k
Spanish Words
~90k
Check out Tiny-Word: A Model Trained on 753k Words
Have fun. ALotta Words for you to enjoy!
japanese-trending-words
Japanese Trending Words Dataset (2006-2025)
Dataset Description
This dataset comprises Japanese trending words from the official annual Japanese Trending Words Awards (流行語大賞) from 2006 to 2025, documenting the cultural, social, and political phenomena that have shaped Japan over the past two decades.
Dataset Summary
Total entries: 593 words
Time period: 2006-2025 (20 years)
Languages: Japanese with English translations
Format: CSV with seven columns: word… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-trending-words.101_billion_arabic_words_dataset_urls
Dataset Card for 101_billion_arabic_words_dataset_urls
This dataset provides the URLs and top-level domains associated with training records in ClusterlabAi/101_billion_arabic_words_dataset. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/101_billion_arabic_words_dataset_urls.turkish-number-words-1m
Turkish Number Words 1M v2
0 ile 999.999 arasındaki her tamsayının Türkçe yazıyla karşılığı.
Doğrulanmış boyut
Train: 980,000
Validation: 10,000
Test: 10,000
Toplam: 1,000,000
Ana görev sütunları: id, number, words
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version, generator_sha256… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-number-words-1m.pt-all-words
Dataset Card for Dicionário Português
It is a list of portuguese words with its inflections
How to use it:
from datasets import load_dataset
remote_dataset = load_dataset("VanessaSchenkel/pt-all-words")
remote_dataset
for-the-small-shield-instruct
For The Small Shield — Instruction Data
The training data to fine-tune an LLM is derived from a 1.2-million-word
manuscript called For The Small Shield (https://github.com/wordsum/For_The_Small_Shield),
which I open-sourced 9 years ago.
For The Small Shield is grimdark, so the QA pairs may be grimdark.
The system role in the training files contains the only words
I wrote in the dataset and are intended to make the model just darkish.
I've used this to fine-tune a Llama model… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-instruct.india-trending-words
Google India Trending Words Dataset (2008-2021, 2023-2024)
Dataset Description
This dataset contains Google trending search terms specific to India from 2008 to 2024 (https://trends.withgoogle.com).
Dataset Summary
Total Entries: 900
Years Covered: 2008-2009, 2011-2021, 2023-2024 (15 years, 2010 and 2022 data not available)
Categories: 18 unique tags
Region: India
Format: CSV
Dataset Structure
Data Fields
word (string): The trending… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/india-trending-words.task159_check_frequency_of_words_in_sentence_pair
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task159_check_frequency_of_words_in_sentence_pair
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task159_check_frequency_of_words_in_sentence_pair.task158_count_frequency_of_words
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task158_count_frequency_of_words
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task158_count_frequency_of_words.task376_reverse_order_of_words
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task376_reverse_order_of_words
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task376_reverse_order_of_words.bok_words_700
