datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
english-words-definitions
English Words Definitions
This dataset contains definitions and important facts about 467k words that appear in the context of English texts.
It has been used to train our high-performance, compact text embedding models mdbr-leaf-ir and mdbr-leaf-mt.
The original list of words stems from here. We have extended it with definitions and important facts about each word using Claude 3.7 Sonnet.
numeral-words
numeral-words
A 3-way parallel digital dataset containing 999,999 spelled-out Hmar & English number words mapped in sequential numerical order (1 to 999,999).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Languages: Hmar (hmr, ISO 639-3, Glottolog: hmar1241), English (en)
Family: Zo Languages
Volume: 999,999 parallel rows (1 to 999,999)
Format: Compressed JSONL (data/train-*.jsonl.gz)
License: Apache-2.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/numeral-words.iraqi_words_finetuning
Iraqi Words
A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a
dependency-free BM25 retriever and a fine-tuning data generator built on top of it.
Why this exists
Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA)
and higher-resource dialects such as Egyptian or Levantine. Lexical resources that
map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or
instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.stop_wordsfor-the-small-shield-chapters
Foreword
The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster.
I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct
I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.anki-words
anki:Kaishi 1.5k Vocabulary Dataset
Japanese vocabulary dataset exported from Anki deck.
Dataset Information
Deck Name: anki:Kaishi 1.5k
Word Count: 478
Format: JSONL (one word per line)
Last Updated: 2026-02-21
Schema
Each word entry contains:
word (string): The Japanese word in kanji/kana
reading (string): The reading in hiragana/katakana
meaning (string): English translation/meaning
jlpt: JLPT level (N5, N4, N3, N2, N1)
N5: Basic level (easiest)
N4:… See the full description on the dataset page: https://huggingface.co/datasets/lysandre/anki-words.digital-sat-words-in-context-llmEval_Counting_Letters_in_WordsLetters in Words Evaluation Dataset
"The strawberry question is pretty much the new Turing Test for future AI" BlakeSergin OP 3mo agohttps://www.reddit.com/r/singularity/comments/1enqk04/how_many_rs_in_strawberry_why_is_this_a_very/
This dataset .json provides a simple yet effective way to assess the basic letter-counting abilities of Large Language Models (LLMs). (Try it on the new SmolLM2 models.) It consists of a set of questions designed to evaluate an LLM's capacity for:
Understanding… See the full description on the dataset page: https://huggingface.co/datasets/MartialTerran/Eval_Counting_Letters_in_Words.english-words-definitions
English Words Definitions
This dataset contains definitions and important facts about 467k words that appear in the context of English texts.
It has been used to train our high-performance, compact text embedding models mdbr-leaf-ir and mdbr-leaf-mt.
The original list of words stems from here. We have extended it with definitions and important facts about each word using Claude 3.7 Sonnet.
thai-trade-words-chiang-mai
คำบนป้าย — Thai trade words of Chiang Mai and Chiang Rai
1,782 Thai trade terms taken from the tags on business listings in
Chiang Mai and Chiang Rai — the words the city uses for what a shop does —
each with an RTGS reading and an English gloss. Plus 403
administrative place names (tambon, amphoe, city, province) with their
readings.
ซ่อมมอเตอร์ไซค์ Som Motoesai motorcycle repair 191 places
ตู้น้ำดื่มหยอดเหรียญ Tunam Duem Yotrian coin-op drinking-water… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-trade-words-chiang-mai.english-words-wordnet
English Words WordNet Dataset
This dataset contains a comprehensive collection of English words paired with their detailed definitions from WordNet. Includes words between 2 and 20 characters in length.
Each line represents a single record containing the following structure:
{
"word": "zwiebacks",
"index_0": {
"pos": "n",
"definition":"slice of sweet raised bread baked again until it is brown and hard and crisp",
"examples": []
}
}
arabizi_toxic_harassment_words
license: other
task_categories:
- text-classification
task_ids:
- hate-speech-detection
language:
- ar
tags:
- arabizi
- dialectal-arabic
- toxicity
- harassment
- social-media-safety
- red-teaming
size_categories:
- n<1K
🛡️ Pro-Grade Arabizi 250- Sentences - Safety & Moderation Dataset
Purchase Full Access
👉 Buy the Complete Dataset on Gumroad
Arabizi Toxicity & Content Moderation Dataset
A dataset containing Arabizi (3araby) toxic, abusive, and… See the full description on the dataset page: https://huggingface.co/datasets/Arabicc/arabizi_toxic_harassment_words.kazakh-swear-words
Kazakh Swear Words Dataset 🇰🇿
Dataset of Kazakh obscene and profane expressions for NLP tasks including text classification, content moderation, toxicity detection, and LLM fine-tuning.
Dataset Description
This is a low-resource language dataset containing Kazakh profanity, swear words, and offensive expressions along with neutral examples for binary classification tasks.
Languages
Kazakh (kk)
Dataset Structure
Data Files
data.jsonl -… See the full description on the dataset page: https://huggingface.co/datasets/Rustem-Kaimolla/kazakh-swear-words.turkish-rare-wordsenglish-words-definitions
English Words Definitions
This dataset contains definitions and important facts about 467k words that appear in the context of English texts.
It has been used to train our high-performance, compact text embedding models mdbr-leaf-ir and mdbr-leaf-mt.
The original list of words stems from here. We have extended it with definitions and important facts about each word using Claude 3.7 Sonnet.
scrambled_wordsRussian-Swear-WordsSource: https://github.com/nickname76/russian-swears
es-en-words
Words
A dataset comprised of 753k words, 90k of them are Spanish, and 660k of them are English.
Key
Value
Entries (words)
753,232
Tokens
3,225,398
Characters
7,022,310
Avg. Tokens Per Entry
~4.2
Avg. Words Per Entry
1
Avg. Chars Per Entry
~9.3
Longest Entry (Tokens)
36
Shortest Entry (Tokens)
1
English Words~660k
Spanish Words
~90k
Check out Tiny-Word: A Model Trained on 753k Words
Have fun. ALotta Words for you to enjoy!
russian-foreign-words
Russian Foreign Words
This dataset is based on the Dictionary of Foreign Words developed by the Institute for Linguistic Studies of the Russian Academy of Sciences.
Usage
The dataset can be loaded using the Hugging Face datasets library.
from datasets import load_dataset
dataset = load_dataset("rustemgareev/russian-foreign-words", split='train')
Dataset Structure
Each entry in the dataset represents a dictionary article and is stored as a JSON object with the… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-foreign-words.words_with_path_tags_version_2_splittedwords_hu_dictThis dataset was generated from a Hungarian dictionary, where 60345 sample given
The command used to generate data :
python3 run.py -i "dicts/hu.txt" -t 8 -f 64 -l hu -c 60345 -na 2 --output_dir "out/words/hu/" --font_dir fonts/hu/ -b 3 -al 0
TRDGHuMu is used for generating text: https://github.com/Mohammed20201991/TextRecognitionDataGeneratorHuMu23
scrambled_words_multiple_choicepali-words-myanmar-script
Pali Words in Myanmar Script (Master Index)
This dataset is a master index of 220,252 unique Pali words written in Myanmar (Burmese) Unicode script, intended for reuse across linguistic, religious, and computational workflows.
Data Fields
Each record contains:
word_id: A stable, sequential integer identifier.
pali_word: A Pali lexical item rendered in Myanmar Unicode script.
Data Processing Methodology
The dataset was constructed using the following steps:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/pali-words-myanmar-script.old_turkish_wordswords_with_path_tags_version_2_validanki-words
anki:Kaishi 1.5k Vocabulary Dataset
Japanese vocabulary dataset exported from Anki deck.
Dataset Information
Deck Name: anki:Kaishi 1.5k
Word Count: 478
Format: JSONL (one word per line)
Last Updated: 2026-02-21
Schema
Each word entry contains:
word (string): The Japanese word in kanji/kana
reading (string): The reading in hiragana/katakana
meaning (string): English translation/meaning
jlpt: JLPT level (N5, N4, N3, N2, N1)
N5: Basic level (easiest)
N4:… See the full description on the dataset page: https://huggingface.co/datasets/Highgroundbkk/anki-words.words-operations-rewards-5k
Dataset Card for words-operations-rewards-5k with 5K entries.
Dataset Summary
License: Apache-2.0. Contains JSONL. Use this for Reward Models.
Solved tasks:
Count Letters;
Write Backwards;
Write Character on a Position;
Repeat Word;
Write In Case;
Change Case on a Position;
Write Numbering;
Connect Characters;
Write a Word from Characters;
Count Syllables;
Example:
{
"message_tree_id": "00000000-0000-0000-0000-000000000004",
"tree_state":… See the full description on the dataset page: https://huggingface.co/datasets/0x22almostEvil/words-operations-rewards-5k.words_with_path_tags_version_26key_wordswords_with_path_tags_version_2_train
