datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kurdish-corpus
Kurdish Corpus
A large-scale multi-source Kurdish language dataset for training language models.
Dataset Statistics
Total Documents: 1,797,686
Total Tokens: 625,716,980
Shards: 4
Built: 2026-05-02
By Language
Language
Documents
Sorani (ckb)
1,274,425
Kurmanji (kmr)
478,540
Zazaki (diq)
34,069
Hawrami (hac)
10,652
By Source Type
Source Type
Documents
News
1,443,750
Web
229,705
Wikipedia
124,231
Usage… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-ai/kurdish-corpus.kurdish-web
Dataset Card for Kurdish Web Corpus (Deduplicated)
Dataset Summary
A multilingual corpus of web text in three Kurdish/Zaza varieties — Kurmanji Kurdish
(kmr_Latn), Sorani Kurdish (ckb_Arab), and Zazaki (diq_Latn) — scraped from Kurdish-language
websites, language-identified with GlotLID, and
deduplicated (exact + MinHash near-duplicate removal). See the "Dataset Creation" section below
for the full pipeline.
Rows, by language config:
config
language
script… See the full description on the dataset page: https://huggingface.co/datasets/muzaffercky/kurdish-web.KurdishCorpus-Clean
⚠️ This is a personal mirror. The actively maintained, canonical version of this dataset is kurdish-tech/KurdishCorpus-clean — that's the one to cite, link, and build on. This copy is kept for history.
The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1
A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji
(Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a
Zazaki baseline. Built for language modeling… See the full description on the dataset page: https://huggingface.co/datasets/alanhasn/KurdishCorpus-Clean.KurdishCorpus-clean
The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1
A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji
(Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a
Zazaki baseline. Built for language modeling, tokenizer training, and
general-purpose Kurdish NLP.
This release contains only openly-licensed or presumptively-free
redistributable content. A parallel research-tier subset (copyrighted
commercial… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/KurdishCorpus-clean.kurdish-unified-corpus
Unified Kurdish Corpus
Dataset Description
This dataset aggregates 758,166 Kurdish text samples from multiple high-quality sources. All text has been preprocessed using asosoft.
Languages
Central Kurdish (ckb)
Kurdish (ku)
Dataset Structure
Columns
text: Preprocessed text content (asosoft applied)
base_dataset: Source dataset name
url: Source URL (NULL if not available)
word_count: Number of words (space-separated)
character_count:… See the full description on the dataset page: https://huggingface.co/datasets/shiima/kurdish-unified-corpus.
