datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KurdishCorpus-Clean
⚠️ This is a personal mirror. The actively maintained, canonical version of this dataset is kurdish-tech/KurdishCorpus-clean — that's the one to cite, link, and build on. This copy is kept for history.
The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1
A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji
(Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a
Zazaki baseline. Built for language modeling… See the full description on the dataset page: https://huggingface.co/datasets/alanhasn/KurdishCorpus-Clean.KurdishCorpus-clean
The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1
A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji
(Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a
Zazaki baseline. Built for language modeling, tokenizer training, and
general-purpose Kurdish NLP.
This release contains only openly-licensed or presumptively-free
redistributable content. A parallel research-tier subset (copyrighted
commercial… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/KurdishCorpus-clean.Kurdish_Multi-Domain_Corpus_KMDC
Kurdish Multi-Domain Corpus (KMDC)
Dataset Description
The Kurdish Multi-Domain Corpus (KMDC) is a large-scale instruction-style dataset designed to support natural language processing (NLP), supervised fine-tuning (SFT), and large language model (LLM) development for Central Kurdish (Sorani). The dataset consists of structured question–response pairs generated through an LLM-guided pipeline that transforms raw Kurdish text into machine-learning-ready… See the full description on the dataset page: https://huggingface.co/datasets/shkomq/Kurdish_Multi-Domain_Corpus_KMDC.
