kurdish-tech/KurdishCorpus-clean
The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1 A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji (Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a Zazaki baseline. Built for language modeling, tokenizer training, and general-purpose Kurdish NLP. This release contains only openly-licensed or presumptively-free redistributable content. A parallel research-tier subset (copyrighted commercial… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/KurdishCorpus-clean.
This repository belongs to kurdish-tech on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
