kurdish-tech/KurdishCorpus-clean
The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1 A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji (Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a Zazaki baseline. Built for language modeling, tokenizer training, and general-purpose Kurdish NLP. This release contains only openly-licensed or presumptively-free redistributable content. A parallel research-tier subset (copyrighted commercial… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/KurdishCorpus-clean.
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Delete corpus_v1_open_release.jsonl
Upload README.md with huggingface_hub
Upload corpus_v2_open_release.jsonl with huggingface_hub
Upload README.md with huggingface_hub
Upload corpus_v1_open_release.jsonl with huggingface_hub
initial commit
