CoolFace
Datasetpublic

kurdish-tech/KurdishCorpus-clean

The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1 A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji (Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a Zazaki baseline. Built for language modeling, tokenizer training, and general-purpose Kurdish NLP. This release contains only openly-licensed or presumptively-free redistributable content. A parallel research-tier subset (copyrighted commercial… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/KurdishCorpus-clean.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
2likes30downloads
8 commits on main
5c446d52mo ago

Upload README.md with huggingface_hub

alanhasn
d9effe92mo ago

Upload README.md with huggingface_hub

alanhasn
562b4792mo ago

Delete corpus_v1_open_release.jsonl

alanhasn
1df17cf2mo ago

Upload README.md with huggingface_hub

alanhasn
e8639a22mo ago

Upload corpus_v2_open_release.jsonl with huggingface_hub

alanhasn
37f615d2mo ago

Upload README.md with huggingface_hub

alanhasn
74018fb2mo ago

Upload corpus_v1_open_release.jsonl with huggingface_hub

alanhasn
454071e2mo ago

initial commit

alanhasn