CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kurdish-ai /kurdish-corpus Kurdish Corpus A large-scale multi-source Kurdish language dataset for training language models. Dataset Statistics Total Documents: 1,797,686 Total Tokens: 625,716,980 Shards: 4 Built: 2026-05-02 By Language Language Documents Sorani (ckb) 1,274,425 Kurmanji (kmr) 478,540 Zazaki (diq) 34,069 Hawrami (hac) 10,652 By Source Type Source Type Documents News 1,443,750 Web 229,705 Wikipedia 124,231 Usage… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-ai/kurdish-corpus.tabulartext-generation1M<n<10M2 likes95 downloads5mo agoHugging Face02muzaffercky /kurdish-web Dataset Card for Kurdish Web Corpus (Deduplicated) Dataset Summary A multilingual corpus of web text in three Kurdish/Zaza varieties — Kurmanji Kurdish (kmr_Latn), Sorani Kurdish (ckb_Arab), and Zazaki (diq_Latn) — scraped from Kurdish-language websites, language-identified with GlotLID, and deduplicated (exact + MinHash near-duplicate removal). See the "Dataset Creation" section below for the full pipeline. Rows, by language config: config language script… See the full description on the dataset page: https://huggingface.co/datasets/muzaffercky/kurdish-web.tabulartext-generation100K<n<1M1 likes91 downloads3mo agoHugging Face03alanhasn /KurdishCorpus-Clean ⚠️ This is a personal mirror. The actively maintained, canonical version of this dataset is kurdish-tech/KurdishCorpus-clean — that's the one to cite, link, and build on. This copy is kept for history. The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1 A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji (Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a Zazaki baseline. Built for language modeling… See the full description on the dataset page: https://huggingface.co/datasets/alanhasn/KurdishCorpus-Clean.tabulartext-generation1M<n<10M3 likes37 downloads1mo agoHugging Face04kurdish-tech /KurdishCorpus-clean The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1 A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji (Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a Zazaki baseline. Built for language modeling, tokenizer training, and general-purpose Kurdish NLP. This release contains only openly-licensed or presumptively-free redistributable content. A parallel research-tier subset (copyrighted commercial… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/KurdishCorpus-clean.tabulartext-generation1M<n<10M2 likes30 downloads2mo agoHugging Face05shiima /kurdish-unified-corpusgated Unified Kurdish Corpus Dataset Description This dataset aggregates 758,166 Kurdish text samples from multiple high-quality sources. All text has been preprocessed using asosoft. Languages Central Kurdish (ckb) Kurdish (ku) Dataset Structure Columns text: Preprocessed text content (asosoft applied) base_dataset: Source dataset name url: Source URL (NULL if not available) word_count: Number of words (space-separated) character_count:… See the full description on the dataset page: https://huggingface.co/datasets/shiima/kurdish-unified-corpus.tabulartext-generation100K<n<1M0 likes5 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.