CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KaraKaraWitch /Kurdish-Underwater-Basketweaving-Forum KaraKaraWitch/Kurdish-Underwater-Basketweaving-Forum Yes this is a 4chan dataset. THIS CONTAINS TOXIC SHIT like (/POL/) content. YOU HAVE BEEN WARNED. KaraKaraWitch & their company dissolves all responsbilities when using this dataset. Text Sample Note: namedconversation is a modification of OAI's conversation format. While identical, namedconversation is not required to stick to system,user,model/assistant verbs. This allows for a much more varied use… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/Kurdish-Underwater-Basketweaving-Forum.text100K<n<1M5 likes175 downloads2y agoHugging Face02kurdish-tech /kurdish-grammar-eval Kurdish Grammar Minimal Pairs (BLiMP-style) — Kurmancî · Soranî A grammar-competence benchmark for Kurdish, built on the BLiMP idea: for each item, a correct sentence is paired with a corrupted version where one specific grammar rule has been deliberately broken. Score a language model by checking whether it assigns higher likelihood to the correct sentence than the corrupted one — accuracy well above 50% means the model learned the rule, not just surface fluency. Built by… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/kurdish-grammar-eval.texttext-classification1K<n<10K1 likes68 downloads1mo agoHugging Face03alanhasn /KurdishCorpus-Clean ⚠️ This is a personal mirror. The actively maintained, canonical version of this dataset is kurdish-tech/KurdishCorpus-clean — that's the one to cite, link, and build on. This copy is kept for history. The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1 A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji (Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a Zazaki baseline. Built for language modeling… See the full description on the dataset page: https://huggingface.co/datasets/alanhasn/KurdishCorpus-Clean.tabulartext-generation1M<n<10M3 likes39 downloads1mo agoHugging Face04kurdish-tech /KurdishCorpus-clean The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1 A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji (Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a Zazaki baseline. Built for language modeling, tokenizer training, and general-purpose Kurdish NLP. This release contains only openly-licensed or presumptively-free redistributable content. A parallel research-tier subset (copyrighted commercial… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/KurdishCorpus-clean.tabulartext-generation1M<n<10M2 likes34 downloads2mo agoHugging Face05open-llm-leaderboard /nazimali__Mistral-Nemo-Kurdish-Instruct-detailsgated Dataset Card for Evaluation run of nazimali/Mistral-Nemo-Kurdish-Instruct Dataset automatically created during the evaluation run of model nazimali/Mistral-Nemo-Kurdish-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nazimali__Mistral-Nemo-Kurdish-Instruct-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face06open-llm-leaderboard /nazimali__Mistral-Nemo-Kurdish-detailsgated Dataset Card for Evaluation run of nazimali/Mistral-Nemo-Kurdish Dataset automatically created during the evaluation run of model nazimali/Mistral-Nemo-Kurdish The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nazimali__Mistral-Nemo-Kurdish-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face07shkomq /Kurdish_Multi-Domain_Corpus_KMDCgated Kurdish Multi-Domain Corpus (KMDC) Dataset Description The Kurdish Multi-Domain Corpus (KMDC) is a large-scale instruction-style dataset designed to support natural language processing (NLP), supervised fine-tuning (SFT), and large language model (LLM) development for Central Kurdish (Sorani). The dataset consists of structured question–response pairs generated through an LLM-guided pipeline that transforms raw Kurdish text into machine-learning-ready… See the full description on the dataset page: https://huggingface.co/datasets/shkomq/Kurdish_Multi-Domain_Corpus_KMDC.textquestion-answering100K<n<1M0 likes5 downloads4mo agoHugging Face08fadhill /kurdish_sorani_curriculum_textn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.