datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Kurdish_Medical_Corpus_KMC
Kurdish Medical Corpus (KMC) V1
The Kurdish Medical Corpus (KMC) V1 is a high-quality, domain-specific dataset designed for instruction-tuning and scientific knowledge discovery in Central Kurdish (Sorani). This dataset consists entirely of human-written medical content, avoiding the common pitfalls of machine-translated corpora in low-resource language research.
Dataset Summary
KMC V1 is concentrated into a single, consolidated JSON file: kurdish_medical_corpus_kmc.json.… See the full description on the dataset page: https://huggingface.co/datasets/shkomq/Kurdish_Medical_Corpus_KMC.Kurdish_Multi-Domain_Corpus_KMDC
Kurdish Multi-Domain Corpus (KMDC)
Dataset Description
The Kurdish Multi-Domain Corpus (KMDC) is a large-scale instruction-style dataset designed to support natural language processing (NLP), supervised fine-tuning (SFT), and large language model (LLM) development for Central Kurdish (Sorani). The dataset consists of structured question–response pairs generated through an LLM-guided pipeline that transforms raw Kurdish text into machine-learning-ready… See the full description on the dataset page: https://huggingface.co/datasets/shkomq/Kurdish_Multi-Domain_Corpus_KMDC.
