datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kakugo-ckb
Kakugo Central Kurdish dataset
[Paper] [Code] [Model]
A synthetically generated conversation dataset for training in Central Kurdish.
This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Central Kurdish. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/kakugo-ckb.corpus-ckb
Dataset Overview
The corpus-ckb is a large-scale text dataset primarily composed of Kurdish texts. It is intended for use in various natural language processing (NLP) tasks such as text classification, language modeling, and machine translation. The dataset is particularly useful for researchers and developers working with Kurdish (Central Kurdish, ckb) language data.
Dataset Details
Dataset Info
Features: The dataset contains a single feature:
text: A string… See the full description on the dataset page: https://huggingface.co/datasets/PawanKrd/corpus-ckb.
