datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smol-ckb
smol-ckb
A small Central Kurdish (Sorani) instruction set in the SmolTalk style: an
English-language prompt scaffold with a Kurdish response. 1,183 rows,
1.64M tokens, built by rendering SmolTalk-style templates over existing Kurdish
seed text.
At a glance
Rows
1,183 (single train split)
File
data.csv (5.5 MB)
Columns
index, prompt, response, prompt_tokens, response_tokens, audience, format, seed_data
Language
Prompts in English, responses in… See the full description on the dataset page: https://huggingface.co/datasets/razhan/smol-ckb.Preprocessed-CVS-24-CKBsafa-ckb-dataset
