datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imdb_ckb
IMDB Kurdish (imdb_ckb)
Central Kurdish (Sorani) movie-review sentiment: 49,595 reviews labelled positive
or negative, translated from the Stanford IMDB review set. Balanced labels, so a
0.50 accuracy baseline is meaningless — report F1.
At a glance
Rows
49,595 — train 24,903 / test 24,692
Columns
text (string), label (0 = negative, 1 = positive)
Files
train.csv (59.4 MB), test.csv (57.5 MB) — CSV, not parquet
Language
Central Kurdish / Sorani… See the full description on the dataset page: https://huggingface.co/datasets/razhan/imdb_ckb.smol-ckb
smol-ckb
A small Central Kurdish (Sorani) instruction set in the SmolTalk style: an
English-language prompt scaffold with a Kurdish response. 1,183 rows,
1.64M tokens, built by rendering SmolTalk-style templates over existing Kurdish
seed text.
At a glance
Rows
1,183 (single train split)
File
data.csv (5.5 MB)
Columns
index, prompt, response, prompt_tokens, response_tokens, audience, format, seed_data
Language
Prompts in English, responses in… See the full description on the dataset page: https://huggingface.co/datasets/razhan/smol-ckb.Preprocessed-CVS-24-CKBCKBPckbmonolingual_ckbsafa-ckb-dataset
