ckb
Datasets
All datasets matching “ckb”BBFFTTyt-ckb
yt-ckb — Kurdish speech from YouTube
~775k short WAV clips of Central Kurdish (Sorani) speech with their transcriptions,
sharded into 56 tar archives (116.4 GB total) plus a metadata.csv index.
At a glance
Files
56 shards + metadata.csv
Total size
116.35 GB
Shard size
2,147,328,000 bytes (~2.0 GiB) each, last one 273,244,160 bytes
Clips
775,172 (with header line: 775,173 lines in metadata.csv)
Index
metadata.csv, 124.4 MB, columns file_name… See the full description on the dataset page: https://huggingface.co/datasets/razhan/yt-ckb.SSSSSSimdb_ckb
IMDB Kurdish (imdb_ckb)
Central Kurdish (Sorani) movie-review sentiment: 49,595 reviews labelled positive
or negative, translated from the Stanford IMDB review set. Balanced labels, so a
0.50 accuracy baseline is meaningless — report F1.
At a glance
Rows
49,595 — train 24,903 / test 24,692
Columns
text (string), label (0 = negative, 1 = positive)
Files
train.csv (59.4 MB), test.csv (57.5 MB) — CSV, not parquet
Language
Central Kurdish / Sorani… See the full description on the dataset page: https://huggingface.co/datasets/razhan/imdb_ckb.script_normalization_ckb
Script normalization CKB — noisy → standard Sorani (script_normalization_ckb)
Nearly 6M sentence pairs: the text column holds Central Kurdish written with
non-standard or distorted characters, summary holds the same sentence in standard
Sorani orthography. Use it for text normalisation, spell correction, or to learn the
character-level mapping rules.
At a glance
Rows
5,997,025 — train 5,330,689 / test 666,336
Columns
text (noisy input), summary… See the full description on the dataset page: https://huggingface.co/datasets/razhan/script_normalization_ckb.fw2-ckb
FineWeb-2 CKB (fw2-ckb)
The Central Kurdish (ckb) slice of FineWeb-2: 501,443 web documents from Common
Crawl, filtered for Kurdish and deduplicated, with the full FineWeb-2 metadata
carried through (dump, URL, date, language score, script, MinHash cluster size).
At a glance
Rows
501,443 — train 495,859 / test 5,584
Columns
text, id, dump, url, date, file_path, language, language_score, language_script, minhash_cluster_size, top_langs
Parquet on… See the full description on the dataset page: https://huggingface.co/datasets/razhan/fw2-ckb.
