CoolFace
20 results

ckb

CKBETA001 /BBFFTT0 likes442 downloads11mo agoHugging Facerazhan /yt-ckb yt-ckb — Kurdish speech from YouTube ~775k short WAV clips of Central Kurdish (Sorani) speech with their transcriptions, sharded into 56 tar archives (116.4 GB total) plus a metadata.csv index. At a glance Files 56 shards + metadata.csv Total size 116.35 GB Shard size 2,147,328,000 bytes (~2.0 GiB) each, last one 273,244,160 bytes Clips 775,172 (with header line: 775,173 lines in metadata.csv) Index metadata.csv, 124.4 MB, columns file_name… See the full description on the dataset page: https://huggingface.co/datasets/razhan/yt-ckb.automatic-speech-recognition100K<n<1M0 likes210 downloads3d agoHugging FaceCKBETA001 /SSSSSS0 likes77 downloads11mo agoHugging Facerazhan /imdb_ckb IMDB Kurdish (imdb_ckb) Central Kurdish (Sorani) movie-review sentiment: 49,595 reviews labelled positive or negative, translated from the Stanford IMDB review set. Balanced labels, so a 0.50 accuracy baseline is meaningless — report F1. At a glance Rows 49,595 — train 24,903 / test 24,692 Columns text (string), label (0 = negative, 1 = positive) Files train.csv (59.4 MB), test.csv (57.5 MB) — CSV, not parquet Language Central Kurdish / Sorani… See the full description on the dataset page: https://huggingface.co/datasets/razhan/imdb_ckb.texttext-classification10K<n<100K1 likes57 downloads3d agoHugging Facerazhan /script_normalization_ckb Script normalization CKB — noisy → standard Sorani (script_normalization_ckb) Nearly 6M sentence pairs: the text column holds Central Kurdish written with non-standard or distorted characters, summary holds the same sentence in standard Sorani orthography. Use it for text normalisation, spell correction, or to learn the character-level mapping rules. At a glance Rows 5,997,025 — train 5,330,689 / test 666,336 Columns text (noisy input), summary… See the full description on the dataset page: https://huggingface.co/datasets/razhan/script_normalization_ckb.text1M<n<10M0 likes54 downloads3d agoHugging Facerazhan /fw2-ckb FineWeb-2 CKB (fw2-ckb) The Central Kurdish (ckb) slice of FineWeb-2: 501,443 web documents from Common Crawl, filtered for Kurdish and deduplicated, with the full FineWeb-2 metadata carried through (dump, URL, date, language score, script, MinHash cluster size). At a glance Rows 501,443 — train 495,859 / test 5,584 Columns text, id, dump, url, date, file_path, language, language_score, language_script, minhash_cluster_size, top_langs Parquet on… See the full description on the dataset page: https://huggingface.co/datasets/razhan/fw2-ckb.tabular100K<n<1M0 likes50 downloads3d agoHugging Face