CoolFace
Datasetpublic

Professor/dholuo-speech-data

Dholuo Speech Data (Pooled) A ~191.5-hour Dholuo (Luo) speech corpus, pooled from two independent sources and filtered to only genuinely transcribed audio. Part of the AfroNet multi-language TTS data effort. Sources Anv-ke/Dholuo — African Next Voices, a pilot data-collection effort in Kenya led by the KenCorpus Consortium (a coalition of Kenyan universities and research centers), funded by the Gates Foundation. 91,672 clips, 186.1h, source = anv_ke. Gated on… See the full description on the dataset page: https://huggingface.co/datasets/Professor/dholuo-speech-data.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes62downloads
9 commits on main
1f49ede1mo ago

Add dataset card

Professor
9262ef41mo ago

Add audio shards

Professor
c3268a31mo ago

Add manifest.jsonl

Professor
245250c1mo ago

Add manifest.parquet

Professor
d51c74f1mo ago

Add dataset card

Professor
5e7aa041mo ago

Add audio shards

Professor
a25f8e21mo ago

Add manifest.jsonl

Professor
21724ed1mo ago

Add manifest.parquet

Professor
a4523e31mo ago

initial commit

Professor