CoolFace
Datasetpublic

Professor/dagbani-speech-data

Dagbani Speech Data (Pooled) A ~96.3-hour Dagbani (Dagbanli) speech corpus, drawn from a single source (WAXAL) and filtered to only genuinely transcribed audio. Part of the AfroNet multi-language TTS data effort. Source WAXAL (google/WaxalNLP), dag_asr config — crowdsourced, image-prompted speech (a shared collection pipeline also used for Dagaare, Ikposo, and Akan's aka_asr in this collection). 17,818 clips, 96.3h, source = waxal. A known upstream bug, verified… See the full description on the dataset page: https://huggingface.co/datasets/Professor/dagbani-speech-data.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes29downloads
5 commits on main
38067a81mo ago

Add dataset card

Professor
a4ac7b91mo ago

Add audio shards

Professor
dbc3a251mo ago

Add manifest.jsonl

Professor
9f35ab91mo ago

Add manifest.parquet

Professor
96e6c111mo ago

initial commit

Professor