Professor/somali-speech-data
Somali Speech Data (Pooled) A ~103.2-hour Somali speech corpus, drawn from a single source (Afrivoice) and filtered to only genuinely transcribed audio. Part of the AfroNet multi-language TTS data effort. Source DigitalUmuganda/Afrivoice (the general, pan-African Afrivoice release — not Afrivoice_Ethiopia, which we've separately ingested for 5 Ethiopian languages) — Somali portion: 22,627 clips, 103.2h, source dataset_id/source = afrivoice. There is also a Somali… See the full description on the dataset page: https://huggingface.co/datasets/Professor/somali-speech-data.
This repository belongs to Professor on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
