datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
navigation-corpus-dagbani-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Dag Speech Segments (sentence splitting)
52799 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani
Full-file CTC forced… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-dagbani-speech.dagbani-bible-audio-text-tts
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
53410 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/dagbani-bible-audio-text-tts.navigation-corpus-speech-full-dagbani
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana TTS Navigation Corpus — Dagbani
Synthetic speech dataset for navigation.
Structure
audio/ – all .wav audio files
text/ – matching .txt files with transcriptions
metadata.csv – full metadata table
navigation-corpus-dagbani-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Dag Speech Segments (sentence splitting)
52799 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani
Full-file CTC forced… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/navigation-corpus-dagbani-speech.sautiledger-market-speech
SautiLedger Market Speech: code-switched Pidgin/Yoruba and Shona with English
45 consented, de-identified recordings (217 seconds, 3.6 minutes) of
code-switched market bookkeeping speech, the kind a trader says to a
voice ledger, in three language groups:
Group
Clip prefix
Clips
Speaker
Example
Nigerian Pidgin + Yoruba + English
slms-pcm-
15
spk-ng-01, adult male, Nigeria
"I don sell three derica of rice five thousand five"
Shona + English (prices in US dollars)… See the full description on the dataset page: https://huggingface.co/datasets/dagbolade/sautiledger-market-speech.navigation-corpus-dagbani-speech
Dag Speech Segments (sentence splitting)
52799 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani
Full-file CTC forced alignment (MMS-300M) for word-level timestamps
Sentence-boundary splits (. ? !) — long sentences re-chunked to 16 words
Leading/trailing silence trimmed with VAD (-40 dBFS threshold)
Filtered: min 1.0s, max 15.0s
Original sample rate preserved
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/navigation-corpus-dagbani-speech.dagbani-speech-data
Dagbani Speech Data (Pooled)
A ~96.3-hour Dagbani (Dagbanli) speech corpus, drawn from a single source
(WAXAL) and filtered to only genuinely transcribed audio. Part of the
AfroNet multi-language TTS data
effort.
Source
WAXAL (google/WaxalNLP),
dag_asr config — crowdsourced, image-prompted speech (a shared collection pipeline
also used for Dagaare, Ikposo, and Akan's aka_asr in this collection). 17,818
clips, 96.3h, source = waxal.
A known upstream bug, verified… See the full description on the dataset page: https://huggingface.co/datasets/Professor/dagbani-speech-data.dagbani-bible-audio-text-tts
Twi 16-Word Speech Segments
53410 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for word-level timestamps
Words grouped into 16-word segments
Leading/trailing silence trimmed with VAD (-40 dBFS threshold)
Filtered: min 1.0s, max 15.0s
Original sample rate preserved (24kHz)
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/dagbani-bible-audio-text-tts.dagaare-speech-data
Dagaare Speech Data (Pooled)
A ~104.3-hour Dagaare (Dagara) speech corpus, drawn from a single source
(WAXAL) and filtered to only genuinely transcribed audio. Part of the
AfroNet multi-language TTS data
effort.
Source
WAXAL (google/WaxalNLP),
dga_asr config — crowdsourced, image-prompted speech (a shared collection pipeline
also used for Dagbani, Ikposo, and Akan's aka_asr in this collection). 18,859
clips, 104.3h, source = waxal.
A known upstream bug, verified… See the full description on the dataset page: https://huggingface.co/datasets/Professor/dagaare-speech-data.
