datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
navigation-corpus-dagbani-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Dag Speech Segments (sentence splitting)
52799 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani
Full-file CTC forced… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-dagbani-speech.dagbani-bible-audio-text-tts
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
53410 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/dagbani-bible-audio-text-tts.navigation-corpus-speech-full-dagbani
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana TTS Navigation Corpus — Dagbani
Synthetic speech dataset for navigation.
Structure
audio/ – all .wav audio files
text/ – matching .txt files with transcriptions
metadata.csv – full metadata table
navigation-corpus-dagbani-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Dag Speech Segments (sentence splitting)
52799 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani
Full-file CTC forced… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/navigation-corpus-dagbani-speech.sautiledger-market-speech
SautiLedger Market Speech: code-switched Pidgin/Yoruba and Shona with English
45 consented, de-identified recordings (217 seconds, 3.6 minutes) of
code-switched market bookkeeping speech, the kind a trader says to a
voice ledger, in three language groups:
Group
Clip prefix
Clips
Speaker
Example
Nigerian Pidgin + Yoruba + English
slms-pcm-
15
spk-ng-01, adult male, Nigeria
"I don sell three derica of rice five thousand five"
Shona + English (prices in US dollars)… See the full description on the dataset page: https://huggingface.co/datasets/dagbolade/sautiledger-market-speech.Dagi-leyu-multilingual-speech-datasetnavigation-corpus-dagbani-speech
Dag Speech Segments (sentence splitting)
52799 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani
Full-file CTC forced alignment (MMS-300M) for word-level timestamps
Sentence-boundary splits (. ? !) — long sentences re-chunked to 16 words
Leading/trailing silence trimmed with VAD (-40 dBFS threshold)
Filtered: min 1.0s, max 15.0s
Original sample rate preserved
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/navigation-corpus-dagbani-speech.ghana-bible-combined-90k-twi-ewe-dagbani
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana Bible Combined 90K Twi Ewe Dagbani
v14-dag-asr-audioleyu-multilingual-speech-datasetdagbani-bible-audio-text-tts
Twi 16-Word Speech Segments
53410 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for word-level timestamps
Words grouped into 16-word segments
Leading/trailing silence trimmed with VAD (-40 dBFS threshold)
Filtered: min 1.0s, max 15.0s
Original sample rate preserved (24kHz)
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/dagbani-bible-audio-text-tts.ghana-nlp-health-UNICEF-asr-dagbani
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
dagbani-tts-bible-full-audio-text
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Dagbani Tts Bible Full Audio Text
dagbani-bible-tts-200
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
test-dagsterpipeline_output_dagsterjtss
