datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ghana-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Language Statistics
Language
Subset
Segments
Duration
Akuapem_Twi
Akuapem_Twi_twi
52,650
63.25h
Anyin
Anyin_any
5,568
13.24h
Asante_Twi
Asante_Twi_twi
143,383
200.02h
Avatime
Avatime_avn
9,956
21.62h
Bassar_Ntcham… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-speech.ghana-speech-ipa
Ghana Speech — Audio with IPA Transcripts
Speech with both transcript forms: the original orthography and the IPA phoneme
sequence read off the audio by ASR. Each language is a subset, with real
train/validation splits.
from datasets import load_dataset
ds = load_dataset("ghanaopendata/ghana-speech-ipa", "Akuapem_Twi_twi", split="train")
ds[0]["audio"] # decoded waveform, 16 kHz
ds[0]["text"] # original orthography
ds[0]["ipa"] # IPA phonemes
369,347 clips · ~747 h… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-speech-ipa.ghana-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Language Statistics
Language
Subset
Segments
Duration
Akuapem_Twi
Akuapem_Twi_twi
52,650
63.25h
Anyin
Anyin_any
5,568
13.24h
Asante_Twi
Asante_Twi_twi
143,383
200.02h
Avatime
Avatime_avn
9,956
21.62h
Bassar_Ntcham… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-speech.kasem-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Kasem Speech-Text Parallel Dataset
Dataset Description
This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.ghana-english-asr-2700hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
🇬🇭 Ghana English ASR Dataset
A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts,
designed for training and fine-tuning Automatic Speech Recognition (ASR) models on
West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-asr-2700hrs.twi-health-asr-gemini-500hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset Gemini (500 hours)
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken
languages, sourced from publicly available video content on health and wellness.
Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr-gemini-500hrs.ghana-english-speech-600hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
🇬🇭 Ghana English ASR Dataset
A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts,
designed for training and fine-tuning Automatic Speech Recognition (ASR) models on
West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-speech-600hrs.new-twi-tts-aligned-ipa
new-twi-tts-aligned + IPA phonemes
ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme
transcription for every clip, produced with
ghananlpcommunity/ghana-speech-phoneme-asr.
Audio included — this is self-contained, no join with the source dataset needed.
Contents
split
clips
hours
phoneme units
mean units/clip
test
16,140
17.24
663,140
41.1
train
145,258
155.21
5,945,389
40.9
Columns
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned-ipa.ga-speech-text-parallel-90k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Ga Speech-Text Parallel Dataset
Dataset Description
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-speech-text-parallel-90k.asante-twi-bible-speech-text
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
twi-trigrams-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.ghana-speech-eval
ghana-speech-eval
⚠️ Evaluation only. This is a held-out benchmark for measuring ASR quality on
Ghanaian languages. Please do not train or fine-tune on it — doing so contaminates
the benchmark and makes reported scores meaningless. Every subset ships a single eval split.
Multi-source ASR evaluation benchmark for Ghanaian languages. All subsets are
16 kHz mono with standardised schema (audio, text, language, country,
length, iso, subset).
Source groups… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-speech-eval.ghana-english-tts-clean2
Ghana English TTS Filtered Clean v2
Filtered subset of ghananlpcommunity/ghana-english-tts-filtered using PANNs CNN14.
Filtering
Second-pass filtering with PANNs CNN14 (soundclassifier with music, applause, and speech tags):
Keep if: music_prob ≤ 0.2 AND applause_prob ≤ 0.2 AND speech_prob ≥ 0.5
Batch size 32 on NVIDIA H200, float16 inference
282,096 / 303,204 kept (93.0%)
Fields
corrected_text: utterance text
bytes: raw WAV bytes (16-bit PCM… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-tts-clean2.ghana-speech-ipa
Ghana Speech — Audio with IPA Transcripts
Speech with both transcript forms: the original orthography and the IPA phoneme
sequence read off the audio by ASR. Each language is a subset, with real
train/validation splits.
from datasets import load_dataset
ds = load_dataset("ghanaopendata/ghana-speech-ipa", "Akuapem_Twi_twi", split="train")
ds[0]["audio"] # decoded waveform, 16 kHz
ds[0]["text"] # original orthography
ds[0]["ipa"] # IPA phonemes
369,347 clips · ~747 h… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-speech-ipa.navigation-corpus-dagbani-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Dag Speech Segments (sentence splitting)
52799 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani
Full-file CTC forced… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-dagbani-speech.navigation-corpus-twi-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Speech Segments (sentence splitting)
52562 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-twi
Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-twi-speech.twi-health-asr
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages,
sourced from publicly available video content on health and wellness.
Created by Mich-Seth Owusu and… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr.ewe-bible-audio-text-tts
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
48775 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ewe-bible-audio-text-tts.ghana-english-speech-ipa
Ghanaian English Speech — Audio with IPA Transcripts
Speech with both transcript forms: the original orthography and the IPA phoneme
sequence read off the audio by ASR. Each language is a subset, with real
train/validation splits.
from datasets import load_dataset
ds = load_dataset("ghanaopendata/ghana-english-speech-ipa", "English_eng", split="train")
ds[0]["audio"] # decoded waveform, 16 kHz
ds[0]["text"] # original orthography
ds[0]["ipa"] # IPA phonemes
52,855… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-speech-ipa.kumawood-speech-transcriptions
Kumawood Speech Transcriptions
Speech segments from Ghanaian films, each paired with the film's human-authored
English subtitle and a machine Twi transcript.
Total number of hours 249.8 hours
Fields
field
meaning
audio
16 kHz mono FLAC segment
text
English subtitle displayed during the segment (human-authored, recovered by OCR)
twi_text
Twi transcript from Google STT (ak) — machine output
twi_words_per_sec
transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.nsanku-tts-benchmark-audioghana-named-entities-tts-twi
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana Named Entities TTS — Twi
A Twi-language speech dataset built from descriptions of Ghana named entities
(people, places, organisations, and concepts). Each audio clip is a synthesised
reading of a passage that describes several… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-named-entities-tts-twi.new-twi-tts-aligned
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi TTS Dataset
A speech dataset of Twi (Akan) extracted from Ghanaian news media broadcasts, designed for training and fine-tuning Text-To-Speech (TTS) models.
📂 Dataset Structure
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned.twi-health-asr-gemini-500hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset Gemini (500 hours)
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken
languages, sourced from publicly available video content on health and wellness.
Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-health-asr-gemini-500hrs.dagbani-bible-audio-text-tts
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
53410 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/dagbani-bible-audio-text-tts.fante-speech-text-multispeaker_lds
Fante Speech-Text Multispeaker Dataset (LDS)
Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations.
Dataset Statistics
Split
Clips
Hours
Talks
Train
29,992
58.32
405
Eval
2,028
4.09
28
Total
32,020
62.41
433
Features
audio: 16 kHz mono FLAC sentence-level clips
text: Fante transcript (sentence-aligned)
talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-speech-text-multispeaker_lds.navigation-corpus-ewe-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ewe Speech Segments (sentence splitting)
49348 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/navigation-corpus-speech-full-ewe
Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-ewe-speech.kumawood-speech-transcriptions
Kumawood Speech Transcriptions
Speech segments from Ghanaian films, each paired with the film's human-authored
English subtitle and a machine Twi transcript.
Total number of hours 249.8 hours
Fields
field
meaning
audio
16 kHz mono FLAC segment
text
English subtitle displayed during the segment (human-authored, recovered by OCR)
twi_text
Twi transcript from Google STT (ak) — machine output
twi_words_per_sec
transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kumawood-speech-transcriptions.ghana-speech-eval
ghana-speech-eval
⚠️ Evaluation only. This is a held-out benchmark for measuring ASR quality on
Ghanaian languages. Please do not train or fine-tune on it — doing so contaminates
the benchmark and makes reported scores meaningless. Every subset ships a single eval split.
Multi-source ASR evaluation benchmark for Ghanaian languages. All subsets are
16 kHz mono with standardised schema (audio, text, language, country,
length, iso, subset).
Source groups… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-speech-eval.twi-agriculture-speech
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Agriculture Speech Dataset
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages,
sourced from publicly available video content on agriculture.Created by Mich-Seth Owusu and published… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-agriculture-speech.
