ghanaopenai/ghana-named-entities-tts-twi
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghana Named Entities TTS — Twi A Twi-language speech dataset built from descriptions of Ghana named entities (people, places, organisations, and concepts). Each audio clip is a synthesised reading of a passage that describes several… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-named-entities-tts-twi.
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana Named Entities TTS — Twi
A Twi-language speech dataset built from descriptions of Ghana named entities (people, places, organisations, and concepts). Each audio clip is a synthesised reading of a passage that describes several named entities.
Dataset description
Columns
Generation pipeline
The dataset was created through a four-step pipeline:
- Source collection —
ghana_named_entities.csvcontains ~253 000 Ghana named entities with English descriptions scraped and curated by GhanaNLP Community.
- Chunking — English descriptions were grouped into passages of 10 entities each using
combine-descriptions.py, producingcombined_descriptions.csv.
- Translation — Passages were translated into Twi with Gemini (
gemini-3.1-flash-lite-preview, thinking levelHIGH) viatranslate.py, yieldingtranslated_twi.csv.
- Speech synthesis — Twi passages were converted to speech with [OmniVoice](https://huggingface.co/k2-fsa/OmniVoice) (
k2-fsa/OmniVoice) running on a Modal T4 GPU viagenerate-speech.py: - Voice instruct:
female, young adult, moderate pitch - Language code:
abr(Akan/Twi dialect code accepted by OmniVoice) - Diffusion steps: 16, guidance scale: 2.0
- Output sample rate: 24 000 Hz
Usage
from datasets import load_dataset
ds = load_dataset("ghananlpcommunity/ghana-named-entities-tts-twi", split="train")
sample = ds[0]
print(sample["text"])
# sample["audio"]["array"] → numpy array at 24 000 HzIntended uses
- Training and evaluating Twi TTS systems
- Training and evaluating Twi ASR systems
- Low-resource African-language speech research
- Named-entity pronunciation modelling for Ghanaian proper nouns
Limitations
- Audio is synthetic (not human-recorded); prosody and pronunciation may contain TTS artefacts.
- Each clip covers multiple entities concatenated into a passage rather than isolated utterances.
- Translation was automated; some Twi renderings may not be idiomatic.
Citation
If you use this dataset please cite:
@dataset{ghananlp2024twi_tts_named_entities,
author = {GhanaNLP Community},
title = {Ghana Named Entities TTS -- Twi},
year = {2024},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/ghananlpcommunity/ghana-named-entities-tts-twi}
}License
This dataset is released under CC BY-NC 4.0 — free to use, share, and adapt for non-commercial research and educational purposes with attribution.
