CoolFace
Datasetpublic

ghanaopenai/ghana-named-entities-tts-twi

This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghana Named Entities TTS — Twi A Twi-language speech dataset built from descriptions of Ghana named entities (people, places, organisations, and concepts). Each audio clip is a synthesised reading of a passage that describes several… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-named-entities-tts-twi.

sourceHugging Facecc-by-nc-4.0updated 3mo agoView on Hugging Face
0likes611downloads
Dataset Card
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.

Ghana Named Entities TTS — Twi

A Twi-language speech dataset built from descriptions of Ghana named entities (people, places, organisations, and concepts). Each audio clip is a synthesised reading of a passage that describes several named entities.

Dataset description

FieldValue
LanguageTwi (ISO 639-1: tw)
DomainNamed entities — Ghanaian proper nouns and institutions
Audio formatWAV, 24 000 Hz, mono
VoiceFemale, young adult, moderate pitch (synthetic)
Examples6 704
SplitSingle train split

Columns

ColumnTypeDescription
idstringUnique identifier (row_0, row_1, …)
audioAudioWAV file decoded at native 24 kHz
textstringTwi transcription aligned to the audio

Generation pipeline

The dataset was created through a four-step pipeline:

  1. 1.Source collection — ghana_named_entities.csv contains ~253 000 Ghana named entities with English descriptions scraped and curated by GhanaNLP Community.
  1. 1.Chunking — English descriptions were grouped into passages of 10 entities each using combine-descriptions.py, producing combined_descriptions.csv.
  1. 1.Translation — Passages were translated into Twi with Gemini (gemini-3.1-flash-lite-preview, thinking level HIGH) via translate.py, yielding translated_twi.csv.
  1. 1.Speech synthesis — Twi passages were converted to speech with [OmniVoice](https://huggingface.co/k2-fsa/OmniVoice) (k2-fsa/OmniVoice) running on a Modal T4 GPU via generate-speech.py:
  2. 2.Voice instruct: female, young adult, moderate pitch
  3. 3.Language code: abr (Akan/Twi dialect code accepted by OmniVoice)
  4. 4.Diffusion steps: 16, guidance scale: 2.0
  5. 5.Output sample rate: 24 000 Hz

Usage

python
from datasets import load_dataset

ds = load_dataset("ghananlpcommunity/ghana-named-entities-tts-twi", split="train")
sample = ds[0]
print(sample["text"])
# sample["audio"]["array"] → numpy array at 24 000 Hz

Intended uses

  • —Training and evaluating Twi TTS systems
  • —Training and evaluating Twi ASR systems
  • —Low-resource African-language speech research
  • —Named-entity pronunciation modelling for Ghanaian proper nouns

Limitations

  • —Audio is synthetic (not human-recorded); prosody and pronunciation may contain TTS artefacts.
  • —Each clip covers multiple entities concatenated into a passage rather than isolated utterances.
  • —Translation was automated; some Twi renderings may not be idiomatic.

Citation

If you use this dataset please cite:

bibtex
@dataset{ghananlp2024twi_tts_named_entities,
  author    = {GhanaNLP Community},
  title     = {Ghana Named Entities TTS -- Twi},
  year      = {2024},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/datasets/ghananlpcommunity/ghana-named-entities-tts-twi}
}

License

This dataset is released under CC BY-NC 4.0 — free to use, share, and adapt for non-commercial research and educational purposes with attribution.