CoolFace
Datasetpublic

fiifinketia/dagbani-bible-audio-text-tts

Twi 16-Word Speech Segments 53410 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for word-level timestamps Words grouped into 16-word segments Leading/trailing silence trimmed with VAD (-40 dBFS threshold) Filtered: min 1.0s, max 15.0s Original sample rate preserved (24kHz) Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/dagbani-bible-audio-text-tts.

sourceHugging Facecc-by-4.0updated 6mo agoView on Hugging Face
0likes20downloads
Dataset Card

Twi 16-Word Speech Segments

53410 speech-text pairs split from long recordings.

Processing pipeline

  1. 1.Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text
  2. 2.Full-file CTC forced alignment (MMS-300M) for word-level timestamps
  3. 3.Words grouped into 16-word segments
  4. 4.Leading/trailing silence trimmed with VAD (-40 dBFS threshold)
  5. 5.Filtered: min 1.0s, max 15.0s
  6. 6.Original sample rate preserved (24kHz)

Usage

python
from datasets import load_dataset
ds = load_dataset("ghananlpcommunity/dagbani-bible-audio-text-tts", split="train")