CoolFace
Datasetpublic

ghananlpcommunity/ghana-female-twi-8sec-splits

This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 8-Word Speech Segments 25951 speech-text pairs split from 30-min recordings. Processing pipeline Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-female-twi-8sec-splits.

sourceHugging Facecc-by-nc-4.0updated 3mo agoView on Hugging Face
0likes195downloads
Dataset Card
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.

Twi 8-Word Speech Segments

25951 speech-text pairs split from 30-min recordings.

Processing pipeline

  1. 1.Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length
  2. 2.Full-file CTC forced alignment (MMS-300M) for word-level timestamps
  3. 3.Words grouped into 16-word (8-gram) segments
  4. 4.Leading/trailing silence trimmed with VAD (-40 dBFS threshold)
  5. 5.Filtered: min 1.0s, max 15.0s
  6. 6.Original sample rate preserved (24kHz)

Usage

python
from datasets import load_dataset
ds = load_dataset("ghananlpcommunity/ghana-female-twi-8sec-splits", split="train")