CoolFace
Datasetpublic

ghananlpcommunity/kumawood-speech-transcriptions

Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.

sourceHugging Facecc-by-nc-4.0updated 25d agoView on Hugging Face
0likes799downloads
Dataset Card

Kumawood Speech Transcriptions

Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript.

Total number of hours 249.8 hours

Fields

fieldmeaning
audio16 kHz mono FLAC segment
textEnglish subtitle displayed during the segment (human-authored, recovered by OCR)
twi_textTwi transcript from Google STT (ak) — machine output
twi_words_per_sectranscript words per second of audio
film, start, end, durationprovenance within the source film

Language direction

Audio is spoken Twi (Akan). text is an English subtitle, i.e. a translation rather than a transcription; twi_text is the machine transcript of what is actually said. The set therefore supports Twi ASR (audio + twi_text) and Twi-to-English speech translation (audio + text).

How segments were selected

  1. 1.Subtitle appearance/disappearance times detected per frame, restricted to a centre-lower window so channel overlays do not trigger detection.
  2. 2.Subtitle text read from the full frame by a vision-language model.
  3. 3.Audio cut to the detected spans, 16 kHz mono.
  4. 4.Non-dialogue English text removed: credits, cast lists, phone numbers, station branding, all-caps captions. Short ambiguous lines were judged individually by an LLM; longer lines kept.
  5. 5.Duplicate uploads removed at film level, including parts contained in a full-movie cut.
  6. 6.Only segments Google STT could transcribe are included; segments it returned nothing for are excluded.
  7. 7.Segments whose Twi words-per-second fell outside 0.5-6.0 were dropped as implausible, along with extreme English subtitle/duration mismatches.

Splits

Held out by film and title family, never by clip: segments from one film share speakers, room acoustics and OCR quirks, so a clip-level split would leak evaluation audio into training. Held-out films are drawn from across the film-size range so evaluation is not concentrated in one production.

Caveats

  • —twi_text is ASR output with a real error rate; it is not verified ground truth.
  • —text is OCR output and contains recognition errors.
  • —Subtitles are condensed translations, so text may not cover everything said.
  • —Segment boundaries follow the subtitle, not speech onset/offset.