ghananlpcommunity/kumawood-speech-transcriptions
Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.
Kumawood Speech Transcriptions
Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript.
Total number of hours 249.8 hours
Fields
Language direction
Audio is spoken Twi (Akan). text is an English subtitle, i.e. a translation rather than a transcription; twi_text is the machine transcript of what is actually said. The set therefore supports Twi ASR (audio + twi_text) and Twi-to-English speech translation (audio + text).
How segments were selected
- Subtitle appearance/disappearance times detected per frame, restricted to a centre-lower window so channel overlays do not trigger detection.
- Subtitle text read from the full frame by a vision-language model.
- Audio cut to the detected spans, 16 kHz mono.
- Non-dialogue English text removed: credits, cast lists, phone numbers, station branding, all-caps captions. Short ambiguous lines were judged individually by an LLM; longer lines kept.
- Duplicate uploads removed at film level, including parts contained in a full-movie cut.
- Only segments Google STT could transcribe are included; segments it returned nothing for are excluded.
- Segments whose Twi words-per-second fell outside 0.5-6.0 were dropped as implausible, along with extreme English subtitle/duration mismatches.
Splits
Held out by film and title family, never by clip: segments from one film share speakers, room acoustics and OCR quirks, so a clip-level split would leak evaluation audio into training. Held-out films are drawn from across the film-size range so evaluation is not concentrated in one production.
Caveats
twi_textis ASR output with a real error rate; it is not verified ground truth.textis OCR output and contains recognition errors.- Subtitles are condensed translations, so
textmay not cover everything said. - Segment boundaries follow the subtitle, not speech onset/offset.
