shona
Datasets
All datasets matching “shona”shona1shona-bible-bdsc-aligned
Shona Bible Speech Alignment Dataset
Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC
source audio made available by Biblica, Inc. through Open.Bible. This release
contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech
segments covering approximately 75.55 hours.
Dataset summary
Language: Shona (sna)
Speaker: narrator 1
Speaker sex: male
Books: 66
Clips: 31,284
Audio: approximately 75.55 hours
Audio format: mono 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-bible-bdsc-aligned.shona-waxal-pseudo-labeled
Shona WAXAL pseudo-labelled speech
This release contains 90,253 Shona speech clips, totalling 441.585 hours. Each clip keeps its original FLAC audio and a Sunbird Whisper pseudo-transcript. These are model outputs, not human reference transcriptions.
What this release contains
The source is the unlabeled Shona ASR split from WAXAL NLP, preserved in the operational checkpoint manassehzw/sna-waxal-annotated-unlabeled. The source checkpoint has no transcripts. This… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-waxal-pseudo-labeled.shona-speechshona-vibevoice-corpus
Shona VibeVoice Corpus
Unified 24 kHz mono Shona speech corpus in the VibeVoice segment schema, with a
mix of single- and multi-utterance rows (short clips concatenated with silence
gaps and multi-segment labels). Compiled by build_and_push.py from:
realtime-speech/shona1
google/WaxalNLP (sna ASR)
google/fleurs (sn_zw)
teeofftechnologies/badrex-shona-whisper-cleaned-16k (train + test)
Kittech/mixed_shona_dataset
realtime-speech/shona2 (train + test + validation)… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/shona-vibevoice-corpus.mixed_shona_dataset
