datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shona1shona-bible-bdsc-aligned
Shona Bible Speech Alignment Dataset
Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC
source audio made available by Biblica, Inc. through Open.Bible. This release
contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech
segments covering approximately 75.55 hours.
Dataset summary
Language: Shona (sna)
Speaker: narrator 1
Speaker sex: male
Books: 66
Clips: 31,284
Audio: approximately 75.55 hours
Audio format: mono 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-bible-bdsc-aligned.shona-waxal-pseudo-labeled
Shona WAXAL pseudo-labelled speech
This release contains 90,253 Shona speech clips, totalling 441.585 hours. Each clip keeps its original FLAC audio and a Sunbird Whisper pseudo-transcript. These are model outputs, not human reference transcriptions.
What this release contains
The source is the unlabeled Shona ASR split from WAXAL NLP, preserved in the operational checkpoint manassehzw/sna-waxal-annotated-unlabeled. The source checkpoint has no transcripts. This… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-waxal-pseudo-labeled.shona-speechshona-vibevoice-corpus
Shona VibeVoice Corpus
Unified 24 kHz mono Shona speech corpus in the VibeVoice segment schema, with a
mix of single- and multi-utterance rows (short clips concatenated with silence
gaps and multi-segment labels). Compiled by build_and_push.py from:
realtime-speech/shona1
google/WaxalNLP (sna ASR)
google/fleurs (sn_zw)
teeofftechnologies/badrex-shona-whisper-cleaned-16k (train + test)
Kittech/mixed_shona_dataset
realtime-speech/shona2 (train + test + validation)… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/shona-vibevoice-corpus.mixed_shona_datasetshona-synthetic-corpusDigitalUmuganda_AfriVoice_shonakittech_shona_datasetshona-speech-datasetinzwi-shona-sample
Inzwi — Shona Speech Corpus (Cleaned Sample)
A small, fully-documented sample of the Inzwi speech corpus: consented, peer-validated
Shona audio paired with ground-truth transcripts, prepared to be AI-ready for
automatic speech recognition (ASR). Built for the POTRAZ AI for Impact (AI4I) Challenge —
Data Track, to demonstrate the Inzwi data pipeline end to end.
This is a representative sample (the validated slice of an early, un-incentivised run),
not the full corpus. It exists… See the full description on the dataset page: https://huggingface.co/datasets/nigeLbasa/inzwi-shona-sample.Shona_test_5hr_v1shona2ruzivo-shona-tts-samples
