datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shona1shona-bible-bdsc-aligned
Shona Bible Speech Alignment Dataset
Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC
source audio made available by Biblica, Inc. through Open.Bible. This release
contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech
segments covering approximately 75.55 hours.
Dataset summary
Language: Shona (sna)
Speaker: narrator 1
Speaker sex: male
Books: 66
Clips: 31,284
Audio: approximately 75.55 hours
Audio format: mono 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-bible-bdsc-aligned.shona-waxal-pseudo-labeled
Shona WAXAL pseudo-labelled speech
This release contains 90,253 Shona speech clips, totalling 441.585 hours. Each clip keeps its original FLAC audio and a Sunbird Whisper pseudo-transcript. These are model outputs, not human reference transcriptions.
What this release contains
The source is the unlabeled Shona ASR split from WAXAL NLP, preserved in the operational checkpoint manassehzw/sna-waxal-annotated-unlabeled. The source checkpoint has no transcripts. This… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-waxal-pseudo-labeled.shona-speechshona-vibevoice-corpus
Shona VibeVoice Corpus
Unified 24 kHz mono Shona speech corpus in the VibeVoice segment schema, with a
mix of single- and multi-utterance rows (short clips concatenated with silence
gaps and multi-segment labels). Compiled by build_and_push.py from:
realtime-speech/shona1
google/WaxalNLP (sna ASR)
google/fleurs (sn_zw)
teeofftechnologies/badrex-shona-whisper-cleaned-16k (train + test)
Kittech/mixed_shona_dataset
realtime-speech/shona2 (train + test + validation)… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/shona-vibevoice-corpus.mixed_shona_datasetshona-synthetic-corpusshona-correctionsDigitalUmuganda_AfriVoice_shonaalpaca_shona_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_shona_taco.kittech_shona_datasetCode-170k-shona
Dataset Description
Code-170k-shona is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Shona, making coding education accessible to Shona speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Shona language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-shona.shona-hate-speech
Dataset Card for Balanced Shona Hate Speech Dataset
Dataset Summary
This dataset contains 2,000 balanced examples of Shona text classified into four categories: NEUTRAL, OFFENSIVE, CONTEXTUAL, and HATE.
Data Sources
Label
Source
Count
NEUTRAL
Literary novel (Imbwa Yemunhu by Ignatius T. Mabasa)
500
OFFENSIVE
Synthetic template-based generation
500
CONTEXTUAL
Synthetic (quoted hate speech, not endorsed)
500
HATE
Synthetic (direct attacks on… See the full description on the dataset page: https://huggingface.co/datasets/omanyasa/shona-hate-speech.inzwi-shona-corpusruzivo-shona-ragshona-speech-datasethausa-shona_sentence-pairs
Hausa-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Shona_Sentence-Pairs
Number of Rows: 829704
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-shona_sentence-pairs.amharic-shona_sentence-pairs
Amharic-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Shona_Sentence-Pairs
Number of Rows: 856822
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-shona_sentence-pairs.inzwi-shona-sample
Inzwi — Shona Speech Corpus (Cleaned Sample)
A small, fully-documented sample of the Inzwi speech corpus: consented, peer-validated
Shona audio paired with ground-truth transcripts, prepared to be AI-ready for
automatic speech recognition (ASR). Built for the POTRAZ AI for Impact (AI4I) Challenge —
Data Track, to demonstrate the Inzwi data pipeline end to end.
This is a representative sample (the validated slice of an early, un-incentivised run),
not the full corpus. It exists… See the full description on the dataset page: https://huggingface.co/datasets/nigeLbasa/inzwi-shona-sample.shona-umbundu_sentence-pairs
Shona-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Umbundu_Sentence-Pairs
Number of Rows: 153467
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-umbundu_sentence-pairs.scales_shona_aishona_asroromo-shona_sentence-pairs
Oromo-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Shona_Sentence-Pairs
Number of Rows: 102932
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-shona_sentence-pairs.english-shona_sentence-pairs_mt560
English-Shona Parallel Dataset
This dataset contains parallel sentences in English and Shona (Zimbabwe).
Dataset Information
Language Pair: English ↔ Shona
Language Code: sna
Country: Zimbabwe
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset, please cite the citation… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-shona_sentence-pairs_mt560.Shona_test_5hr_v1Shona_test_5hrshona2argo-core-parquet-2021-2025alpaca-shona-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-shona-cleaned.fulah-shona_sentence-pairs
Fulah-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fulah-Shona_Sentence-Pairs
Number of Rows: 122877
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-shona_sentence-pairs.
