CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01realtime-speech /shona1audio10K<n<100K1 likes407 downloads2y agoHugging Face02manassehzw /shona-bible-bdsc-aligned Shona Bible Speech Alignment Dataset Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC source audio made available by Biblica, Inc. through Open.Bible. This release contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech segments covering approximately 75.55 hours. Dataset summary Language: Shona (sna) Speaker: narrator 1 Speaker sex: male Books: 66 Clips: 31,284 Audio: approximately 75.55 hours Audio format: mono 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-bible-bdsc-aligned.audioautomatic-speech-recognition10K<n<100K3 likes267 downloads17d agoHugging Face03manassehzw /shona-waxal-pseudo-labeled Shona WAXAL pseudo-labelled speech This release contains 90,253 Shona speech clips, totalling 441.585 hours. Each clip keeps its original FLAC audio and a Sunbird Whisper pseudo-transcript. These are model outputs, not human reference transcriptions. What this release contains The source is the unlabeled Shona ASR split from WAXAL NLP, preserved in the operational checkpoint manassehzw/sna-waxal-annotated-unlabeled. The source checkpoint has no transcripts. This… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-waxal-pseudo-labeled.audioautomatic-speech-recognition10K<n<100K3 likes190 downloads25d agoHugging Face04badrex /shona-speechaudio10K<n<100K8 likes175 downloads11mo agoHugging Face05takuM23 /shona-vibevoice-corpus Shona VibeVoice Corpus Unified 24 kHz mono Shona speech corpus in the VibeVoice segment schema, with a mix of single- and multi-utterance rows (short clips concatenated with silence gaps and multi-segment labels). Compiled by build_and_push.py from: realtime-speech/shona1 google/WaxalNLP (sna ASR) google/fleurs (sn_zw) teeofftechnologies/badrex-shona-whisper-cleaned-16k (train + test) Kittech/mixed_shona_dataset realtime-speech/shona2 (train + test + validation)… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/shona-vibevoice-corpus.audioautomatic-speech-recognition10K<n<100K0 likes135 downloads3mo agoHugging Face06Kittech /mixed_shona_datasetaudiotext-generationn<1K4 likes83 downloads2y agoHugging Face07Starsm91 /shona-synthetic-corpusaudion<1K0 likes82 downloads19d agoHugging Face08Starsm91 /shona-correctionstextn<1K0 likes68 downloads22d agoHugging Face09Beijuka /DigitalUmuganda_AfriVoice_shonaaudio10K<n<100K2 likes67 downloads2y agoHugging Face10saillab /alpaca_shona_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_shona_taco.text10K<n<100K5 likes48 downloads2y agoHugging Face11Kittech /kittech_shona_datasetaudiotranslationn<1K2 likes36 downloads2y agoHugging Face12michsethowusu /Code-170k-shona Dataset Description Code-170k-shona is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Shona, making coding education accessible to Shona speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Shona language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-shona.texttext-generation100K<n<1M7 likes36 downloads11mo agoHugging Face13omanyasa /shona-hate-speech Dataset Card for Balanced Shona Hate Speech Dataset Dataset Summary This dataset contains 2,000 balanced examples of Shona text classified into four categories: NEUTRAL, OFFENSIVE, CONTEXTUAL, and HATE. Data Sources Label Source Count NEUTRAL Literary novel (Imbwa Yemunhu by Ignatius T. Mabasa) 500 OFFENSIVE Synthetic template-based generation 500 CONTEXTUAL Synthetic (quoted hate speech, not endorsed) 500 HATE Synthetic (direct attacks on… See the full description on the dataset page: https://huggingface.co/datasets/omanyasa/shona-hate-speech.texttext-classification1K<n<10K0 likes35 downloads5mo agoHugging Face14mmakanda /inzwi-shona-corpustabular100K<n<1M0 likes35 downloads3mo agoHugging Face15cybux /ruzivo-shona-ragtext100K<n<1M1 likes31 downloads3mo agoHugging Face16shunyalabs /shona-speech-datasetaudio1K<n<10K2 likes28 downloads1y agoHugging Face17michsethowusu /hausa-shona_sentence-pairs Hausa-Shona_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Hausa-Shona_Sentence-Pairs Number of Rows: 829704 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-shona_sentence-pairs.text100K<n<1M0 likes25 downloads1y agoHugging Face18michsethowusu /amharic-shona_sentence-pairs Amharic-Shona_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Shona_Sentence-Pairs Number of Rows: 856822 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-shona_sentence-pairs.text100K<n<1M0 likes25 downloads1y agoHugging Face19nigeLbasa /inzwi-shona-sample Inzwi — Shona Speech Corpus (Cleaned Sample) A small, fully-documented sample of the Inzwi speech corpus: consented, peer-validated Shona audio paired with ground-truth transcripts, prepared to be AI-ready for automatic speech recognition (ASR). Built for the POTRAZ AI for Impact (AI4I) Challenge — Data Track, to demonstrate the Inzwi data pipeline end to end. This is a representative sample (the validated slice of an early, un-incentivised run), not the full corpus. It exists… See the full description on the dataset page: https://huggingface.co/datasets/nigeLbasa/inzwi-shona-sample.audioautomatic-speech-recognitionn<1K0 likes23 downloads2mo agoHugging Face20michsethowusu /shona-umbundu_sentence-pairs Shona-Umbundu_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Shona-Umbundu_Sentence-Pairs Number of Rows: 153467 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-umbundu_sentence-pairs.text100K<n<1M0 likes22 downloads1y agoHugging Face21scaleszw /scales_shona_aitextn<1K1 likes20 downloads2y agoHugging Face22realtime-speech /shona_asrtextn<1K0 likes20 downloads2y agoHugging Face23michsethowusu /oromo-shona_sentence-pairs Oromo-Shona_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Oromo-Shona_Sentence-Pairs Number of Rows: 102932 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-shona_sentence-pairs.text100K<n<1M0 likes19 downloads1y agoHugging Face24michsethowusu /english-shona_sentence-pairs_mt560 English-Shona Parallel Dataset This dataset contains parallel sentences in English and Shona (Zimbabwe). Dataset Information Language Pair: English ↔ Shona Language Code: sna Country: Zimbabwe Original Source: OPUS MT560 Dataset Dataset Structure The dataset contains parallel sentences that can be used for: Machine translation training Cross-lingual NLP tasks Language model fine-tuning Citation If you use this dataset, please cite the citation… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-shona_sentence-pairs_mt560.text100K<n<1M3 likes19 downloads1y agoHugging Face25Beijuka /Shona_test_5hr_v1audio1K<n<10K1 likes18 downloads2y agoHugging Face26Beijuka /Shona_test_5hrtext1K<n<10K0 likes17 downloads2y agoHugging Face27realtime-speech /shona2audio1K<n<10K1 likes17 downloads2y agoHugging Face28shonadoyle /argo-core-parquet-2021-2025tabular100M<n<1B0 likes17 downloads3mo agoHugging Face29saillab /alpaca-shona-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-shona-cleaned.text10K<n<100K1 likes16 downloads2y agoHugging Face30michsethowusu /fulah-shona_sentence-pairs Fulah-Shona_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Fulah-Shona_Sentence-Pairs Number of Rows: 122877 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-shona_sentence-pairs.text100K<n<1M0 likes15 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.