datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Custom_common_voice_dataset_using_RVC
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
Custom common_voice_v11 corpus with a custom voice was was created using RVC(Retrieval-Based Voice Conversion)
The model underwent 200 epochs of training, utilizing a total of 1 hour of audio clips. The data was scraped from Youtube.
The audio in the custom generated dataset is of a YouTuber named
Ajay Pandey
Description
license: cc0-1.0
language:
- hi… See the full description on the dataset page: https://huggingface.co/datasets/Aniket-Tathe-08/Custom_common_voice_dataset_using_RVC.hausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset
common-voice-kinyarwanda-english-dataset
Kinyarwanda-English Commonvoice dataset
A compilation of Kinyarwanda-english dataset to be used to train multi-lingual ASR
Note: The audio dataset shall be added in the future
commonvoice-mnCommonVoices20_ro
Common Voices Corpus 20.0 (Romanian)
Common Voices is an open-source dataset of speech recordings created by
Mozilla to improve speech recognition technologies.
It consists of crowdsourced voice samples in multiple languages, contributed by volunteers worldwide.
Challenges: The raw dataset included numerous recordings with incorrect transcriptions
or those requiring adjustments, such as sampling rate modifications, conversion to .wav format, and other refinements
essential… See the full description on the dataset page: https://huggingface.co/datasets/TransferRapid/CommonVoices20_ro.mozilla-common-voice-23-bel-texts-exportcommonvoicebadini
Northern Kurdish (Arabic Script) ASR Dataset
Dataset Description
Northern Kurdish is the most widely spoken variant of the Kurdish language and is used across all parts of Kurdistan. Although it is mainly written today in the Latin script, it was historically written in the Arabic script. The Arabic script is still used for this dialect in Southern Kurdistan, particularly in the Duhok province of the Kurdistan Regional Government (KRG).Similarly, the primary writing… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/commonvoicebadini.CommonVoiceAzcommon_voice_13_0_zh_pseudo_labelledcommon-voicecommon_voice_13_0_mn_pseudo_test_smallcommon_voice_13_0_thai_small_pseudo_labelledcommon_voice_16_1_spanish_test_set
Dataset Card for Common Voice Corpus 16 Spanish Dataset
Acknowledgement
The dataset belongs to COMMON VOICE MOZILLA FOUNDATION.
I just uploaded the spanish test set (from HERE : https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1/tree/main)
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Languages
Spanish
How to use
The datasets library allows you to load and pre-process… See the full description on the dataset page: https://huggingface.co/datasets/omarsou/common_voice_16_1_spanish_test_set.commonvoice_without_audiocommon_voice_16_1_hi_pseudo_labelledcommon_voice_19_uk_croppedcommon_voice_16_0_fa_pseudo_labelledcommon_voice_16_1_sw_pseudo_labelledcommon_voice_16_1_sw2_pseudo_labelledcommon_voice_16_1_zh-CN_pseudo_labelled_3common_voice_13_0_ru_pseudo_labelledcommon_voice_17_0_es_pseudo_labelledcommon_voice_17_0_pseudo_labelled_frCommon-voice-urdu-11CommonVoice17-Clone
