datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.CommonVoices20_ro
Common Voices Corpus 20.0 (Romanian)
Common Voices is an open-source dataset of speech recordings created by
Mozilla to improve speech recognition technologies.
It consists of crowdsourced voice samples in multiple languages, contributed by volunteers worldwide.
Challenges: The raw dataset included numerous recordings with incorrect transcriptions
or those requiring adjustments, such as sampling rate modifications, conversion to .wav format, and other refinements
essential… See the full description on the dataset page: https://huggingface.co/datasets/TransferRapid/CommonVoices20_ro.commonvoicebadini
Northern Kurdish (Arabic Script) ASR Dataset
Dataset Description
Northern Kurdish is the most widely spoken variant of the Kurdish language and is used across all parts of Kurdistan. Although it is mainly written today in the Latin script, it was historically written in the Arabic script. The Arabic script is still used for this dialect in Southern Kurdistan, particularly in the Duhok province of the Kurdistan Regional Government (KRG).Similarly, the primary writing… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/commonvoicebadini.
