datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.Wikimedia-Speech-Irish
Dataset Details
Synthetic audio dataset, created using Azure text-to-speech service.
The bilingual text is a portion of the Wikimedia dataset, consisting of 7,545 text segments.
The dataset includes two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural).
The speech data comprises approximately 34 hours and 23 minutes (34:23:12) spread across 15,090 utterances.
Dataset Structure
Dataset({
features: ['audio', 'text_ga'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Wikimedia-Speech-Irish.The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden
Dataset Card for "The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden"
Homepage: Spoken Wikipedia CorporaRepository: Bitbucket RepositoryPaper: Publication at LREC 2016Leaderboard: Interspeech 2018 PaperPoint of Contact: nats@nats.gitlab.io
Dataset Summary
The Spoken Wikipedia Corpora (SWC) is a collection of aligned spoken Wikipedia articles including articles in Dutch. It includes approximately 210 hours of audio, transcriptions, and metadata. The corpus is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Jaspernl/The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden.wikipedia_spanish
Dataset Card for wikipedia_spanish
Dataset Summary
According to the project page of the WikiProject Spoken Wikipedia:
The WikiProject Spoken Wikipedia aims to produce recordings of Wikipedia articles being read aloud. Therefore, the WIKIPEDIA SPANISH CORPUS is a dataset created from the Spanish version of the WikiProject Spoken Wikipedia, called Wikipedia Grabada
The WIKIPEDIA SPANISH CORPUS aims to be used in the Automatic Speech Recognition (ASR) task. It is a gender… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/wikipedia_spanish.Segmented-Moroccan-Darija-Wiki-Audio-Dataset
Dataset Card for Segmented Moroccan Darija Wiki Dataset
Dataset Summary
This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon).
Each audio is split into segments of up to 30 seconds to make it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/anaszil/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.WikIPA
WikIPA
Dataset Description
WikIPA is a multilingual benchmark dataset designed for speech-to-IPA (STIPA) transcription, linking spoken audio with International Phonetic Alphabet (IPA) transcriptions.
The dataset integrates two large-scale community-driven resources:
WikiPron — human-curated IPA pronunciations extracted from Wiktionary
Lingua Libre — crowdsourced recordings of spoken lexical items
By connecting these two resources, WikIPA provides a dataset that links… See the full description on the dataset page: https://huggingface.co/datasets/pierluigic/WikIPA.wikitongues-darija
Wikitongues-Darija
This is a small test dataset for Automatic Speech Recognition in Darija language, built from 2 captioned videos of the WikiTongues project:
nawal
anass
Process:
each webm video has been converted to monochannel 16khz wav files with ffmpeg :
ffmpeg -i WIKITONGUES-_Nawal_speaking_Moroccan_Arabic.webm.1080p.vp9.webm -ar 16000 -ac 1 nawal.wav
each audio has been cut in samples of less than 30 seconds audio according to the captions timestamps. The script may be… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/wikitongues-darija.Segmented-Moroccan-Darija-Wiki-Audio-Dataset
Dataset Card for Segmented Moroccan Darija Wiki Dataset
Dataset Summary
This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon).
Each audio is split into segments of up to 30 seconds to make it suitable… See the full description on the dataset page: https://huggingface.co/datasets/H20-sys/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.
