datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ciempiess_light
Dataset Card for ciempiess_light
Dataset Summary
The CIEMPIESS LIGHT is a Radio Corpus designed to create acoustic models for automatic speech recognition and it is made up by recordings of spontaneous conversations in Mexican Spanish between a radio moderator and his guests. It is an enhanced version of the CIEMPIESS Corpus (LDC item LDC2015S07).
CIEMPIESS LIGHT is "light" because it doesn't include much of the files of the first version of CIEMPIESS and it is "enhanced"… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_light.librivox_spanish
Dataset Card for librivox_spanish
Dataset Summary
Librivox is a non-commercial, non-profit and ad-free project that is dedicated to make all books in the public domain available, for free, in audio format on the internet. According to this, we downloaded 300 titles in Spanish to create the LIBRIVOX SPANISH CORPUS.
The LIBRIVOX SPANISH CORPUS has a duration of 73 hours and it is constituted by audio files between 3 and 10 seconds long, manually segmented. Transcription are… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/librivox_spanish.Vedavani-Dataset
Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
Vedavani is the first benchmark dataset for automatic speech recognition (ASR) on Vedic Sanskrit poetry, consisting of richly annotated verses from the Rig Veda and Atharva Veda. This corpus captures the unique prosodic structure, phonetic complexity, and chanting style found in traditional Vedic recitation.
🔗 Paper: Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry (ACL 2025)📁 GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/CielAstral/Vedavani-Dataset.voxforge_spanish
Dataset Card for voxforge_spanish
Dataset Summary
VoxForge was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac). They promise they will make available all submitted audio files under the GPL license, and then 'compile' them into acoustic models for use with Open Source speech recognition engines such as CMU Sphinx, ISIP, Julius and HTK. According to this, we downloaded the Spanish recordings of… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/voxforge_spanish.ciempiess_testThe CIEMPIESS TEST Corpus is a gender balanced corpus destined to test acoustic models for the speech recognition task. The corpus was manually transcribed and it contains audio recordings from 10 male and 10 female speakers. The CIEMPIESS TEST is one of the three corpora included at the LDC's \"CIEMPIESS Experimentation\" (LDC2019S07).wikipedia_spanish
Dataset Card for wikipedia_spanish
Dataset Summary
According to the project page of the WikiProject Spoken Wikipedia:
The WikiProject Spoken Wikipedia aims to produce recordings of Wikipedia articles being read aloud. Therefore, the WIKIPEDIA SPANISH CORPUS is a dataset created from the Spanish version of the WikiProject Spoken Wikipedia, called Wikipedia Grabada
The WIKIPEDIA SPANISH CORPUS aims to be used in the Automatic Speech Recognition (ASR) task. It is a gender… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/wikipedia_spanish.tele_con_ciencia
Dataset Card for tele_con_ciencia
Dataset Summary
According to the Facebook page of Tele con Ciencia:
"Nuestra misión es la comunicación pública de la ciencia y la tecnología mexicana. El objetivo,
la participación activa de todos los mexicanos en las áreas del descubrimiento científico y el
desarrollo tecnológico."
"Our mission is to spread the achievements of the Mexican Science and Technology. The main goal
is to promote the active participation of mexican people in… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/tele_con_ciencia.tedx_spanish
Dataset Card for tedx_spanish
Dataset Summary
According to the TEDx website:
In the spirit of ideas worth spreading, TEDx is a program of local, self-organized events that
bring people together to share a TED-like experience. At a TEDx event, TEDTalks video and live
speakers combine to spark deep discussion and connection in a small group. These local,
self-organized events are branded TEDx, where x = independently organized TED event. The TED
Conference provides… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/tedx_spanish.ciempiess_complementary
Dataset Card for ciempiess_complementary
Dataset Summary
The CIEMPIESS COMPLEMENTARY is a phonetically balanced corpus of isolated Spanish words spoken by people of Central Mexico. It was designed to solve one particular issue when training automatic speech recognition (ASR) systems in the Spanish of Central Mexico. This problem appears when someone collects some training data, but the system complains because it does not find enough instances of one or more particular… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_complementary.swiss-german-city-sentences_v2
Swiss German City Sentences v2
Synthetic Swiss German speech dataset with city name sentences across multiple dialects.
ciempiess_balance
Dataset Card for ciempiess_balance
Dataset Summary
The CIEMPIESS BALANCE Corpus is designed to match with the CIEMPIESS LIGHT Corpus (LDC2017S23). So, "Balance" means that if the CIEMPIESS BALANCE is combined with the CIEMPIESS LIGHT, one will get a gender balanced corpus. To appreciate this, one need to know that the CIEMPIESS LIGHT is by itself, a gender unbalanced corpus of approximately 25% of female speakers and 75% of male speakers. So, the CIEMPIESS BALANCE is a… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_balance.ciempiess_fem
Dataset Card for ciempiess_fem
Dataset Summary
Since the publication of the CIEMPIESS Corpus (LDC2015S07) in 2015 we have noticed that there is a lack of female speakers in the sources where we traditionally take audio to create new CIEMPIESS datasets. That is why we decided to create a corpus that helps to balance future gender unbalanced datasets.
The CIEMPIESS FEM Corpus was created by recordings and human transcripts of 21 different women. 16 of these women are… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_fem.ciempiess_light
Dataset Card for ciempiess_light
Dataset Summary
The CIEMPIESS LIGHT is a Radio Corpus designed to create acoustic models for automatic speech recognition and it is made up by recordings of spontaneous conversations in Mexican Spanish between a radio moderator and his guests. It is an enhanced version of the CIEMPIESS Corpus (LDC item LDC2015S07).
CIEMPIESS LIGHT is "light" because it doesn't include much of the files of the first version of CIEMPIESS and it is "enhanced"… See the full description on the dataset page: https://huggingface.co/datasets/YunisAFS/ciempiess_light.ePark_xue_xi_ci_biao_learning_vocabulary
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary.moore_audio_data~70h of Mooré paired audio+text data from jw.org and jwsoup
Headers: audio, text
