datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ciempiess_light
Dataset Card for ciempiess_light
Dataset Summary
The CIEMPIESS LIGHT is a Radio Corpus designed to create acoustic models for automatic speech recognition and it is made up by recordings of spontaneous conversations in Mexican Spanish between a radio moderator and his guests. It is an enhanced version of the CIEMPIESS Corpus (LDC item LDC2015S07).
CIEMPIESS LIGHT is "light" because it doesn't include much of the files of the first version of CIEMPIESS and it is "enhanced"… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_light.librivox_spanish
Dataset Card for librivox_spanish
Dataset Summary
Librivox is a non-commercial, non-profit and ad-free project that is dedicated to make all books in the public domain available, for free, in audio format on the internet. According to this, we downloaded 300 titles in Spanish to create the LIBRIVOX SPANISH CORPUS.
The LIBRIVOX SPANISH CORPUS has a duration of 73 hours and it is constituted by audio files between 3 and 10 seconds long, manually segmented. Transcription are… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/librivox_spanish.voxforge_spanish
Dataset Card for voxforge_spanish
Dataset Summary
VoxForge was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac). They promise they will make available all submitted audio files under the GPL license, and then 'compile' them into acoustic models for use with Open Source speech recognition engines such as CMU Sphinx, ISIP, Julius and HTK. According to this, we downloaded the Spanish recordings of… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/voxforge_spanish.ciempiess_testThe CIEMPIESS TEST Corpus is a gender balanced corpus destined to test acoustic models for the speech recognition task. The corpus was manually transcribed and it contains audio recordings from 10 male and 10 female speakers. The CIEMPIESS TEST is one of the three corpora included at the LDC's \"CIEMPIESS Experimentation\" (LDC2019S07).wikipedia_spanish
Dataset Card for wikipedia_spanish
Dataset Summary
According to the project page of the WikiProject Spoken Wikipedia:
The WikiProject Spoken Wikipedia aims to produce recordings of Wikipedia articles being read aloud. Therefore, the WIKIPEDIA SPANISH CORPUS is a dataset created from the Spanish version of the WikiProject Spoken Wikipedia, called Wikipedia Grabada
The WIKIPEDIA SPANISH CORPUS aims to be used in the Automatic Speech Recognition (ASR) task. It is a gender… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/wikipedia_spanish.tele_con_ciencia
Dataset Card for tele_con_ciencia
Dataset Summary
According to the Facebook page of Tele con Ciencia:
"Nuestra misión es la comunicación pública de la ciencia y la tecnología mexicana. El objetivo,
la participación activa de todos los mexicanos en las áreas del descubrimiento científico y el
desarrollo tecnológico."
"Our mission is to spread the achievements of the Mexican Science and Technology. The main goal
is to promote the active participation of mexican people in… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/tele_con_ciencia.tedx_spanish
Dataset Card for tedx_spanish
Dataset Summary
According to the TEDx website:
In the spirit of ideas worth spreading, TEDx is a program of local, self-organized events that
bring people together to share a TED-like experience. At a TEDx event, TEDTalks video and live
speakers combine to spark deep discussion and connection in a small group. These local,
self-organized events are branded TEDx, where x = independently organized TED event. The TED
Conference provides… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/tedx_spanish.ciempiess_complementary
Dataset Card for ciempiess_complementary
Dataset Summary
The CIEMPIESS COMPLEMENTARY is a phonetically balanced corpus of isolated Spanish words spoken by people of Central Mexico. It was designed to solve one particular issue when training automatic speech recognition (ASR) systems in the Spanish of Central Mexico. This problem appears when someone collects some training data, but the system complains because it does not find enough instances of one or more particular… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_complementary.swiss-german-city-sentences_v2
Swiss German City Sentences v2
Synthetic Swiss German speech dataset with city name sentences across multiple dialects.
ciempiess_balance
Dataset Card for ciempiess_balance
Dataset Summary
The CIEMPIESS BALANCE Corpus is designed to match with the CIEMPIESS LIGHT Corpus (LDC2017S23). So, "Balance" means that if the CIEMPIESS BALANCE is combined with the CIEMPIESS LIGHT, one will get a gender balanced corpus. To appreciate this, one need to know that the CIEMPIESS LIGHT is by itself, a gender unbalanced corpus of approximately 25% of female speakers and 75% of male speakers. So, the CIEMPIESS BALANCE is a… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_balance.ciempiess_fem
Dataset Card for ciempiess_fem
Dataset Summary
Since the publication of the CIEMPIESS Corpus (LDC2015S07) in 2015 we have noticed that there is a lack of female speakers in the sources where we traditionally take audio to create new CIEMPIESS datasets. That is why we decided to create a corpus that helps to balance future gender unbalanced datasets.
The CIEMPIESS FEM Corpus was created by recordings and human transcripts of 21 different women. 16 of these women are… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_fem.ciempiess_light
Dataset Card for ciempiess_light
Dataset Summary
The CIEMPIESS LIGHT is a Radio Corpus designed to create acoustic models for automatic speech recognition and it is made up by recordings of spontaneous conversations in Mexican Spanish between a radio moderator and his guests. It is an enhanced version of the CIEMPIESS Corpus (LDC item LDC2015S07).
CIEMPIESS LIGHT is "light" because it doesn't include much of the files of the first version of CIEMPIESS and it is "enhanced"… See the full description on the dataset page: https://huggingface.co/datasets/YunisAFS/ciempiess_light.az-asr-youtube-136h
Labelling
field
value
label_origin
asr:google
speech_register
spontaneous
channel
wideband-16k
provenance
documented
Google ASR output over YouTube audio. No human labels anywhere in it.
label_origin distinguishes text that existed before the audio (script,
exact by construction) from text written by a listener (human-transcript, high
but edited) from machine output (asr:<vendor>, bounded by that vendor's error
rate). The older label_type field is retained… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-youtube-136h.az-asr-calls-190h
Azerbaijani call-centre speech, Soniox pseudo-labels
Real Azerbaijani call-centre audio transcribed by Soniox
stt-async-v5. 169,273 clips, 190.63 h, 8 kHz telephony.
The labels are machine output and carry a measured ceiling. Soniox scores
46.65% WER against human transcripts of this same kind of audio. A model
trained on these labels learns to agree with Soniox -- including where Soniox is
wrong, and including its habit of dropping words.
That is not hypothetical. Scoring the… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-calls-190h.az-asr-calls-human-140h
Azerbaijani call-centre speech, human-transcribed
Real Azerbaijani call-centre audio with transcripts written by people
listening to it. 29,982 clips, 139.75 h, 8 kHz telephony.
This is the most valuable corpus in the collection and the smallest. It is the
only one that is both human-transcribed and spontaneous telephony -- the register
Chinar-F8 actually targets. Adding 134.7 h of it to the training mix moved WER on
human-transcribed calls from 43.39% to 35.08%, the largest… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-calls-human-140h.az-asr-human-430h
Azerbaijani read speech (LocalDoc + FLEURS)
Azerbaijani read speech assembled from LocalDoc/azerbaijani_asr and
LocalDoc/fleurs-azerbaijani-asr. 371,515 clips, 429.85 h, 16 kHz.
Clean, accurately labelled, and the wrong register for telephony. It is
literature and schoolbooks read aloud, plus FLEURS sentences. Useful for
vocabulary and general acoustics; it will not teach a model what a spontaneous
phone conversation sounds like.
A related experiment is worth knowing about… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-human-430h.az-asr-voa-305h
Labelling
field
value
label_origin
script
speech_register
broadcast
channel
wideband-16k
provenance
inferred
Editorial broadcast text aligned to VOA audio. Terminal punctuation at 77.4% (against 98.1% for LocalDoc) is consistent with segments cut from continuous broadcast rather than at sentence ends.
Adding this to a call model made it worse. Chinar-F8 v4 mixed in 147.8 h of it and strict WER on human-transcribed calls moved 44.23% -> 56.69%. Register, not… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-voa-305h.moore_audio_data~70h of Mooré paired audio+text data from jw.org and jwsoup
Headers: audio, text
az-asr-youtube-195h
Labelling
field
value
label_origin
asr:google
speech_register
spontaneous
channel
wideband-16k
provenance
documented
Google ASR output over YouTube audio. Largely a fuller repack of az-asr-youtube-136h from the same nine channels, so the two overlap heavily and should not be summed.
label_origin distinguishes text that existed before the audio (script,
exact by construction) from text written by a listener (human-transcript, high
but edited) from machine… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-youtube-195h.
