CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ciempiess /ciempiess_light Dataset Card for ciempiess_light Dataset Summary The CIEMPIESS LIGHT is a Radio Corpus designed to create acoustic models for automatic speech recognition and it is made up by recordings of spontaneous conversations in Mexican Spanish between a radio moderator and his guests. It is an enhanced version of the CIEMPIESS Corpus (LDC item LDC2015S07). CIEMPIESS LIGHT is "light" because it doesn't include much of the files of the first version of CIEMPIESS and it is "enhanced"… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_light.audioautomatic-speech-recognition10K<n<100K2 likes207 downloads2y agoHugging Face02ciempiess /librivox_spanish Dataset Card for librivox_spanish Dataset Summary Librivox is a non-commercial, non-profit and ad-free project that is dedicated to make all books in the public domain available, for free, in audio format on the internet. According to this, we downloaded 300 titles in Spanish to create the LIBRIVOX SPANISH CORPUS. The LIBRIVOX SPANISH CORPUS has a duration of 73 hours and it is constituted by audio files between 3 and 10 seconds long, manually segmented. Transcription are… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/librivox_spanish.audioautomatic-speech-recognition10K<n<100K5 likes187 downloads2y agoHugging Face03ciempiess /voxforge_spanish Dataset Card for voxforge_spanish Dataset Summary VoxForge was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac). They promise they will make available all submitted audio files under the GPL license, and then 'compile' them into acoustic models for use with Open Source speech recognition engines such as CMU Sphinx, ISIP, Julius and HTK. According to this, we downloaded the Spanish recordings of… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/voxforge_spanish.audioautomatic-speech-recognition10K<n<100K4 likes122 downloads2y agoHugging Face04ciempiess /ciempiess_testThe CIEMPIESS TEST Corpus is a gender balanced corpus destined to test acoustic models for the speech recognition task. The corpus was manually transcribed and it contains audio recordings from 10 male and 10 female speakers. The CIEMPIESS TEST is one of the three corpora included at the LDC's \"CIEMPIESS Experimentation\" (LDC2019S07).audioautomatic-speech-recognition1K<n<10K3 likes116 downloads3y agoHugging Face05ciempiess /wikipedia_spanish Dataset Card for wikipedia_spanish Dataset Summary According to the project page of the WikiProject Spoken Wikipedia: The WikiProject Spoken Wikipedia aims to produce recordings of Wikipedia articles being read aloud. Therefore, the WIKIPEDIA SPANISH CORPUS is a dataset created from the Spanish version of the WikiProject Spoken Wikipedia, called Wikipedia Grabada The WIKIPEDIA SPANISH CORPUS aims to be used in the Automatic Speech Recognition (ASR) task. It is a gender… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/wikipedia_spanish.audioautomatic-speech-recognition10K<n<100K1 likes110 downloads2y agoHugging Face06ciempiess /tele_con_ciencia Dataset Card for tele_con_ciencia Dataset Summary According to the Facebook page of Tele con Ciencia: "Nuestra misión es la comunicación pública de la ciencia y la tecnología mexicana. El objetivo, la participación activa de todos los mexicanos en las áreas del descubrimiento científico y el desarrollo tecnológico." "Our mission is to spread the achievements of the Mexican Science and Technology. The main goal is to promote the active participation of mexican people in… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/tele_con_ciencia.audioautomatic-speech-recognition10K<n<100K0 likes100 downloads2y agoHugging Face07ciempiess /tedx_spanish Dataset Card for tedx_spanish Dataset Summary According to the TEDx website: In the spirit of ideas worth spreading, TEDx is a program of local, self-organized events that bring people together to share a TED-like experience. At a TEDx event, TEDTalks video and live speakers combine to spark deep discussion and connection in a small group. These local, self-organized events are branded TEDx, where x = independently organized TED event. The TED Conference provides… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/tedx_spanish.audioautomatic-speech-recognition10K<n<100K1 likes96 downloads2y agoHugging Face08ciempiess /ciempiess_complementary Dataset Card for ciempiess_complementary Dataset Summary The CIEMPIESS COMPLEMENTARY is a phonetically balanced corpus of isolated Spanish words spoken by people of Central Mexico. It was designed to solve one particular issue when training automatic speech recognition (ASR) systems in the Spanish of Central Mexico. This problem appears when someone collects some training data, but the system complains because it does not find enough instances of one or more particular… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_complementary.audioautomatic-speech-recognitionn<1K1 likes88 downloads2y agoHugging Face09i4ds /swiss-german-city-sentences_v2 Swiss German City Sentences v2 Synthetic Swiss German speech dataset with city name sentences across multiple dialects. audioautomatic-speech-recognition10K<n<100K1 likes76 downloads4mo agoHugging Face10ciempiess /ciempiess_balance Dataset Card for ciempiess_balance Dataset Summary The CIEMPIESS BALANCE Corpus is designed to match with the CIEMPIESS LIGHT Corpus (LDC2017S23). So, "Balance" means that if the CIEMPIESS BALANCE is combined with the CIEMPIESS LIGHT, one will get a gender balanced corpus. To appreciate this, one need to know that the CIEMPIESS LIGHT is by itself, a gender unbalanced corpus of approximately 25% of female speakers and 75% of male speakers. So, the CIEMPIESS BALANCE is a… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_balance.audioautomatic-speech-recognition1K<n<10K1 likes64 downloads2y agoHugging Face11ciempiess /ciempiess_fem Dataset Card for ciempiess_fem Dataset Summary Since the publication of the CIEMPIESS Corpus (LDC2015S07) in 2015 we have noticed that there is a lack of female speakers in the sources where we traditionally take audio to create new CIEMPIESS datasets. That is why we decided to create a corpus that helps to balance future gender unbalanced datasets. The CIEMPIESS FEM Corpus was created by recordings and human transcripts of 21 different women. 16 of these women are… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_fem.audioautomatic-speech-recognition1K<n<10K1 likes63 downloads2y agoHugging Face12YunisAFS /ciempiess_light Dataset Card for ciempiess_light Dataset Summary The CIEMPIESS LIGHT is a Radio Corpus designed to create acoustic models for automatic speech recognition and it is made up by recordings of spontaneous conversations in Mexican Spanish between a radio moderator and his guests. It is an enhanced version of the CIEMPIESS Corpus (LDC item LDC2015S07). CIEMPIESS LIGHT is "light" because it doesn't include much of the files of the first version of CIEMPIESS and it is "enhanced"… See the full description on the dataset page: https://huggingface.co/datasets/YunisAFS/ciempiess_light.audioautomatic-speech-recognition10K<n<100K0 likes42 downloads7mo agoHugging Face13cillegio /az-asr-youtube-136hgated Labelling field value label_origin asr:google speech_register spontaneous channel wideband-16k provenance documented Google ASR output over YouTube audio. No human labels anywhere in it. label_origin distinguishes text that existed before the audio (script, exact by construction) from text written by a listener (human-transcript, high but edited) from machine output (asr:<vendor>, bounded by that vendor's error rate). The older label_type field is retained… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-youtube-136h.textautomatic-speech-recognition10K<n<100K0 likes17 downloads6d agoHugging Face14cillegio /az-asr-calls-190hgated Azerbaijani call-centre speech, Soniox pseudo-labels Real Azerbaijani call-centre audio transcribed by Soniox stt-async-v5. 169,273 clips, 190.63 h, 8 kHz telephony. The labels are machine output and carry a measured ceiling. Soniox scores 46.65% WER against human transcripts of this same kind of audio. A model trained on these labels learns to agree with Soniox -- including where Soniox is wrong, and including its habit of dropping words. That is not hypothetical. Scoring the… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-calls-190h.textautomatic-speech-recognition100K<n<1M0 likes15 downloads6d agoHugging Face15cillegio /az-asr-calls-human-140hgated Azerbaijani call-centre speech, human-transcribed Real Azerbaijani call-centre audio with transcripts written by people listening to it. 29,982 clips, 139.75 h, 8 kHz telephony. This is the most valuable corpus in the collection and the smallest. It is the only one that is both human-transcribed and spontaneous telephony -- the register Chinar-F8 actually targets. Adding 134.7 h of it to the training mix moved WER on human-transcribed calls from 43.39% to 35.08%, the largest… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-calls-human-140h.textautomatic-speech-recognition10K<n<100K0 likes15 downloads6d agoHugging Face16cillegio /az-asr-human-430hgated Azerbaijani read speech (LocalDoc + FLEURS) Azerbaijani read speech assembled from LocalDoc/azerbaijani_asr and LocalDoc/fleurs-azerbaijani-asr. 371,515 clips, 429.85 h, 16 kHz. Clean, accurately labelled, and the wrong register for telephony. It is literature and schoolbooks read aloud, plus FLEURS sentences. Useful for vocabulary and general acoustics; it will not teach a model what a spontaneous phone conversation sounds like. A related experiment is worth knowing about… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-human-430h.textautomatic-speech-recognition100K<n<1M0 likes13 downloads6d agoHugging Face17cillegio /az-asr-voa-305hgated Labelling field value label_origin script speech_register broadcast channel wideband-16k provenance inferred Editorial broadcast text aligned to VOA audio. Terminal punctuation at 77.4% (against 98.1% for LocalDoc) is consistent with segments cut from continuous broadcast rather than at sentence ends. Adding this to a call model made it worse. Chinar-F8 v4 mixed in 147.8 h of it and strict WER on human-transcribed calls moved 44.23% -> 56.69%. Register, not… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-voa-305h.tabularautomatic-speech-recognition100K<n<1M0 likes10 downloads6d agoHugging Face18CITADEL-BF-Center /moore_audio_datagated~70h of Mooré paired audio+text data from jw.org and jwsoup Headers: audio, text audioautomatic-speech-recognition10K<n<100K1 likes8 downloads3mo agoHugging Face19cillegio /az-asr-youtube-195hgated Labelling field value label_origin asr:google speech_register spontaneous channel wideband-16k provenance documented Google ASR output over YouTube audio. Largely a fuller repack of az-asr-youtube-136h from the same nine channels, so the two overlap heavily and should not be summed. label_origin distinguishes text that existed before the audio (script, exact by construction) from text written by a listener (human-transcript, high but edited) from machine… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-youtube-195h.textautomatic-speech-recognition10K<n<100K0 likes5 downloads6d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.