CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openslr /librispeech_asr Dataset Card for librispeech_asr Dataset Summary LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. Supported Tasks and Leaderboards automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/openslr/librispeech_asr.audioautomatic-speech-recognition100K<n<1M245 likes51k downloads1y agoHugging Face02openslr /openslrOpenSLR is a site devoted to hosting speech and language resources, such as training corpora for speech recognition, and software related to speech recognition. We intend to be a convenient place for anyone to put resources that they have created, so that they can be downloaded publicly.automatic-speech-recognition1K<n<10K31 likes519 downloads2y agoHugging Face03Sadique5 /openslr_quranic_asraudio10K<n<100K3 likes516 downloads2y agoHugging Face04voice-biomarkers /openslr-140-hq-Kazakh Kazakh Speech Dataset (KSD) Identifier: SLR140 Source: https://www.openslr.org/140/ Summary: High-quality open source Kazakh speech corpus developed by the Department of Artificial Intelligence and Big Data of Al-Farabi Kazakh National University (554 hours) Category: Speech License: Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0 US) About this resource: High-quality open source Kazakh speech corpus. The corpus contains about 554 hours of transcribed audio recordings… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-140-hq-Kazakh.audio100K<n<1M4 likes330 downloads2y agoHugging Face05thantzinphyo /burmese-speech-refined-openslr-80 Burmese Speech Refined OpenSLR-80 Summary This dataset is a speech dataset developed based on the original OpenSLR Dataset (SLR80), with the text and audio data carefully reviewed and further refined for Burmese language applications. In the original OpenSLR Dataset, the Burmese text was transcribed based on how the words were pronounced in the corresponding audio recordings. In this dataset, the original audio and text data were used as a reference, and the text… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/burmese-speech-refined-openslr-80.audioautomatic-speech-recognition1K<n<10K2 likes313 downloads22d agoHugging Face06edifier99 /sinhala-openslr-111haudio100K<n<1M0 likes248 downloads4mo agoHugging Face07Kimang18 /asr-khm-ddd-fleurs-openslr210K<n<100K0 likes212 downloads5mo agoHugging Face08SPEAK-ASR /openslr-sinhala-asr-normaudio10K<n<100K0 likes201 downloads7mo agoHugging Face09SPEAK-ASR /openslr-sinhala-asraudio100K<n<1M2 likes198 downloads7mo agoHugging Face10JKA-NLP /OpenSLR-126 Dataset Card for "OpenSLR-126" More Information needed audio100K<n<1M0 likes177 downloads3y agoHugging Face11openslr /librispeech_lmLanguage modeling resources to be used in conjunction with the LibriSpeech ASR corpus.text-generation10M<n<100M2 likes171 downloads3y agoHugging Face12SPEAK-ASR /openslr-sinhala-asr-depricated-versionaudio100K<n<1M0 likes160 downloads9mo agoHugging Face13KrorngAI /openslr-librispeech_asr-clean100-00000-19999-km-translateaudio10K<n<100K0 likes133 downloads2mo agoHugging Face14Pratik /Gujarati_OpenSLROpenSLR is a site devoted to hosting speech and language resources, such as training corpora for speech recognition, and software related to speech recognition. They intend to be a convenient place for anyone to put resources that they have created, so that they can be downloaded publicly. They aim to provide a central, hassle-free place for others to put their speech resources. see there http://www.openslr.org/contributions.html #Supported Task Automatic Speech Recognition #Languages… See the full description on the dataset page: https://huggingface.co/datasets/Pratik/Gujarati_OpenSLR.1 likes127 downloads5y agoHugging Face15chuuhtetnaing /myanmar-speech-dataset-openslr-80Please visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Speech Dataset (OpenSLR-80) This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual OpenSLR dataset. For the complete multilingual dataset and additional information, please visit the original dataset repository of OpenSLR HuggingFace page. Original Source OpenSLR is a site devoted to hosting speech and language resources, such as training… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-openslr-80.audiotext-to-speech1K<n<10K7 likes121 downloads2y agoHugging Face16deepdml /openslr65-tamil OpenSLR-65 – Tamil Transcribed Speech Source: https://www.openslr.org/65/ This dataset contains transcribed high-quality audio of Tamil sentences recorded by volunteers. It is part of the OpenSLR collection of free speech resources for low-resource languages. The data was collected via the Appen (formerly Figure Eight / CrowdFlower) crowdsourcing platform and is intended for use in training automatic speech recognition (ASR) and text-to-speech (TTS) systems. Data… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr65-tamil.audioautomatic-speech-recognition1K<n<10K0 likes121 downloads7mo agoHugging Face17spktsagar /openslr-nepali-asr-cleanedThis data set contains transcribed audio data for Nepali. The data set consists of flac files, and a TSV file. The file utt_spk_text.tsv contains a FileID, anonymized UserID and the transcription of audio in the file. The data set has been manually quality checked, but there might still be errors. The audio files are sampled at rate of 16KHz, and leading and trailing silences are trimmed using torchaudio's voice activity detection.1 likes114 downloads2y agoHugging Face18voice-biomarkers /openslr-147-hq-Nahuatl Veracruz Orizaba Nahuatl Endangered Language Identifier: SLR147 Summary: Audio corpus of Orizaba (Veracruz) Nahuatl speech (Glottocode: oriz1235; ISO 639-3: nlv) Category: Speech License: Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) About this resource: The substantive material of this deposit was gathered over a 13-month period from February 2022 to March 2023. It comprised 657 files totaling approximately 119 hours, 26 minutes, 59 seconds of material. All but 81… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-147-hq-Nahuatl.audioautomatic-speech-recognitionn<1K1 likes114 downloads2y agoHugging Face19novalalthoff /su-openslr su-openslr (Sundanese ASR subset) This dataset bundles Sundanese audio with a metadata.csv. Files audio/ — WAV files (su_id_female/, su_id_male/) metadata.csv — columns like path, sentence, etc. Usage from datasets import load_dataset, Audio ds = load_dataset("novalalthoff/su-openslr", data_files="metadata.csv", split="train") ds = ds.cast_column("path", Audio(sampling_rate=16000)) 0 likes105 downloads11mo agoHugging Face20SPEAK-ASR /openslr-sinhala-asr-preprocessed100K<n<1M1 likes102 downloads7mo agoHugging Face21ocisd4 /openslr_asraudio100K<n<1M0 likes100 downloads2y agoHugging Face22voice-biomarkers /openslr-32-hq-SA-languages-Afrikaans High quality TTS data for four South African languages - Afrikaans Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Afrikaans License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.audioautomatic-speech-recognition1K<n<10K5 likes90 downloads2y agoHugging Face23projecte-aina /openslr-slr69-ca-trimmed-denoised Dataset Card for openslr-slr69-ca-denoised This is a post-processed version of the Catalan subset belonging to the Open Speech and Language Resources (OpenSLR) speech dataset. Specifically the subset OpenSLR-69. The original HF🤗 SLR-69 dataset is located here. Same license is maintained: Attribution-ShareAlike 4.0 International. Dataset Details Dataset Description We processed the data of the Catalan OpenSLR with the following recipe: Trimming: Long… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/openslr-slr69-ca-trimmed-denoised.audiotext-to-speech1K<n<10K0 likes89 downloads3y agoHugging Face24JKA-NLP /OpenSLR-79 Dataset Card for "OpenSLR-79" More Information needed audio1K<n<10K0 likes82 downloads3y agoHugging Face25KrorngAI /fleurs-km-kh-openslr-SLR42 Combined openslr/openslr-SLR42 and google/fleurs-km-kh. Each audio is shorter than or equal to 30 seconds. Sampling rate is 16,000 Each transcription is already normalized by tha.normalize.processor (pip install tha) audio1K<n<10K0 likes80 downloads7mo agoHugging Face26xezpeleta /openslr76 Dataset Card for "openslr76" More Information needed audio1K<n<10K0 likes79 downloads1y agoHugging Face27nolimitsxl /open_slr_lang_resourceaudio1K<n<10K0 likes78 downloads8d agoHugging Face28KrorngAI /fleurs_openslr42_mpwtNOTE: If your colab crashes, please use pip install --upgrade --quiet datasets[audio]==3.6.0 to install datasets[audio] version 3.6.0. This dataset combined google/fleurs, openslr/openslr42, and cleaned seanghay/khmer_mpwt_speech. Severals processes are executed: clean up seanghay/khmer_mpwt_speech: manually correct wrong transcriptions over 2058 rows normalize transcription: remove invisible white space; process ៗ, numbers, currencies, date into khmer text; and separate each word by space… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/fleurs_openslr42_mpwt.audioautomatic-speech-recognition1K<n<10K1 likes70 downloads11mo agoHugging Face29iamTangsang /OpenSLR54-Nepali-ASRaudio100K<n<1M0 likes66 downloads2y agoHugging Face30phonsobon /openslr42-khmer-malegated OpenSLR SLR42 Khmer Male Speech This dataset is a processed version of the OpenSLR SLR42 Khmer speech dataset. Dataset Description This dataset contains approximately 2,906 Khmer speech recordings with corresponding Khmer transcriptions. Each example contains: audio: Khmer speech recording text: Khmer transcription Dataset Structure Column Type Description audio Audio Khmer speech recording text String Khmer transcription… See the full description on the dataset page: https://huggingface.co/datasets/phonsobon/openslr42-khmer-male.audioautomatic-speech-recognition1K<n<10K0 likes61 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.