CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openslr /librispeech_asr Dataset Card for librispeech_asr Dataset Summary LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. Supported Tasks and Leaderboards automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/openslr/librispeech_asr.audioautomatic-speech-recognition100K<n<1M245 likes51k downloads1y agoHugging Face02thantzinphyo /burmese-speech-refined-openslr-80 Burmese Speech Refined OpenSLR-80 Summary This dataset is a speech dataset developed based on the original OpenSLR Dataset (SLR80), with the text and audio data carefully reviewed and further refined for Burmese language applications. In the original OpenSLR Dataset, the Burmese text was transcribed based on how the words were pronounced in the corresponding audio recordings. In this dataset, the original audio and text data were used as a reference, and the text… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/burmese-speech-refined-openslr-80.audioautomatic-speech-recognition1K<n<10K2 likes313 downloads22d agoHugging Face03chuuhtetnaing /myanmar-speech-dataset-openslr-80Please visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Speech Dataset (OpenSLR-80) This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual OpenSLR dataset. For the complete multilingual dataset and additional information, please visit the original dataset repository of OpenSLR HuggingFace page. Original Source OpenSLR is a site devoted to hosting speech and language resources, such as training… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-openslr-80.audiotext-to-speech1K<n<10K7 likes121 downloads2y agoHugging Face04deepdml /openslr65-tamil OpenSLR-65 – Tamil Transcribed Speech Source: https://www.openslr.org/65/ This dataset contains transcribed high-quality audio of Tamil sentences recorded by volunteers. It is part of the OpenSLR collection of free speech resources for low-resource languages. The data was collected via the Appen (formerly Figure Eight / CrowdFlower) crowdsourcing platform and is intended for use in training automatic speech recognition (ASR) and text-to-speech (TTS) systems. Data… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr65-tamil.audioautomatic-speech-recognition1K<n<10K0 likes121 downloads7mo agoHugging Face05voice-biomarkers /openslr-147-hq-Nahuatl Veracruz Orizaba Nahuatl Endangered Language Identifier: SLR147 Summary: Audio corpus of Orizaba (Veracruz) Nahuatl speech (Glottocode: oriz1235; ISO 639-3: nlv) Category: Speech License: Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) About this resource: The substantive material of this deposit was gathered over a 13-month period from February 2022 to March 2023. It comprised 657 files totaling approximately 119 hours, 26 minutes, 59 seconds of material. All but 81… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-147-hq-Nahuatl.audioautomatic-speech-recognitionn<1K1 likes114 downloads2y agoHugging Face06voice-biomarkers /openslr-32-hq-SA-languages-Afrikaans High quality TTS data for four South African languages - Afrikaans Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Afrikaans License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.audioautomatic-speech-recognition1K<n<10K5 likes90 downloads2y agoHugging Face07KrorngAI /fleurs_openslr42_mpwtNOTE: If your colab crashes, please use pip install --upgrade --quiet datasets[audio]==3.6.0 to install datasets[audio] version 3.6.0. This dataset combined google/fleurs, openslr/openslr42, and cleaned seanghay/khmer_mpwt_speech. Severals processes are executed: clean up seanghay/khmer_mpwt_speech: manually correct wrong transcriptions over 2058 rows normalize transcription: remove invisible white space; process ៗ, numbers, currencies, date into khmer text; and separate each word by space… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/fleurs_openslr42_mpwt.audioautomatic-speech-recognition1K<n<10K1 likes70 downloads11mo agoHugging Face08phonsobon /openslr42-khmer-malegated OpenSLR SLR42 Khmer Male Speech This dataset is a processed version of the OpenSLR SLR42 Khmer speech dataset. Dataset Description This dataset contains approximately 2,906 Khmer speech recordings with corresponding Khmer transcriptions. Each example contains: audio: Khmer speech recording text: Khmer transcription Dataset Structure Column Type Description audio Audio Khmer speech recording text String Khmer transcription… See the full description on the dataset page: https://huggingface.co/datasets/phonsobon/openslr42-khmer-male.audioautomatic-speech-recognition1K<n<10K0 likes61 downloads1mo agoHugging Face09deepdml /openslr80-burmese OpenSLR-80 – Brumese Transcribed Speech Source: https://www.openslr.org/80/ This dataset contains transcribed high-quality audio of Burmese sentences recorded by female volunteers. It is part of the OpenSLR collection of free speech resources for low-resource languages. The data was collected via the Appen (formerly Figure Eight / CrowdFlower) crowdsourcing platform and is intended for use in training automatic speech recognition (ASR) and text-to-speech (TTS) systems.… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr80-burmese.audioautomatic-speech-recognition1K<n<10K0 likes57 downloads7mo agoHugging Face101rsh /gujarati-f-openslr Gujarati OpenSLR Female Interspeech data downloaded from https://www.openslr.org/resources/78/gu_in_female.zip Dataset Details Gujarati Data (Most of the entries are <30 seconds and hence Whisper Models can be used for accurate timestamp prediction) Also, the audio seems to have been spoken by a single female. audioautomatic-speech-recognition1K<n<10K1 likes56 downloads2y agoHugging Face11deepdml /openslr-32-hq-SA-languages SLR32 – High Quality TTS Data for Four South African Languages Identifier: SLR32License: CC BY-SA 4.0Source: https://www.openslr.org/32/ This dataset contains multi-speaker high quality transcribed audio data for four languages of South Africa: Afrikaans (af_za), Sesotho (st_za), Setswana (tn_za) and isiXhosa (xh_za). The dataset consists of WAV files and a TSV file transcribing the audio. In each folder the file line_index.tsv contains a FileID (which in turn encodes the UserID)… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr-32-hq-SA-languages.audioautomatic-speech-recognition1K<n<10K0 likes56 downloads7mo agoHugging Face12Kukedlc /openslr61-es-ar-full openslr61-es-ar-full OpenSLR 61 (Crowdsourced high-quality Argentinian Spanish) consolidado COMPLETO con linaje. Incluye male + female + weather messages argentinos. Linaje (trazabilidad por sample) source: siempre "openslr61" subset: "main" (frases generales) o "weather" (mensajes de clima) gender: "m" / "f" speaker_id: ID anonimizado original del speaker file_id: ID original del archivo OpenSLR license: CC-BY-SA-4.0 Schema campo tipo… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/openslr61-es-ar-full.audiotext-to-speech1K<n<10K0 likes52 downloads4mo agoHugging Face13JeevanDai /OpenSLR54-Nepali-ASR-parquet OpenSLR 54: Large Nepali ASR training data set (unmodified parquet repackaging) This is an unofficial repackaging of the official OpenSLR 54 release (SLR54, https://www.openslr.org/54/), converted to parquet so it can be streamed with 🤗 datasets. It is not affiliated with or endorsed by OpenSLR or the original authors. All credit for the data belongs to the original creators (see Citation). What's inside 157,905 utterances, 16 shards: one per original zip… See the full description on the dataset page: https://huggingface.co/datasets/JeevanDai/OpenSLR54-Nepali-ASR-parquet.audioautomatic-speech-recognition100K<n<1M0 likes49 downloads2d agoHugging Face14voice-biomarkers /openslr-32-hq-SA-languages-Setswana High quality TTS data for four South African languages - Setswana Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Setswana License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Setswana.audioautomatic-speech-recognition1K<n<10K1 likes45 downloads2y agoHugging Face15vrclc /openslr63 SLR63: Crowdsourced high-quality Malayalam multi-speaker speech data set This data set contains transcribed high-quality audio of Malayalam sentences recorded by volunteers. The data set consists of wave files, and a TSV file (line_index.tsv). The file line_index.tsv contains a anonymized FileID and the transcription of audio in the file. The data set has been manually quality checked, but there might still be errors. Please report any issues in the following issue tracker on… See the full description on the dataset page: https://huggingface.co/datasets/vrclc/openslr63.audioautomatic-speech-recognition1K<n<10K3 likes42 downloads3y agoHugging Face16voice-biomarkers /openslr-32-hq-SA-languages-Sesotho High quality TTS data for four South African languages - Sesotho Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Sesotho License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Sesotho.audioautomatic-speech-recognition1K<n<10K2 likes36 downloads2y agoHugging Face17deepdml /openslr42-khmer-tts OpenSLR42 – High Quality TTS Data for Khmer This dataset contains high-quality transcribed audio data for Khmer (km-KH).It is the HuggingFace mirror of OpenSLR Resource #42. Identifier: SLR42 Summary: Multi-speaker TTS data for Khmer Category: Speech License: CC BY-SA 4.0 Original source: https://www.openslr.org/42/ Collected by: Google Copyright: 2016, 2017, 2018 Google LLC What's inside Field Description filename Original filename (without… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr42-khmer-tts.audioautomatic-speech-recognition1K<n<10K0 likes35 downloads7mo agoHugging Face18Max5ive /openslr-32-hq-SA-languages-Sesotho High quality TTS data for four South African languages - Sesotho Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Sesotho License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/openslr-32-hq-SA-languages-Sesotho.audioautomatic-speech-recognition1K<n<10K0 likes35 downloads7mo agoHugging Face19djsamseng /openslr-khmer-tts-asr OpenSLR - Khmer TTS / ASR Dataset import datasets ds = datasets.load_dataset("djsamseng/openslr-khmer-tts-asr") ds["train"][1018] { 'audio': { array': array([3.05175781e-05, 3.05175781e-05, 3.05175781e-05, ..., 6.10351562e-05, 0.00000000e+00, 0.00000000e+00]), 'sampling_rate': 48000 }, 'english': 'Current news in the country', 'khmer': 'ព័ត៌មាន ទាន់ ហេតុការណ៍ ក្នុង ប្រទេស', 'transliteration': 'poatemean toan hetokar knong protes', 'speaker': '3154'… See the full description on the dataset page: https://huggingface.co/datasets/djsamseng/openslr-khmer-tts-asr.audiotext-to-speech1K<n<10K0 likes35 downloads6mo agoHugging Face20daniel-dona /openslr-slr67See: https://www.openslr.org/67/ audioautomatic-speech-recognition10K<n<100K0 likes26 downloads2y agoHugging Face21voice-biomarkers /openslr-32-hq-SA-languages-isiXhosa High quality TTS data for four South African languages - isiXhosa Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - isiXhosa License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-isiXhosa.audioautomatic-speech-recognition1K<n<10K1 likes23 downloads2y agoHugging Face22oademuwagun /tts-yor-openslr Yoruba Speech Dataset (OpenSLR 86 - Female Speakers) A high-quality Yoruba speech dataset from female speakers, sourced from OpenSLR 86. Usage from datasets import load_dataset dataset = load_dataset("oademuwagun/yoruba-speech-female") Dataset Structure ├── metadata.csv └── wavs/ ├── yof_06136_00616155826.wav ├── yof_04310_00085429033.wav └── ... Metadata Format Column Description file_name Path to audio file transcription… See the full description on the dataset page: https://huggingface.co/datasets/oademuwagun/tts-yor-openslr.audiotext-to-speech1K<n<10K0 likes22 downloads7mo agoHugging Face23JeevanDai /OpenSLR43-Nepali-TTS-parquet OpenSLR 43: High quality TTS data for Nepali (unmodified parquet repackaging) This is an unofficial repackaging of the official OpenSLR 43 release (SLR43, https://www.openslr.org/43/), converted to parquet for 🤗 datasets. It is not affiliated with or endorsed by OpenSLR, Google or the original authors. All credit for the data belongs to the original creators (see Citation). What's inside 2,064 utterances, 18 female speakers, ~2.8 h, 48 kHz mono; recorded in… See the full description on the dataset page: https://huggingface.co/datasets/JeevanDai/OpenSLR43-Nepali-TTS-parquet.audioautomatic-speech-recognition1K<n<10K0 likes16 downloads1d agoHugging Face24Aananda-giri /openSLR-Nepali OpenSLR Nepali Speech Dataset (Preprocessed) Dataset Description This is a preprocessed version of the Nepali speech dataset from OpenSLR, ready for training speech models including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). Dataset Statistics Total Audio Files: 118,231 Total Duration: 57.34 hours Sample Rate: 16kHz Channels: Mono Format: WAV Preprocessing Applied Text Preprocessing: Text cleaning and… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/openSLR-Nepali.audioautomatic-speech-recognition10K<n<100K0 likes12 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.