CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openslr /librispeech_asr Dataset Card for librispeech_asr Dataset Summary LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. Supported Tasks and Leaderboards automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/openslr/librispeech_asr.audioautomatic-speech-recognition100K<n<1M245 likes54k downloads1y agoHugging Face02TheAgenticDataCompany /open-yap-1k Open Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use Today we're releasing Open Yap 1K: 1,000 hours of dual-channel English conversation, capturing how people speak together naturally in real-world environments recorded in 48kHz. The dataset ships free for both commercial and research use. The sample on the Hugging Face Hub - 8.9 hours, 16 conversations, CC-BY-4.0, listenable in the dataset viewer. The full corpus - 1,000 hours, 1,602… See the full description on the dataset page: https://huggingface.co/datasets/TheAgenticDataCompany/open-yap-1k.audioaudio-to-audion<1K104 likes4.9k downloads16d agoHugging Face03multilingual-tts /open-bible OpenBibleTTS OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license. Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources Source: Open Bible (CC BY-SA) Languages Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.audiotext-to-speech1M<n<10M1 likes2.7k downloads3mo agoHugging Face04AfriSpeech /open-bible-speech-african Open Bible Resources — African Languages Spoken-audio Bible recordings aligned to verse-level text for 19 African languages — roughly 1,741 hours of audio across ~552,907 audio–text pairs (~357 GB). This dataset is the African-language subset of davidguzmanr/open-bible-resources, re-hosted here by AfriSpeech to make the African languages easy to find and use on their own. The audio and text are unchanged from the source; only the non-African configurations have been removed. All… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/open-bible-speech-african.audioautomatic-speech-recognition100K<n<1M3 likes2.5k downloads2mo agoHugging Face05espnet /ace-opencpop-segments Citation Information @misc{shi2024singingvoicedatascalingup, title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing}, author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe}, year={2024}, eprint={2401.17619}, archivePrefix={arXiv}, primaryClass={cs.SD}, url={https://arxiv.org/abs/2401.17619}, } audiotext-to-audio100K<n<1M8 likes1.1k downloads2y agoHugging Face06hf-audio /open-asr-leaderboard-multilingual-datasets ASR Leaderboard Datasets This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS). How to Load To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>. from datasets import load_dataset # Load the FLEURS dataset for Bulgarian fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg") print(fleurs_bg) # Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.audioautomatic-speech-recognition100K<n<1M4 likes1.1k downloads2mo agoHugging Face07VoiceArena /MonsoonASR-Open-ASR-leaderboard-en-IN Voice Arena Monsoon en-IN (public test) Part of the Open ASR Leaderboard, in the main board's default column set, so it contributes to the headline Average WER for every model listed. A conversational Indian English ASR test set that records who was speaking, not only what was said. Every clip carries twelve speaker attributes — gender, age, native district and state, education, occupation, income band, handset — so a difference between two systems can be traced to a group of… See the full description on the dataset page: https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-en-IN.audioautomatic-speech-recognition1K<n<10K3 likes875 downloads24d agoHugging Face08OpenFormosa /parliament Parliament Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the zh_tw split of disco-eth/WorldSpeech. It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD parliamentary proceedings. This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style text filtering pipeline. Audio is preserved from the upstream dataset and cast as a Hugging Face Audio(sampling_rate=24000) feature. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/parliament.audioautomatic-speech-recognition100K<n<1M2 likes720 downloads4mo agoHugging Face09openslr /openslrOpenSLR is a site devoted to hosting speech and language resources, such as training corpora for speech recognition, and software related to speech recognition. We intend to be a convenient place for anyone to put resources that they have created, so that they can be downloaded publicly.automatic-speech-recognition1K<n<10K31 likes520 downloads2y agoHugging Face10openbank-uz /youtube_transcriptions Dataset Description A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings. Use Cases Automatic Speech Recognition (ASR) for Uzbek Text-to-Speech (TTS) synthesis for Uzbek Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS) Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.audioautomatic-speech-recognition100K<n<1M2 likes439 downloads6mo agoHugging Face11SKNahin /open-large-bengali-asr-data Open Large Bengali ASR Data This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps). Datasets: commonvoice openslr madasr shrutilipi flerus kathbath indictts ucla gali audioautomatic-speech-recognition1M<n<10M13 likes418 downloads2y agoHugging Face12rishiraj /open-large-bengali-asr-data Open Large Bengali ASR Data This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps). Datasets: commonvoice audioautomatic-speech-recognition100K<n<1M0 likes406 downloads3mo agoHugging Face13thantzinphyo /burmese-speech-refined-openslr-80 Burmese Speech Refined OpenSLR-80 Summary This dataset is a speech dataset developed based on the original OpenSLR Dataset (SLR80), with the text and audio data carefully reviewed and further refined for Burmese language applications. In the original OpenSLR Dataset, the Burmese text was transcribed based on how the words were pronounced in the corresponding audio recordings. In this dataset, the original audio and text data were used as a reference, and the text… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/burmese-speech-refined-openslr-80.audioautomatic-speech-recognition1K<n<10K2 likes306 downloads18d agoHugging Face14OpenFormosa /common_voice_25_zh-TW Common Voice Scripted Speech 25.0 - Chinese (Taiwan) This repository mirrors the Chinese (Taiwan) (zh-TW) portion of Mozilla Common Voice Scripted Speech 25.0 in Hugging Face datasets format. The audio has been embedded into Parquet and exposed as a Hugging Face Audio feature at 48 kHz. Official source page: Mozilla Data Collective - Common Voice Scripted Speech 25.0 - Chinese (Taiwan) Dataset Details Field Value Dataset ID cmn2g7eaj01fio10769r1m96n… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/common_voice_25_zh-TW.audioautomatic-speech-recognition100K<n<1M2 likes298 downloads4mo agoHugging Face15VoiceArena /MonsoonASR-Open-ASR-leaderboard-hi-IN Voice Arena Monsoon hi (public test) Part of the Open ASR Leaderboard, on the Multilingual tab, where a model is ranked only if it supports every selected language. A conversational Hindi ASR test set that records who was speaking, not only what was said. Every clip carries twelve speaker attributes — gender, age, native district and state, education, occupation, income band, handset — so a difference between two systems can be traced to a group of speakers instead of… See the full description on the dataset page: https://huggingface.co/datasets/VoiceArena/MonsoonASR-Open-ASR-leaderboard-hi-IN.audioautomatic-speech-recognitionn<1K2 likes297 downloads24d agoHugging Face16mrfakename /open-yap-1k Open Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use Today we're releasing Open Yap 1K: 1,000 hours of dual-channel English conversation, capturing how people speak together naturally in real-world environments recorded in 48kHz. The dataset ships free for both commercial and research use. The sample on the Hugging Face Hub - 8.9 hours, 16 conversations, CC-BY-4.0, listenable in the dataset viewer. The full corpus - 1,000 hours, 1,602… See the full description on the dataset page: https://huggingface.co/datasets/mrfakename/open-yap-1k.audioaudio-to-audion<1K3 likes253 downloads13d agoHugging Face17lab260 /openstt_balalaika OpenSTT Annotated by Balalaika [!IMPORTANT] Official dataset for our INTERSPEECH 2026 paper "A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika. If you use this resource, please cite it. A curated Russian speech dataset for advanced speech generative tasks. Overview OpenSTT… See the full description on the dataset page: https://huggingface.co/datasets/lab260/openstt_balalaika.tabulartext-to-speech100K<n<1M3 likes226 downloads3mo agoHugging Face18OpenVoiceOS /stt-sampler-v1 stt-sampler-v1 Licensing: clips inherit their source dataset's license — CC-BY-4.0 for MInDS-14 and FLEURS clips, CC-BY-NC-SA-4.0 for Speech-MASSIVE clips (source_dataset column identifies each clip's origin). A small, balanced, representative multilingual ASR eval sampler for the OVOS Plugin Arena: 100 clips per language x 20 locales = 2000 clips, 16 kHz mono float32, one config per locale (load_dataset("OpenVoiceOS/stt-sampler-v1", "<lang>")). Designed to seed every STT… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/stt-sampler-v1.audioautomatic-speech-recognition1K<n<10K0 likes209 downloads1mo agoHugging Face19Edge0 /ark-asr-3b-open-asr-leaderboard-results ARK-ASR-3B Open ASR Leaderboard Results Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English short-form hf-audio/open-asr-leaderboard splits. These manifests were generated on a local 8x RTX 4090 machine and scored with the shared Open ASR Leaderboard scorer: PYTHONPATH=. python - <<'PY' from normalizer.eval_utils import score_results score_results( 'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official', 'AutoArk-AI/ARK-ASR-3B', ) PY Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.tabularautomatic-speech-recognition10K<n<100K12 likes195 downloads3mo agoHugging Face20baryonlabs /open-ko-s2s-eval-artifacts Open Ko-S2S 평가 산출물 (감사용) ⚠️ KsponSpeech 참조 전사는 해시로 대체돼 있습니다 KsponSpeech 는 AI Hub 배포 데이터로 재배포 제한이 있을 수 있어, kspon 런의 ref 컬럼을 ref_sha256 으로 대체했습니다(전사 원문 미포함). 모델 출력(hyp)과 채점 결과(cer_err/cer_len/cer)는 우리 산출물이라 그대로 공개합니다. Zeroth 런은 원본이 CC BY 4.0(OpenSLR #40)이라 ref 원문을 그대로 담고 있습니다. 라이선스 보유자의 검증 절차 AI Hub 에서 KsponSpeech 를 정당하게 받은 분은 다음으로 우리 수치를 검증할 수 있습니다. 리더보드 저장소의 eval/datasets_ko.py 에서 clean_kspon() 을 가져옵니다. 자기 사본의 원 전사에 clean_kspon() 을 적용합니다. 결과가 목록이면… See the full description on the dataset page: https://huggingface.co/datasets/baryonlabs/open-ko-s2s-eval-artifacts.tabularautomatic-speech-recognitionn<1K0 likes190 downloads1mo agoHugging Face21chuuhtetnaing /myanmar-speech-dataset-openslr-80Please visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Speech Dataset (OpenSLR-80) This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual OpenSLR dataset. For the complete multilingual dataset and additional information, please visit the original dataset repository of OpenSLR HuggingFace page. Original Source OpenSLR is a site devoted to hosting speech and language resources, such as training… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-openslr-80.audiotext-to-speech1K<n<10K7 likes118 downloads1y agoHugging Face22deepdml /openslr65-tamil OpenSLR-65 – Tamil Transcribed Speech Source: https://www.openslr.org/65/ This dataset contains transcribed high-quality audio of Tamil sentences recorded by volunteers. It is part of the OpenSLR collection of free speech resources for low-resource languages. The data was collected via the Appen (formerly Figure Eight / CrowdFlower) crowdsourcing platform and is intended for use in training automatic speech recognition (ASR) and text-to-speech (TTS) systems. Data… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr65-tamil.audioautomatic-speech-recognition1K<n<10K0 likes111 downloads7mo agoHugging Face23voice-biomarkers /openslr-147-hq-Nahuatl Veracruz Orizaba Nahuatl Endangered Language Identifier: SLR147 Summary: Audio corpus of Orizaba (Veracruz) Nahuatl speech (Glottocode: oriz1235; ISO 639-3: nlv) Category: Speech License: Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) About this resource: The substantive material of this deposit was gathered over a 13-month period from February 2022 to March 2023. It comprised 657 files totaling approximately 119 hours, 26 minutes, 59 seconds of material. All but 81… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-147-hq-Nahuatl.audioautomatic-speech-recognitionn<1K1 likes110 downloads2y agoHugging Face24KathleenKunLiu /open-yap-1k Open Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use Today we're releasing Open Yap 1K: 1,000 hours of dual-channel English conversation, capturing how people speak together naturally in real-world environments recorded in 48kHz. The dataset ships free for both commercial and research use. The sample on the Hugging Face Hub - 8.9 hours, 16 conversations, CC-BY-4.0, listenable in the dataset viewer. The full corpus - 1,000 hours, 1,602… See the full description on the dataset page: https://huggingface.co/datasets/KathleenKunLiu/open-yap-1k.audioaudio-to-audion<1K2 likes105 downloads14d agoHugging Face25tsdocode /open-vi-dialog-synthetic-100h OpenDialog Vietnamese Synthetic Dialogue 100h Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments. 12,000 chunks 30 seconds per chunk 100.0 hours total Each item contains S1/S2 speaker labels, turn timings, target text, relationship, pronouns, environment, topic, mood, and source reference IDs. Audio renderer: vLLM-Omni VoxCPM2 Audio format: mono WAV, 48 kHz, 30 seconds per chunk This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.audiotext-to-speech10K<n<100K0 likes103 downloads1mo agoHugging Face26phonsobon /openslr42-khmer-malegated OpenSLR SLR42 Khmer Male Speech This dataset is a processed version of the OpenSLR SLR42 Khmer speech dataset. Dataset Description This dataset contains approximately 2,906 Khmer speech recordings with corresponding Khmer transcriptions. Each example contains: audio: Khmer speech recording text: Khmer transcription Dataset Structure Column Type Description audio Audio Khmer speech recording text String Khmer transcription… See the full description on the dataset page: https://huggingface.co/datasets/phonsobon/openslr42-khmer-male.audioautomatic-speech-recognition1K<n<10K0 likes94 downloads29d agoHugging Face27voice-biomarkers /openslr-32-hq-SA-languages-Afrikaans High quality TTS data for four South African languages - Afrikaans Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Afrikaans License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.audioautomatic-speech-recognition1K<n<10K5 likes89 downloads2y agoHugging Face28KrorngAI /fleurs_openslr42_mpwtNOTE: If your colab crashes, please use pip install --upgrade --quiet datasets[audio]==3.6.0 to install datasets[audio] version 3.6.0. This dataset combined google/fleurs, openslr/openslr42, and cleaned seanghay/khmer_mpwt_speech. Severals processes are executed: clean up seanghay/khmer_mpwt_speech: manually correct wrong transcriptions over 2058 rows normalize transcription: remove invisible white space; process ៗ, numbers, currencies, date into khmer text; and separate each word by space… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/fleurs_openslr42_mpwt.audioautomatic-speech-recognition1K<n<10K1 likes88 downloads11mo agoHugging Face29OpenT2S /DuplexOmni-Data DuplexOmni Data This dataset accompanies DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction. GitHub: MuyeHuang/DuplexOmni arXiv: 2606.09186 Files The root directory includes the Writer-Director source files: inbound.director.jsonl outbound.director.jsonl The synthesized parquet training shards are about 9 TB in total, so uploading them is expected to be very slow. They will be uploaded incrementally; each completed… See the full description on the dataset page: https://huggingface.co/datasets/OpenT2S/DuplexOmni-Data.textany-to-any10K<n<100K0 likes68 downloads2mo agoHugging Face301rsh /gujarati-f-openslr Gujarati OpenSLR Female Interspeech data downloaded from https://www.openslr.org/resources/78/gu_in_female.zip Dataset Details Gujarati Data (Most of the entries are <30 seconds and hence Whisper Models can be used for accurate timestamp prediction) Also, the audio seems to have been spoken by a single female. audioautomatic-speech-recognition1K<n<10K1 likes62 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.