CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bond005 /sova_rudevices Dataset Card for sova_rudevices Dataset Summary SOVA Dataset is free public STT/ASR dataset. It consists of several parts, one of them is SOVA RuDevices. This part is an acoustic corpus of approximately 100 hours of 16kHz Russian live speech with manual annotating, prepared by SOVA.ai team. Authors do not divide the dataset into train, validation and test subsets. Therefore, I was compelled to prepare this splitting. The training subset includes more than 82 hours, the… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sova_rudevices.audioautomatic-speech-recognition10K<n<100K13 likes730 downloads4y agoHugging Face02bond005 /rulibrispeech Dataset Card for "rulibrispeech" More Information needed audio10K<n<100K2 likes686 downloads4y agoHugging Face03bond005 /sberdevices_golos_10h_crowd Dataset Card for sberdevices_golos_10h_crowd Dataset Summary Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated. Authors divide all dataset into train and test subsets.… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_10h_crowd.audioautomatic-speech-recognition10K<n<100K7 likes681 downloads4y agoHugging Face04bond005 /sberdevices_golos_100h_farfield Dataset Card for sberdevices_golos_100h_farfield Dataset Summary Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated. Authors divide all dataset into train and test… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_100h_farfield.audioautomatic-speech-recognition10K<n<100K6 likes326 downloads4y agoHugging Face05bonelag /vi1000h Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus Dataset Summary Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team. Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling. This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/bonelag/vi1000h.audio100K<n<1M1 likes196 downloads15d agoHugging Face06bond005 /podlodka_speechaudion<1K1 likes159 downloads2y agoHugging Face07bonalor /synthetic_maritime_radio_communication MARTTS: Maritime Radio Text-To-Speech Synthetic Corpus Synthetic VHF Maritime Communication Data for Robust ASR Evaluation Dataset accompanying the paper:A Text-to-Speech Framework for Generating Synthetic Maritime Radio Communications in ASR Evaluation Dataset Summary MARTTS is an open-source synthetic speech corpus designed to evaluate and stress-test Automatic Speech Recognition (ASR) systems operating in maritime VHF radiotelephony environments. The… See the full description on the dataset page: https://huggingface.co/datasets/bonalor/synthetic_maritime_radio_communication.audiotext-classificationn<1K4 likes146 downloads9mo agoHugging Face08bond005 /audioset-nonspeech Audioset-Nonspeech Audioset-Nonspeech is a processed version of the well-known agkphysics/AudioSet dataset. The processing was performed to remove all audio recordings that may contain clearly distinguishable human speech, leaving only non-speech audio recordings. The resulting Audioset-Nonspeech dataset can be used not only for audio event classification but also for augmentation (mixing with a specified signal-to-noise ratio) of speech recordings when training speech recognition… See the full description on the dataset page: https://huggingface.co/datasets/bond005/audioset-nonspeech.audioaudio-classification10K<n<100K3 likes140 downloads1y agoHugging Face09bonelag /ccvoice VieTTS Multi-Voice 24kHz Vietnamese Speech Dataset Bộ dữ liệu giọng nói tiếng Việt đa ngữ điệu (Multi-Speaker) quy mô lớn chất lượng cao chuẩn 24kHz 16-bit Mono PCM WAV, được biên tập và chuẩn hoá 100% tiếng Việt phục vụ huấn luyện (training & fine-tuning) các mô hình chuyển văn bản thành giọng nói (Text-to-Speech) hiện đại như Kokoro-Vietnamese, VieNeu-TTS, F5-TTS, VITS, Matcha-TTS, StyleTTS 2 cũng như đánh giá nhận dạng tiếng nói (ASR). 1. Tổng quan dữ liệu… See the full description on the dataset page: https://huggingface.co/datasets/bonelag/ccvoice.audiotext-to-speech100K<n<1M0 likes128 downloads14d agoHugging Face10bond005 /taiga_speechaudio100K<n<1M0 likes97 downloads2y agoHugging Face11boniromou /zh-yue-tts-datasetaudio10K<n<100K1 likes53 downloads2y agoHugging Face1234data /asvspoof5-bonafideaudio1K<n<10K0 likes41 downloads2mo agoHugging Face13bond005 /sberdevices_golos_10h_crowd_noised_2db Dataset Card for "sberdevices_golos_10h_crowd_noised_2db" More Information needed audio10K<n<100K0 likes40 downloads3y agoHugging Face14bond005 /taiga_speech_v2 Dataset Card for "taiga_speech_v2" More Information needed audio10K<n<100K4 likes39 downloads2y agoHugging Face1534data /xmad-bonafideaudio1K<n<10K0 likes32 downloads2mo agoHugging Face16kalyaninaganaboyina2006 /arabic-deepfake-audio-bonafideaudio10K<n<100K0 likes32 downloads2mo agoHugging Face1734data /brspeech-bonafideaudio1K<n<10K0 likes31 downloads2mo agoHugging Face18BonusLockSMith /1776-track-clips 1776 Track Clips Fast, wordless, loop-safe background music built for Shorts, Reels, and edits. This dataset contains 1776 unique audio clips designed for creators, editors, and developers who need clean, reusable background music without lyrics. What’s included 1776 unique clips No lyrics (wordless hooks) Clean, loop-safe structure Optimized for short-form video Editor-first design Preview See the preview video(s) in this repository for examples across… See the full description on the dataset page: https://huggingface.co/datasets/BonusLockSMith/1776-track-clips.audio1K<n<10K0 likes28 downloads9mo agoHugging Face1934data /pyara-bonafideaudio1K<n<10K0 likes28 downloads2mo agoHugging Face20zak-bonn /uwb_atcc Dataset Card for UWB-ATCC corpus Dataset Summary The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data… See the full description on the dataset page: https://huggingface.co/datasets/zak-bonn/uwb_atcc.audioautomatic-speech-recognition10K<n<100K0 likes23 downloads2mo agoHugging Face21bonlime /golos-testaudio10K<n<100K0 likes21 downloads5mo agoHugging Face22zak-bonn /atcosim_corpus Dataset Card for ATCOSIM corpus Dataset Summary The ATCOSIM Air Traffic Control Simulation Speech corpus is a speech database of air traffic control (ATC) operator speech, provided by Graz University of Technology (TUG) and Eurocontrol Experimental Centre (EEC). It consists of ten hours of speech data, which were recorded during ATC real-time simulations using a close-talk headset microphone. The utterances are in English language and pronounced by ten non-native… See the full description on the dataset page: https://huggingface.co/datasets/zak-bonn/atcosim_corpus.audioautomatic-speech-recognition1K<n<10K0 likes18 downloads2mo agoHugging Face23AhmedAshrafMarzouk /arabic-deepfake-audio-bonafideaudio10K<n<100K0 likes14 downloads5mo agoHugging Face24AhmedAshrafMarzouk /arabic-deepfake-audio-bonafide-testaudio1K<n<10K0 likes12 downloads5mo agoHugging Face25Lancelot457357 /Napoleon_Bonaparte_voice_dataaudion<1K0 likes11 downloads3y agoHugging Face26biauser /bonitoaudion<1K0 likes7 downloads3y agoHugging Face27sdadasfgdfgfdg /Bonziaudion<1K0 likes7 downloads3y agoHugging Face28boniromou /zh-yue-tts-dataset-100audion<1K2 likes7 downloads2y agoHugging Face29biadrivex /bonitoaudion<1K0 likes6 downloads3y agoHugging Face30keinelust /bondar-taisa-pavetrany-zamak-na-dvaih-kacjaryna-jagorava Bondar Taisa, Pavetrany zamak na dvaih, Kacjaryna Jagorava Metadata Original Title (Cyrillic): Бондар Таіса, Паветраны замак на дваіх, Кацярына Ягорава Transliterated Title: Bondar Taisa, Pavetrany zamak na dvaih, Kacjaryna Jagorava Audio Files: 174 MP3 files Format: Belarusian audiobook Description This is a Belarusian audiobook dataset containing 174 audio tracks. License Please check the original source for licensing information. audion<1K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.