CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bond005 /sova_rudevices Dataset Card for sova_rudevices Dataset Summary SOVA Dataset is free public STT/ASR dataset. It consists of several parts, one of them is SOVA RuDevices. This part is an acoustic corpus of approximately 100 hours of 16kHz Russian live speech with manual annotating, prepared by SOVA.ai team. Authors do not divide the dataset into train, validation and test subsets. Therefore, I was compelled to prepare this splitting. The training subset includes more than 82 hours, the… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sova_rudevices.audioautomatic-speech-recognition10K<n<100K13 likes741 downloads4y agoHugging Face02bond005 /sberdevices_golos_10h_crowd Dataset Card for sberdevices_golos_10h_crowd Dataset Summary Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated. Authors divide all dataset into train and test subsets.… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_10h_crowd.audioautomatic-speech-recognition10K<n<100K7 likes697 downloads4y agoHugging Face03bond005 /sberdevices_golos_100h_farfield Dataset Card for sberdevices_golos_100h_farfield Dataset Summary Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated. Authors divide all dataset into train and test… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_100h_farfield.audioautomatic-speech-recognition10K<n<100K6 likes331 downloads4y agoHugging Face04bonalor /synthetic_maritime_radio_communication MARTTS: Maritime Radio Text-To-Speech Synthetic Corpus Synthetic VHF Maritime Communication Data for Robust ASR Evaluation Dataset accompanying the paper:A Text-to-Speech Framework for Generating Synthetic Maritime Radio Communications in ASR Evaluation Dataset Summary MARTTS is an open-source synthetic speech corpus designed to evaluate and stress-test Automatic Speech Recognition (ASR) systems operating in maritime VHF radiotelephony environments. The… See the full description on the dataset page: https://huggingface.co/datasets/bonalor/synthetic_maritime_radio_communication.audiotext-classificationn<1K4 likes147 downloads9mo agoHugging Face05bonelag /ccvoice VieTTS Multi-Voice 24kHz Vietnamese Speech Dataset Bộ dữ liệu giọng nói tiếng Việt đa ngữ điệu (Multi-Speaker) quy mô lớn chất lượng cao chuẩn 24kHz 16-bit Mono PCM WAV, được biên tập và chuẩn hoá 100% tiếng Việt phục vụ huấn luyện (training & fine-tuning) các mô hình chuyển văn bản thành giọng nói (Text-to-Speech) hiện đại như Kokoro-Vietnamese, VieNeu-TTS, F5-TTS, VITS, Matcha-TTS, StyleTTS 2 cũng như đánh giá nhận dạng tiếng nói (ASR). 1. Tổng quan dữ liệu… See the full description on the dataset page: https://huggingface.co/datasets/bonelag/ccvoice.audiotext-to-speech100K<n<1M0 likes132 downloads16d agoHugging Face06zak-bonn /uwb_atcc Dataset Card for UWB-ATCC corpus Dataset Summary The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data… See the full description on the dataset page: https://huggingface.co/datasets/zak-bonn/uwb_atcc.audioautomatic-speech-recognition10K<n<100K0 likes22 downloads2mo agoHugging Face07zak-bonn /atcosim_corpus Dataset Card for ATCOSIM corpus Dataset Summary The ATCOSIM Air Traffic Control Simulation Speech corpus is a speech database of air traffic control (ATC) operator speech, provided by Graz University of Technology (TUG) and Eurocontrol Experimental Centre (EEC). It consists of ten hours of speech data, which were recorded during ATC real-time simulations using a close-talk headset microphone. The utterances are in English language and pronounced by ten non-native… See the full description on the dataset page: https://huggingface.co/datasets/zak-bonn/atcosim_corpus.audioautomatic-speech-recognition1K<n<10K0 likes17 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.