datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sova_rudevices
Dataset Card for sova_rudevices
Dataset Summary
SOVA Dataset is free public STT/ASR dataset. It consists of several parts, one of them is SOVA RuDevices. This part is an acoustic corpus of approximately 100 hours of 16kHz Russian live speech with manual annotating, prepared by SOVA.ai team.
Authors do not divide the dataset into train, validation and test subsets. Therefore, I was compelled to prepare this splitting. The training subset includes more than 82 hours, the… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sova_rudevices.rulibrispeech
Dataset Card for "rulibrispeech"
More Information needed
sberdevices_golos_10h_crowd
Dataset Card for sberdevices_golos_10h_crowd
Dataset Summary
Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated.
Authors divide all dataset into train and test subsets.… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_10h_crowd.sberdevices_golos_100h_farfield
Dataset Card for sberdevices_golos_100h_farfield
Dataset Summary
Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated.
Authors divide all dataset into train and test… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_100h_farfield.vi1000h
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/bonelag/vi1000h.podlodka_speechsynthetic_maritime_radio_communication
MARTTS: Maritime Radio Text-To-Speech Synthetic Corpus
Synthetic VHF Maritime Communication Data for Robust ASR Evaluation
Dataset accompanying the paper:A Text-to-Speech Framework for Generating Synthetic Maritime Radio Communications in ASR Evaluation
Dataset Summary
MARTTS is an open-source synthetic speech corpus designed to evaluate and stress-test Automatic Speech Recognition (ASR) systems operating in maritime VHF radiotelephony environments.
The… See the full description on the dataset page: https://huggingface.co/datasets/bonalor/synthetic_maritime_radio_communication.audioset-nonspeech
Audioset-Nonspeech
Audioset-Nonspeech is a processed version of the well-known agkphysics/AudioSet dataset. The processing was performed to remove all audio recordings that may contain clearly distinguishable human speech, leaving only non-speech audio recordings. The resulting Audioset-Nonspeech dataset can be used not only for audio event classification but also for augmentation (mixing with a specified signal-to-noise ratio) of speech recordings when training speech recognition… See the full description on the dataset page: https://huggingface.co/datasets/bond005/audioset-nonspeech.ccvoice
VieTTS Multi-Voice 24kHz Vietnamese Speech Dataset
Bộ dữ liệu giọng nói tiếng Việt đa ngữ điệu (Multi-Speaker) quy mô lớn chất lượng cao chuẩn 24kHz 16-bit Mono PCM WAV, được biên tập và chuẩn hoá 100% tiếng Việt phục vụ huấn luyện (training & fine-tuning) các mô hình chuyển văn bản thành giọng nói (Text-to-Speech) hiện đại như Kokoro-Vietnamese, VieNeu-TTS, F5-TTS, VITS, Matcha-TTS, StyleTTS 2 cũng như đánh giá nhận dạng tiếng nói (ASR).
1. Tổng quan dữ liệu… See the full description on the dataset page: https://huggingface.co/datasets/bonelag/ccvoice.taiga_speechzh-yue-tts-datasetasvspoof5-bonafidesberdevices_golos_10h_crowd_noised_2db
Dataset Card for "sberdevices_golos_10h_crowd_noised_2db"
More Information needed
taiga_speech_v2
Dataset Card for "taiga_speech_v2"
More Information needed
xmad-bonafidearabic-deepfake-audio-bonafidebrspeech-bonafide1776-track-clips
1776 Track Clips
Fast, wordless, loop-safe background music built for Shorts, Reels, and edits.
This dataset contains 1776 unique audio clips designed for creators, editors, and developers who need clean, reusable background music without lyrics.
What’s included
1776 unique clips
No lyrics (wordless hooks)
Clean, loop-safe structure
Optimized for short-form video
Editor-first design
Preview
See the preview video(s) in this repository for examples across… See the full description on the dataset page: https://huggingface.co/datasets/BonusLockSMith/1776-track-clips.pyara-bonafideuwb_atcc
Dataset Card for UWB-ATCC corpus
Dataset Summary
The UWB-ATCC Corpus is provided provided by University of West Bohemia, Department of Cybernetics. The corpus contains recordings of communication between air traffic controllers and pilots. The speech is manually transcribed and labeled with the information about the speaker (pilot/controller, not the full identity of the person). The corpus is currently small (20 hours) but we plan to search for additional data… See the full description on the dataset page: https://huggingface.co/datasets/zak-bonn/uwb_atcc.golos-testatcosim_corpus
Dataset Card for ATCOSIM corpus
Dataset Summary
The ATCOSIM Air Traffic Control Simulation Speech corpus is a speech database of air traffic control (ATC) operator speech, provided by Graz University of Technology (TUG) and Eurocontrol Experimental Centre (EEC). It consists of ten hours of speech data, which were recorded during ATC real-time simulations using a close-talk headset microphone. The utterances are in English language and pronounced by ten non-native… See the full description on the dataset page: https://huggingface.co/datasets/zak-bonn/atcosim_corpus.arabic-deepfake-audio-bonafidearabic-deepfake-audio-bonafide-testNapoleon_Bonaparte_voice_databonitoBonzizh-yue-tts-dataset-100bonitobondar-taisa-pavetrany-zamak-na-dvaih-kacjaryna-jagorava
Bondar Taisa, Pavetrany zamak na dvaih, Kacjaryna Jagorava
Metadata
Original Title (Cyrillic): Бондар Таіса, Паветраны замак на дваіх, Кацярына Ягорава
Transliterated Title: Bondar Taisa, Pavetrany zamak na dvaih, Kacjaryna Jagorava
Audio Files: 174 MP3 files
Format: Belarusian audiobook
Description
This is a Belarusian audiobook dataset containing 174 audio tracks.
License
Please check the original source for licensing information.
