CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NCSpeech /YO-CPT-ru YO-CPT-ru YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.audiotext-to-speech1M<n<10M16 likes11k downloads2mo agoHugging Face02langswap /dialogs-ru-emotional-conversations Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational Russian speech, designed for dialog-oriented and emotional text-to-speech. Unlike existing Russian corpora — mostly single-speaker read speech or large but low-quality web-mined audio — Dialogs was recorded by professional theatre actors performing scripted dialogs face-to-face, capturing natural turn-taking, timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.audiotext-to-speechn<1K18 likes1.7k downloads2mo agoHugging Face03istupakov /russian_librispeech Russian LibriSpeech (RuLS) Identifier: SLR96 from openslr.org Summary: This dataset is based on LibriVox audiobooks Category: Speech License: The dataset is Public Domain in the USA. About this resource: Russian LibriSpeech (RuLS) dataset is based on LibriVox's public domain audio books (see BOOKS.TXT for the list of included books) and contains about 98 hours of audio data. audioautomatic-speech-recognition10K<n<100K6 likes928 downloads1y agoHugging Face04bond005 /sova_rudevices Dataset Card for sova_rudevices Dataset Summary SOVA Dataset is free public STT/ASR dataset. It consists of several parts, one of them is SOVA RuDevices. This part is an acoustic corpus of approximately 100 hours of 16kHz Russian live speech with manual annotating, prepared by SOVA.ai team. Authors do not divide the dataset into train, validation and test subsets. Therefore, I was compelled to prepare this splitting. The training subset includes more than 82 hours, the… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sova_rudevices.audioautomatic-speech-recognition10K<n<100K13 likes741 downloads4y agoHugging Face05ai4bharat /Rural_Women_Bhojpuri Rural Bhojpuri ASR Dataset Dataset Description This dataset is curated to foster the development of inclusive Automatic Speech Recognition (ASR) systems, with a special focus on the underrepresented voices of rural Bhojpuri women. It contains audio clips in both Bhojpuri and Hindi, collected from real-world and synthetic sources, designed to train and evaluate ASR models that can accurately recognize diverse speech patterns. This work is part of the research presented in… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Rural_Women_Bhojpuri.audioautomatic-speech-recognition10K<n<100K6 likes270 downloads1y agoHugging Face06SynDataLab-EN /omnivoice-ru Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,999 Total 19,999 audiotext-to-speech10K<n<100K0 likes160 downloads6mo agoHugging Face07ivkond /synthetic-speech-diarization-ru synthetic-speech-diarization-ru Synthetic speech diarization dataset in Parquet format. Dataset Details Number of tracks: 2000 Sampling rate: 16000 Hz Audio format: Embedded in Parquet files (Audio feature compatible) Storage: Parquet format for efficient loading Dataset Structure The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format. Features audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/ivkond/synthetic-speech-diarization-ru.tabularautomatic-speech-recognition1K<n<10K0 likes93 downloads10mo agoHugging Face08turnipseason /paralingua_ru Russian Paralinguistic Annotation Dataset Датасет паралингвистической разметки спикеров из трёх русскоязычных корпусов: biggest_ru_book, DeepSpeech и Golos. Что размечалось Каждое аудио размечалось вручную по следующим характеристикам: Поле Описание Пример значений gender Пол спикера мужской, женский age_group Возрастная группа молодой, взрослый, пожилой voice_pitch Высота голоса низкий, средний, высокий loudness Громкость тихий, нормальный… See the full description on the dataset page: https://huggingface.co/datasets/turnipseason/paralingua_ru.tabulartext-to-speech100K<n<1M8 likes84 downloads4mo agoHugging Face09lab260 /biggest_ru_book_balalaika Biggest-Ru-Book Annotated by Balalaika [!IMPORTANT] Official dataset for our INTERSPEECH 2026 paper "A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika. If you use this resource, please cite it. A curated Russian speech dataset for advanced speech generative tasks. Overview… See the full description on the dataset page: https://huggingface.co/datasets/lab260/biggest_ru_book_balalaika.tabulartext-to-speech100K<n<1M3 likes81 downloads3mo agoHugging Face10AigizK /notebooklm_rus NotebookLM Russian Podcast Dataset Датасет содержит записи подкастов, сгенерированных с помощью Google NotebookLM на русском языке. Описание Голоса: 2 диктора — мужской и женский Общая длительность: 77 ч 23 мин 22 сек Количество эпизодов: 417 Формат аудио: WAV, 24 kHz, моно Язык: русский Структура датасета Поле Тип Описание audio Audio Аудиозапись эпизода (24 kHz, моно) transcription string Полная текстовая расшифровка эпизода segments string… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/notebooklm_rus.audiotext-to-speechn<1K3 likes79 downloads6mo agoHugging Face11niobures /synthetic-speech-diarization-ru synthetic-speech-diarization-ru Synthetic speech diarization dataset in Parquet format. Dataset Details Number of tracks: 2000 Sampling rate: 16000 Hz Audio format: Embedded in Parquet files (Audio feature compatible) Storage: Parquet format for efficient loading Dataset Structure The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format. Features audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/niobures/synthetic-speech-diarization-ru.tabularautomatic-speech-recognition1K<n<10K0 likes73 downloads5mo agoHugging Face12Thomcles /YodaLingua-Russiangated YodaLingua-Russian YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Russian portion of the multilingual YodaLingua collection. 🧾 Dataset Overview Property Value Total clips 67,482 audio–transcription pairs Total duration 192 hours Speakers 2,611 distinct speakers Audio format MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Russian.audiotext-to-speech10K<n<100K2 likes23 downloads5mo agoHugging Face13rushilrawat /garhwali-speech Garhwali Speech Companion to Garhwali Corpus. This repository has separate configs for Project VAANI and Meta Omnilingual speech; choose one source config at a time because their splits and transcript histories differ. Contents Combined configs: 113,363 source rows, 113,350 unique audio hashes, 154.65 hours, and about 16.91 GiB of source audio. Meta Omnilingual: 2,927 additional recordings, 19.14 hours (train 2,329, validation 298, test 300). Overlap audit: 10… See the full description on the dataset page: https://huggingface.co/datasets/rushilrawat/garhwali-speech.audioautomatic-speech-recognition100K<n<1M0 likes22 downloads10h agoHugging Face14Devvrat024 /Rural_Women_Bhojpuri Rural Bhojpuri ASR Dataset Dataset Description This dataset is curated to foster the development of inclusive Automatic Speech Recognition (ASR) systems, with a special focus on the underrepresented voices of rural Bhojpuri women. It contains audio clips in both Bhojpuri and Hindi, collected from real-world and synthetic sources, designed to train and evaluate ASR models that can accurately recognize diverse speech patterns. This work is part of the research presented in… See the full description on the dataset page: https://huggingface.co/datasets/Devvrat024/Rural_Women_Bhojpuri.audioautomatic-speech-recognition10K<n<100K0 likes18 downloads6mo agoHugging Face15fosters /shata_rustaveli_vitsyaz_u_tygravai_shkury_all Віцязь у тыгравай скуры Аўтар / Author: Шата РуставеліМова / Language: Беларуская (Belarusian) Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд. Частка калекцыі Belarusian Audiobooks (native). Радкоў у датасеце 1,091 Працягласць 3 гадз 24 хв Частата дыскрэтызацыі 44100 Hz Каналы мона Даўжыня фрагмента да 30 с Структура Кожны радок змяшчае: audio — аўдыёфрагмент (native SR, мона… See the full description on the dataset page: https://huggingface.co/datasets/fosters/shata_rustaveli_vitsyaz_u_tygravai_shkury_all.audioautomatic-speech-recognition1K<n<10K0 likes18 downloads3mo agoHugging Face16fosters /shata_rustaveli_vitsyaz_u_tygravai_shkury_output_original Віцязь у тыгравай скуры — арыгінальнае аўдыё Аўтар / Author: Шата РуставеліМова / Language: Беларуская (Belarusian) Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці. Частка калекцыі Ministerskija — корпус беларускіх аўдыёкніг. Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя): shata_rustaveli_vitsyaz_u_tygravai_shkury_output Доўгасць аўдыё 3h31m Радкоў у датасеце 1,027 Структура Кожны радок змяшчае: audio —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/shata_rustaveli_vitsyaz_u_tygravai_shkury_output_original.audioautomatic-speech-recognition1K<n<10K0 likes16 downloads4mo agoHugging Face17RakhatM /kk-ru-pharma-ttsgated Kazakh/Russian Pharmaceutical TTS Corpus A synthetic speech corpus of pharmaceutical / clinical phrases in Kazakh (kk) and Russian (ru), synthesized with a multilingual Orpheus TTS model. Designed for ASR auto-adaptation experiments: the splits cover seen / unseen speakers and matched / unseen evaluation conditions for benchmarking domain and speaker generalization in low-resource medical ASR. Splits Split Clips Purpose train 27,182 Training. RU voices: Elena… See the full description on the dataset page: https://huggingface.co/datasets/RakhatM/kk-ru-pharma-tts.audioautomatic-speech-recognition10K<n<100K1 likes15 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.