datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YO-CPT-ru
YO-CPT-ru
YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily
quality-filtered corpus of Russian speech mined from YouTube (via YODAS2)
and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an
ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level
forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a
speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.dialogs-ru-emotional-conversations
Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus
Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational
Russian speech, designed for dialog-oriented and emotional text-to-speech.
Unlike existing Russian corpora — mostly single-speaker read speech or large but
low-quality web-mined audio — Dialogs was recorded by professional theatre actors
performing scripted dialogs face-to-face, capturing natural turn-taking,
timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio.
russian_librispeech
Russian LibriSpeech (RuLS)
Identifier: SLR96 from openslr.org
Summary: This dataset is based on LibriVox audiobooks
Category: Speech
License: The dataset is Public Domain in the USA.
About this resource:
Russian LibriSpeech (RuLS) dataset is based on LibriVox's public domain audio books (see BOOKS.TXT for the list of included books) and contains about 98 hours of audio data.
sova_rudevices
Dataset Card for sova_rudevices
Dataset Summary
SOVA Dataset is free public STT/ASR dataset. It consists of several parts, one of them is SOVA RuDevices. This part is an acoustic corpus of approximately 100 hours of 16kHz Russian live speech with manual annotating, prepared by SOVA.ai team.
Authors do not divide the dataset into train, validation and test subsets. Therefore, I was compelled to prepare this splitting. The training subset includes more than 82 hours, the… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sova_rudevices.Rural_Women_Bhojpuri
Rural Bhojpuri ASR Dataset
Dataset Description
This dataset is curated to foster the development of inclusive Automatic Speech Recognition (ASR) systems, with a special focus on the underrepresented voices of rural Bhojpuri women. It contains audio clips in both Bhojpuri and Hindi, collected from real-world and synthetic sources, designed to train and evaluate ASR models that can accurately recognize diverse speech patterns.
This work is part of the research presented in… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Rural_Women_Bhojpuri.omnivoice-ru
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
9,999
Total
19,999
russian-call-center-speech-ru
📖 Описание на русском
ОписаниеКрупный датасет реальных записей колл-центров на русском языке.Телефонное качество, разговоры «клиент–оператор».Подходит для обучения систем ASR (распознавание речи), NLP, голосовых ассистентов и анализа диалогов.
Технические характеристики
Язык: русский
Общая продолжительность: ~832 часа
Формат: MP3
Каналы: моно (клиент и оператор в одном канале)
Частота дискретизации: 8000 Гц
Битрейт: 32 кбит/с
Метаданные: отсутствуют… See the full description on the dataset page: https://huggingface.co/datasets/MaratDV/russian-call-center-speech-ru.SpeechParaling-Bench
🤗 About This Repo
This repository contains the SpeechParaling-Bench dataset for evaluating paralinguistic-aware speech generation. The benchmark is designed to assess how well Large Audio-Language Models (LALMs) can generate speech with appropriate paralinguistic features in real-world interaction scenarios.Key Statistics:
2,000+ speech samples (Chinese & English parallel)
3 evaluation tasks
13 paralinguistic dimensions
100+ paralinguistic features
🩷… See the full description on the dataset page: https://huggingface.co/datasets/Ruohan2/SpeechParaling-Bench.synthetic-speech-diarization-ru
synthetic-speech-diarization-ru
Synthetic speech diarization dataset in Parquet format.
Dataset Details
Number of tracks: 2000
Sampling rate: 16000 Hz
Audio format: Embedded in Parquet files (Audio feature compatible)
Storage: Parquet format for efficient loading
Dataset Structure
The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format.
Features
audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/ivkond/synthetic-speech-diarization-ru.notebooklm_rus
NotebookLM Russian Podcast Dataset
Датасет содержит записи подкастов, сгенерированных с помощью Google NotebookLM на русском языке.
Описание
Голоса: 2 диктора — мужской и женский
Общая длительность: 77 ч 23 мин 22 сек
Количество эпизодов: 417
Формат аудио: WAV, 24 kHz, моно
Язык: русский
Структура датасета
Поле
Тип
Описание
audio
Audio
Аудиозапись эпизода (24 kHz, моно)
transcription
string
Полная текстовая расшифровка эпизода
segments
string… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/notebooklm_rus.synthetic-speech-diarization-ru
synthetic-speech-diarization-ru
Synthetic speech diarization dataset in Parquet format.
Dataset Details
Number of tracks: 2000
Sampling rate: 16000 Hz
Audio format: Embedded in Parquet files (Audio feature compatible)
Storage: Parquet format for efficient loading
Dataset Structure
The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format.
Features
audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/niobures/synthetic-speech-diarization-ru.garhwali-speech
Garhwali Speech
Companion to Garhwali Corpus. This
repository has separate configs for Project VAANI and Meta Omnilingual speech;
choose one source config at a time because their splits and transcript histories differ.
Contents
Combined configs: 113,363 source rows, 113,350 unique audio hashes, 154.65 hours, and about 16.91 GiB of source audio.
Meta Omnilingual: 2,927 additional recordings, 19.14 hours (train 2,329, validation 298, test 300).
Overlap audit: 10… See the full description on the dataset page: https://huggingface.co/datasets/rushilrawat/garhwali-speech.human-robot-conversation-russian
Human-Robot Dataset
The dataset comprises 660+ hours of Russian speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-russian.human-robot-conversation-russian
Human-Robot Conversation Dataset (Russian) - 660+ Hours
Dataset (Russian) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-russian.Russian_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 1,025 hours of processed Russian (RU) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Russian_Call_Center_Audio_Dataset_Dual_Channel.russian-speech-dataset
Russian Speech Dataset
The Russian Speech Dataset is a structured speech audio dataset designed to deliver high-quality audio data for machine learning and AI-driven voice systems. It includes 91 hours of audio data distributed across 641 files, provided in MP3 and WAV formats with a total size of 307 MB.
This well-organized audio dataset ensures balanced voice data, with 50% female and 50% male speakers, and a broad age distribution from 18 to 50+ years. The dataset language is… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/russian-speech-dataset.Russian-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 1,025 hours of processed Russian (RU) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Russian-Call-Center-Audio-Dataset-Single-Channel.YodaLingua-Russian
YodaLingua-Russian
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Russian portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
67,482 audio–transcription pairs
Total duration
192 hours
Speakers
2,611 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Russian.Rural_Women_Bhojpuri
Rural Bhojpuri ASR Dataset
Dataset Description
This dataset is curated to foster the development of inclusive Automatic Speech Recognition (ASR) systems, with a special focus on the underrepresented voices of rural Bhojpuri women. It contains audio clips in both Bhojpuri and Hindi, collected from real-world and synthetic sources, designed to train and evaluate ASR models that can accurately recognize diverse speech patterns.
This work is part of the research presented in… See the full description on the dataset page: https://huggingface.co/datasets/Devvrat024/Rural_Women_Bhojpuri.shata_rustaveli_vitsyaz_u_tygravai_shkury_all
Віцязь у тыгравай скуры
Аўтар / Author: Шата РуставеліМова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
1,091
Працягласць
3 гадз 24 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона… See the full description on the dataset page: https://huggingface.co/datasets/fosters/shata_rustaveli_vitsyaz_u_tygravai_shkury_all.shata_rustaveli_vitsyaz_u_tygravai_shkury_output_original
Віцязь у тыгравай скуры — арыгінальнае аўдыё
Аўтар / Author: Шата РуставеліМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
shata_rustaveli_vitsyaz_u_tygravai_shkury_output
Доўгасць аўдыё
3h31m
Радкоў у датасеце
1,027
Структура
Кожны радок змяшчае:
audio —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/shata_rustaveli_vitsyaz_u_tygravai_shkury_output_original.kk-ru-pharma-tts
Kazakh/Russian Pharmaceutical TTS Corpus
A synthetic speech corpus of pharmaceutical / clinical phrases in Kazakh (kk) and Russian (ru), synthesized with a multilingual Orpheus TTS model. Designed for ASR auto-adaptation experiments: the splits cover seen / unseen speakers and matched / unseen evaluation conditions for benchmarking domain and speaker generalization in low-resource medical ASR.
Splits
Split
Clips
Purpose
train
27,182
Training. RU voices: Elena… See the full description on the dataset page: https://huggingface.co/datasets/RakhatM/kk-ru-pharma-tts.minds14-telephony-ru
MINDS-14 telephony — Russian (ru-RU)
Processed banking IVR telephony speech derived from
PolyAI/minds14 (ru-RU / related config).
Credits & license
Original corpus: PolyAI MInDS-14Gerz et al., 2021 — arXiv:2104.08524.
License: CC BY-4.0 (commercial use allowed with attribution).
Snapshot
Language: ru (locale ru-RU)
Samples: 534 clips (~1.296 h) after GT remaster
Audio: OGG/Opus mono, native 8 kHz telephony
Split: all test (eval-only remaster;… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/minds14-telephony-ru.
