datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
russian_librispeech
Russian LibriSpeech (RuLS)
Identifier: SLR96 from openslr.org
Summary: This dataset is based on LibriVox audiobooks
Category: Speech
License: The dataset is Public Domain in the USA.
About this resource:
Russian LibriSpeech (RuLS) dataset is based on LibriVox's public domain audio books (see BOOKS.TXT for the list of included books) and contains about 98 hours of audio data.
notebooklm_rus
NotebookLM Russian Podcast Dataset
Датасет содержит записи подкастов, сгенерированных с помощью Google NotebookLM на русском языке.
Описание
Голоса: 2 диктора — мужской и женский
Общая длительность: 77 ч 23 мин 22 сек
Количество эпизодов: 417
Формат аудио: WAV, 24 kHz, моно
Язык: русский
Структура датасета
Поле
Тип
Описание
audio
Audio
Аудиозапись эпизода (24 kHz, моно)
transcription
string
Полная текстовая расшифровка эпизода
segments
string… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/notebooklm_rus.uzbek-asr-train-manifests
Uzbek ASR Training Manifests
The exact training, validation and test splits behind
rustam1221/uzbek-asr-gigaam:
974 hours of Uzbek speech drawn from seven public corpora, filtered, text-normalized,
and split by speaker.
No audio is copied. Each row is a pointer — a parquet file plus a row
index in the upstream dataset — and the training dataloader decodes the audio
when the batch is built. That keeps the whole corpus definition at 200 MB
instead of roughly a terabyte of… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-train-manifests.russian-speech-dataset
Russian Speech Dataset
The Russian Speech Dataset is a structured speech audio dataset designed to deliver high-quality audio data for machine learning and AI-driven voice systems. It includes 91 hours of audio data distributed across 641 files, provided in MP3 and WAV formats with a total size of 307 MB.
This well-organized audio dataset ensures balanced voice data, with 50% female and 50% male speakers, and a broad age distribution from 18 to 50+ years. The dataset language is… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/russian-speech-dataset.russian-speech-recognition-dataset
Russian Telephone Dialogues Dataset - 338 Hours
The Russian speech dataset includes 338 hours of telephone dialogues in Russian from 460 native speakers, offering high-quality audio recordings with detailed annotations (text, speaker ID, gender, age) to support speech recognition systems, natural language processing, and deep learning models for building accurate Russian dialogue and audio datasets. - Get the data
Dataset characteristics:
Characteristic
Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/russian-speech-recognition-dataset.YodaLingua-Russian
YodaLingua-Russian
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Russian portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
67,482 audio–transcription pairs
Total duration
192 hours
Speakers
2,611 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Russian.garhwali-speech
Garhwali Speech
Companion to Garhwali Corpus,
which is currently private while its redacted catalog and rights-reviewed text
release are reconciled. This dataset packages the Garhwali subset of Project
VAANI audio with its provider transcripts, source metadata, and clearly
separated experimental SraVaani drafts.
Contents
110,436 source audio rows (14.54 GiB audio).
5,894 provider-transcribed rows, including the
provider's train/validation/test splits.
104,542… See the full description on the dataset page: https://huggingface.co/datasets/rushilrawat/garhwali-speech.shata_rustaveli_vitsyaz_u_tygravai_shkury_all
Віцязь у тыгравай скуры
Аўтар / Author: Шата РуставеліМова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
1,091
Працягласць
3 гадз 24 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона… See the full description on the dataset page: https://huggingface.co/datasets/fosters/shata_rustaveli_vitsyaz_u_tygravai_shkury_all.shata_rustaveli_vitsyaz_u_tygravai_shkury_output_original
Віцязь у тыгравай скуры — арыгінальнае аўдыё
Аўтар / Author: Шата РуставеліМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
shata_rustaveli_vitsyaz_u_tygravai_shkury_output
Доўгасць аўдыё
3h31m
Радкоў у датасеце
1,027
Структура
Кожны радок змяшчае:
audio —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/shata_rustaveli_vitsyaz_u_tygravai_shkury_output_original.
