datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
infore2_audiobooks
unofficial mirror of InfoRe Technology public dataset №2
official announcement: https://www.facebook.com/groups/j2team.community/permalink/1010834009248719/
415h, 315k samples, vietnamese audiobooks of chinese wǔxiá 武俠 & xiānxiá 仙俠
bộ dữ liệu bóc ra từ YouTube đọc truyện võ hiệp & tiên hiệp, áp dụng kĩ thuật đối chiếu văn bản để dán nhãn tự động
official download:… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/infore2_audiobooks.audiobooks170 hours of aligned audiobooks taken from tatkniga.ru. There are 4 speakers with 17+ hours of audio and 20 speakers in total. All the books are in free access and most of them in public domain.
sova_rudevices_audiobooks
Dataset instance structure
{'audio': {'path': '/path/to/wav.wav',
'array': array([wav numpy array]), dtype=float32),
'sampling_rate': 16000},
'transcription': 'транскрипция'}
Dataset audio info
16000 Hz
wav
mono
Russian speech from audiobooks
Citation
@misc{sova2021rudevices,
author = {Zubarev, Egor and Moskalets, Timofey and SOVA.ai},
title = {SOVA RuDevices Dataset: free public STT/ASR dataset with manually annotated live speech},
publisher =… See the full description on the dataset page: https://huggingface.co/datasets/dangrebenkin/sova_rudevices_audiobooks.audiobooks-xxlqwen3-tts-zh-audiobooksturkish-tts-audiobooks
Turkish TTS Audiobooks
Turkish read-speech corpus for text-to-speech training, built from Turkish
audiobook and spoken-article recordings by an automatic pipeline: VAD
segmentation → technical QC → acoustic event tagging → DNSMOS → speaker
embedding/consistency → double-pass Whisper ASR → text policy → leakage-free
splitting. Audio is 16 kHz mono lossless FLAC embedded in the Parquet shards.
The pipeline that produced it — every stage, every threshold, the export and
audit… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-tts-audiobooks.bulgarian-audiobooks-tts-400hBulgarian Female Audiobook TTS Dataset (400 Hours)
Description
This dataset contains over 400 hours of high-quality Bulgarian speech audio, specifically curated for training and fine-tuning Text-to-Speech (TTS) models. The data is sourced from various audiobooks and has undergone a rigorous filtering and cleaning process to ensure high training stability.
Dataset Specifications
Total Duration: ~400 hours (post-filtering)
Total Segments: ~200,000
Segment Length: 4 – 12 seconds
Language:… See the full description on the dataset page: https://huggingface.co/datasets/beleata74/bulgarian-audiobooks-tts-400h.audiobooks
Crimean Tatar Audiobooks
Dataset Summary
Crimean Tatar Audiobooks is a speech dataset sourced from different sources (public radio stations/youtube channels) containing audiobooks in Crimean Tatar. The dataset comprises recordings of different native fiction books, all read by a single female speaker (for now). The dataset is intended for text-to-speech (TTS) research and development in the Crimean Tatar language.
25:34:53 of aligned speech in 22,517 segments… See the full description on the dataset page: https://huggingface.co/datasets/gaydmi/audiobooks.narrated-audiobooks-brsbs
Narrated Bashkir Audiobooks — BRSBS (Bashkorttele)
≈ 78.7 hours of human-narrated audiobooks, primarily in the Bashkir language, drawn from public-domain literary works and folk epics. Recorded as accessible "talking books" by the Bashkir Republican Special Library for the Blind (BRSBS) and released for language preservation and AI/ML research.
🌐 Languages of this card: English · Башҡортса · Русский
This dataset is part of a larger series published under the Bashkorttele… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/narrated-audiobooks-brsbs.uyghur-audiobooks
Uyghur Audiobooks Dataset
维吾尔语有声书数据集,包含多个经典有声书作品。
数据集内容
语言: 维吾尔语 (Uyghur)
格式: MP3 音频 + TXT 文本
数量: 228 个有声书作品
文件结构
remaining_data.tar.gz # 包含所有228个文件夹的压缩包
使用方法
下载后解压:
tar -xzvf remaining_data.tar.gz
许可证
根据原始数据许可证
hindi_audiobooksukrainian-tts-audiobooks-24khz
Ukrainian Audiobook TTS Dataset (24 kHz)
Description
Ukrainian speech dataset for TTS and ASR tasks.
Source Dataset
https://huggingface.co/datasets/Yehor/audiobooks-xxl
Processing Pipeline
MusicDetection filtering — removed samples with background music/noise
Audio processing (Sidon) — resampled 16 kHz → 24 kHz, converted to mono
Transcription — generated with nvidia/canary-1b-v2
Dataset Structure
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/Mikhailo/ukrainian-tts-audiobooks-24khz.audiobooksazerbaijani-audiobooksasr-tts-rioplatense-audiobooks-argentina
ASR Dataset - Rioplatense Spanish Audiobooks (Educ.ar Lecturas grabadas)
Description
This dataset contains short audio segments extracted from Rioplatense (Argentinian) Spanish audiobooks, together with their aligned transcriptions.The audio comes from the public collection "Lecturas grabadas" published by Educ.ar, and has been automatically segmented and matched to the corresponding text.
Important: This is a derivative dataset built on top of materials created and… See the full description on the dataset page: https://huggingface.co/datasets/frizynn/asr-tts-rioplatense-audiobooks-argentina.bg-audiobooks-tts
Bulgarian Audiobook Speech Dataset
A high-quality Bulgarian speech dataset derived from audiobooks narrated by Plamen Sivov,
suitable for text-to-speech (TTS) and automatic speech recognition (ASR) tasks.
Dataset Summary
Property
Value
Language
Bulgarian (bg)
Total Duration
15.2 hours
Total Clips
10,627
Speaker
Plamen Sivov (single speaker)
Source
YouTube audiobooks
Sample Rate
24,000 Hz
Audio Format
WAV, mono, 16-bit PCM
Clip Duration
3–15… See the full description on the dataset page: https://huggingface.co/datasets/raditotev/bg-audiobooks-tts.buriy_audiobooks_2_valaudiobooks_ua_test
About dataset
It is a dataset of ukrainian audiobooksEach sample contain an approximately 8 seconds od ukrainian speech
Loading script
>>> load_dataset("Zarakun/audiobooks_ua_test")
Dataset structure
**Every example has the following:
audio - the waveformrate - the sampling rate of the waveformfile_id - the id of the speakerduration - the duration of the video in secondssentence - the transcript of the video
tajik-classic-audiobooks
Tajik Classic Audiobooks
Chapter-level audiobooks of classic Tajik/Persian literature, narrated by a single fine-tuned
Chatterbox-multilingual Tajik TTS voice (synthetic). Text was cleaned to spoken form (numbers
spelled out, no digits/Latin) and segmented per chapter. Each book folder holds mp3/ch_NNN.mp3
plus the source chapters_index.json and QC report.
Books included: shakuri_khuroson, shakuri_panturkizm, sadi_guliston, sadi_buston, nizami_layli, nizami_makhzan_khusrav… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/tajik-classic-audiobooks.AudioBooksInstructGemini2.5
AudioBooksInstructGemini2.5
Датасет инструкций для обучения аудио-языковых моделей, сгенерированный с помощью Gemini 2.5 Flash на основе ToneBooks.
Описание
Датасет содержит аудиозаписи из русскоязычных аудиокниг с автоматически сгенерированными вопросами и ответами для обучения моделей следовать инструкциям.
Типы задач
Задача
Описание
Примерный %
instruction_following
Общие инструкции по работе с аудио (перевод, перефразирование, анализ)
~25%… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/AudioBooksInstructGemini2.5.tajik-audiobooks
Saidzoda Lab — Gated Research Dataset
Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz).
Dataset contents, provenance, and statistics are not publicly disclosed.
Access is granted manually on request.
tajik-audiobooks-chapters
Saidzoda Lab — Gated Research Dataset
Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz).
Dataset contents, provenance, and statistics are not publicly disclosed.
Access is granted manually on request.
ting-audiobooksaudiobook-summariesaudiobooks-archiveaudiobooksar_msa_audiobooks6SOVA-audiobooks-100kaudiobooks-collection
