audiobooks
infore2_audiobooks
unofficial mirror of InfoRe Technology public dataset №2
official announcement: https://www.facebook.com/groups/j2team.community/permalink/1010834009248719/
415h, 315k samples, vietnamese audiobooks of chinese wǔxiá 武俠 & xiānxiá 仙俠
bộ dữ liệu bóc ra từ YouTube đọc truyện võ hiệp & tiên hiệp, áp dụng kĩ thuật đối chiếu văn bản để dán nhãn tự động
official download:… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/infore2_audiobooks.audiobooks-xxlaudiobooks170 hours of aligned audiobooks taken from tatkniga.ru. There are 4 speakers with 17+ hours of audio and 20 speakers in total. All the books are in free access and most of them in public domain.
sova_rudevices_audiobooks
Dataset instance structure
{'audio': {'path': '/path/to/wav.wav',
'array': array([wav numpy array]), dtype=float32),
'sampling_rate': 16000},
'transcription': 'транскрипция'}
Dataset audio info
16000 Hz
wav
mono
Russian speech from audiobooks
Citation
@misc{sova2021rudevices,
author = {Zubarev, Egor and Moskalets, Timofey and SOVA.ai},
title = {SOVA RuDevices Dataset: free public STT/ASR dataset with manually annotated live speech},
publisher =… See the full description on the dataset page: https://huggingface.co/datasets/dangrebenkin/sova_rudevices_audiobooks.qwen3-tts-zh-audiobooksturkish-tts-audiobooks
Turkish TTS Audiobooks
Turkish read-speech corpus for text-to-speech training, built from Turkish
audiobook and spoken-article recordings by an automatic pipeline: VAD
segmentation → technical QC → acoustic event tagging → DNSMOS → speaker
embedding/consistency → double-pass Whisper ASR → text policy → leakage-free
splitting. Audio is 16 kHz mono lossless FLAC embedded in the Parquet shards.
The pipeline that produced it — every stage, every threshold, the export and
audit… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-tts-audiobooks.
