datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Indonesian-ASR-11-Class-Dataset
Indonesian ASR 11-Class Dataset
Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts.
Dataset summary
Audio files: 104,500 WAV files
Real/human recordings: 104,368
Synthetic repair files: 132
Sentence classes: 11 Indonesian sentence categories
Canonical balanced sentence slots: 209 (11 categories × 19 retained slots)
Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs*
Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.tajik-classic-audiobooks
Tajik Classic Audiobooks
Chapter-level audiobooks of classic Tajik/Persian literature, narrated by a single fine-tuned
Chatterbox-multilingual Tajik TTS voice (synthetic). Text was cleaned to spoken form (numbers
spelled out, no digits/Latin) and segmented per chapter. Each book folder holds mp3/ch_NNN.mp3
plus the source chapters_index.json and QC report.
Books included: shakuri_khuroson, shakuri_panturkizm, sadi_guliston, sadi_buston, nizami_layli, nizami_makhzan_khusrav… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/tajik-classic-audiobooks.
