datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sinhala-tts-dataset-archive-20260429-082457
Sinhala TTS Dataset
Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare.
Stats
Metric
Value
Utterances
218
Train
208
Val
10
Hours
0.51
Mean duration
8.5s
Sample rate
22050 Hz
Pipeline
Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 ->
Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB)
Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.Sinhala_speech_to_textsinhala-ctc-111h
Sinhala ASR – Consolidated OpenSLR (SLR52)
Dataset Summary
This dataset is a consolidated and cleaned version of the Sinhala Automatic Speech Recognition (ASR) dataset from OpenSLR (SLR52).
The original OpenSLR release distributes the data across multiple subsets (0–9, a–f).
This repository merges all subsets into a single unified dataset containing approximately 111 hours of speech audio.
Dataset Description
Consolidation
All OpenSLR SLR52… See the full description on the dataset page: https://huggingface.co/datasets/IAmNotAnanth/sinhala-ctc-111h.
