datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sinhala-tts-dataset
Sinhala TTS Dataset
Clean, segmented single-speaker Sinhala speech from the "Unlimited History" YouTube series by @sunchare. Built for TTS fine-tuning (F5-TTS, VITS, etc.).
Dataset Versions
cc_v1 — Full dataset (63 videos)
Metric
Value
Utterances
22,441
Train / Val
21,319 / 1,122
Hours
23.23h
Mean duration
3.73s
Duration range
3.0s – 19.74s
Sample rate
22,050 Hz
Videos processed
63
Avg keep rate
84.4%… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset.sinhala-tts-dataset-archive-20260429-082457
Sinhala TTS Dataset
Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare.
Stats
Metric
Value
Utterances
218
Train
208
Val
10
Hours
0.51
Mean duration
8.5s
Sample rate
22050 Hz
Pipeline
Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 ->
Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB)
Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.sinhala-bank-speech
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This dataset contains 100 audio files, one male voice in the format .wav.
The domain of this dataset is Banking.Only Language is Sinhalese(Sinhala,si)
Total Duration: 700.283 seconds.
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by… See the full description on the dataset page: https://huggingface.co/datasets/IshanSuga/sinhala-bank-speech.Sinhala_speech_to_textsinhala-ctc-111h
Sinhala ASR – Consolidated OpenSLR (SLR52)
Dataset Summary
This dataset is a consolidated and cleaned version of the Sinhala Automatic Speech Recognition (ASR) dataset from OpenSLR (SLR52).
The original OpenSLR release distributes the data across multiple subsets (0–9, a–f).
This repository merges all subsets into a single unified dataset containing approximately 111 hours of speech audio.
Dataset Description
Consolidation
All OpenSLR SLR52… See the full description on the dataset page: https://huggingface.co/datasets/IAmNotAnanth/sinhala-ctc-111h.
