datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phoneme-ctc-english-60h-balanced
Phoneme CTC — English 60h (Balanced & Normalized)
A cleaned, normalized and phoneme-balanced version of
bobboyms/phoneme-ctc-english-60h-noisy,
for training phoneme recognition models (CTC) — e.g. as the native acoustic
model behind pronunciation-feedback systems.
What's different from the source dataset
Label noise removed
Roman numerals dropped — eSpeak reads ii/iv/… as "Roman two/four",
producing labels that don't match the audio.
Non-English phonemes dropped… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/phoneme-ctc-english-60h-balanced.sinhala-ctc-111h
Sinhala ASR – Consolidated OpenSLR (SLR52)
Dataset Summary
This dataset is a consolidated and cleaned version of the Sinhala Automatic Speech Recognition (ASR) dataset from OpenSLR (SLR52).
The original OpenSLR release distributes the data across multiple subsets (0–9, a–f).
This repository merges all subsets into a single unified dataset containing approximately 111 hours of speech audio.
Dataset Description
Consolidation
All OpenSLR SLR52… See the full description on the dataset page: https://huggingface.co/datasets/IAmNotAnanth/sinhala-ctc-111h.
