humair025/UrduMegaSpeech
UrduMegaSpeech-1M Dataset Summary UrduMegaSpeech-1M is a large-scale Urdu-English parallel speech corpus designed for automatic speech recognition (ASR), text-to-speech (TTS), and speech translation tasks. This dataset contains high-quality audio recordings paired with Urdu transcriptions and English source text, along with quality metrics for each sample. Dataset Composition Language: Urdu (transcriptions), English (source text) Total Samples:… See the full description on the dataset page: https://huggingface.co/datasets/humair025/UrduMegaSpeech.
This repository belongs to humair025 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
