datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Common_Voice_Corpus_22_0_Urdu
Common Voice Corpus 22.0 - Urdu
This dataset contains the Urdu subset of the Mozilla Common Voice 22.0 corpus, released in June 2025.It consists of crowdsourced speech recordings and their corresponding text transcriptions, collected to support open-source speech technology.
Dataset Summary
The Common Voice Corpus 22.0 Urdu dataset provides high-quality speech data for automatic speech recognition (ASR), speaker identification, and linguistic research in Urdu.It includes… See the full description on the dataset page: https://huggingface.co/datasets/azeem-ahmed/Common_Voice_Corpus_22_0_Urdu.munch_urdu_preview
🎧 Munch Preview Dataset
📖 Table of Contents
Dataset Description
Dataset Structure
Dataset Creation
Usage
Considerations
CitationContact
📋 Dataset Description
Overview
Munch Preview is a carefully curated preview dataset containing ** high-quality Urdu text-to-speech samples** from both versions of the Munch dataset family. This lightweight version allows researchers, developers, and practitioners to quickly explore and prototype with… See the full description on the dataset page: https://huggingface.co/datasets/humair025/munch_urdu_preview.Urdu-Munch-Lina
Urdu-Munch-Lina
Processed version of zuhri025/Urdu-Munch with LinaCodec encoding.
Dataset Structure
This dataset contains 2 batches of audio data encoded with LinaCodec:
Urdu-Munch-Lina/
├── .meta/ # Metadata (not data files)
│ ├── progress.json # Processing progress
│ └── metadata.json # Dataset metadata
├── batch_0/
│ ├── data-*.arrow # Actual data files
│ └── dataset_info.json
├── batch_1/
└── ...
Available batches: [1, 2]… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/Urdu-Munch-Lina.UrduSpeech
UrduSpeech
Dataset Summary
UrduSpeech is a high-quality multi-style speech corpus for Urdu and Kashmiri languages, designed for text-to-speech (TTS), automatic speech recognition (ASR), and expressive speech synthesis tasks. This dataset contains professionally recorded audio with diverse speaking styles, emotional expressions, and gender representation.
Dataset Composition
Languages: Urdu, Kashmiri
Total Samples: ~51.6K audio-text pairs
Train: 46.5K samples… See the full description on the dataset page: https://huggingface.co/datasets/humairawan/UrduSpeech.common-voice-urdu-processed
🎙️ Common Voice Urdu (Processed)
Ready-to-use Urdu speech dataset for fine-tuning ASR models
Mozilla Common Voice → Preprocessed → Whisper-Ready ✨
📊 Dataset at a Glance
Split
Samples
Use
🏋️ Train
7,339
Model training
🔧 Validation
5,046
Hyperparameter tuning
🧪 Test
5,091
Final evaluation
Total
17,476
💡 Audio is pre-resampled to 16kHz — plug directly into Whisper!
🚀 Quick Start
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/khawajaaliarshad/common-voice-urdu-processed.UrduMegaSpeech
UrduMegaSpeech-1M
Dataset Summary
UrduMegaSpeech-1M is a large-scale Urdu-English parallel speech corpus designed for automatic speech recognition (ASR), text-to-speech (TTS), and speech translation tasks. This dataset contains high-quality audio recordings paired with Urdu transcriptions and English source text, along with quality metrics for each sample.
Dataset Composition
Language: Urdu (transcriptions), English (source text)
Total Samples: ~1M+ audio-text… See the full description on the dataset page: https://huggingface.co/datasets/humair025/UrduMegaSpeech.urdu-tts-mini
Dataset Card for Urdu-TTS-Mini
A curated Urdu speech dataset for Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) research. Audio segments are extracted from publicly available YouTube speech content, processed through a multi-stage quality pipeline, and annotated with Urdu transcriptions. This is a mini release intended to validate the preprocessing pipeline and establish a quality baseline for future large-scale versions.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/salisai/urdu-tts-mini.Urdu-ONYX-WAV
Urdu-ONYX-WAV
Urdu-ONYX-WAV is a high-quality Urdu Text-to-Speech (TTS) dataset consisting of audio recordings and corresponding transcripts. This dataset has been specifically prepared for training TTS models and conducting research in Urdu speech synthesis.
📊 Dataset Structure
This dataset is distributed across multiple parts due to size constraints:
Main repository: Base dataset with initial samples
part2: Additional 2.56 GB of audio data (6 Arrow files)
part3:… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV.Urdu_Podcast_Audio_DatasetDataset Description:
This dataset is a large-scale collection of 2,297 hours of processed Urdu podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It captures real-world interactions across diverse topics and formats. The dataset preserves natural speech patterns, speaker variability, and authentic podcast environments, making it highly… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Urdu_Podcast_Audio_Dataset.Urdu-NSW
Urdu-NSW Dataset
Dataset Summary
The Urdu-NSW dataset is a high-quality collection of synthetic conversational Urdu audio samples paired with transcripts, designed for Text-to-Speech (TTS) model training, evaluation, and fine-tuning. The dataset contains approximately 46,594 audio clips generated using OpenAI's Audio API with the "amuch" voice model.
Each entry includes:
High-quality audio in WAV format (22,050 Hz, mono, 16-bit)
Natural, conversational Urdu text… See the full description on the dataset page: https://huggingface.co/datasets/humairawan/Urdu-NSW.urdu-organic-collection
Urdu Organic Speech Collection
Consolidated organic Urdu speech dataset containing 337,874 clips (~337.2 hours)
in one unified repository: training (326,923 clips / ~323.5h), validation
(5,609 clips / ~6.6h), and a locked test split (5,342 clips / ~7.1h).
The corpus combines the previously published organic Urdu collection with the
audited 208h organic Urdu corpus — a large, translator-independent organic
dataset that was speaker-labeled and quality-audited before inclusion.… See the full description on the dataset page: https://huggingface.co/datasets/theusamaaslam/urdu-organic-collection.Urdu_Podcast_Audio_Dataset_Dual_Channel
Dataset Description
This dataset is a large-scale collection of 2,297 hours of processed Urdu dual-channel podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It captures real-world podcast conversations across diverse topics and formats. The dataset is organized in a dual-channel format, where corresponding speaker audio… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Urdu_Podcast_Audio_Dataset_Dual_Channel.urdu-speech-samples
Urdu Speech Samples
This sample shows Urdu speech with transcript alignment and simple audio metadata. It is meant to help buyers review language fit and capture quality before requesting a larger sample or production delivery.
What This Shows
Urdu speech audio with paired text
A compact view of transcript and metadata structure
Audio format signals for procurement review
Dataset Specifications
Field
Value
Modality
Audio
Language
Urdu… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/urdu-speech-samples.Urdu-aud0
Urdu-aud0 Dataset
Dataset Summary
The Urdu-aud0 dataset is a high-quality collection of synthetic Urdu audio samples paired with transcripts, designed for Text-to-Speech (TTS) model training, evaluation, and fine-tuning. The dataset contains approximately 45,000 audio clips generated using OpenAI's Audio API with the "dan" voice model.
Each entry includes:
High-quality audio in WAV format (22,050 Hz, mono, 16-bit)
Corresponding Urdu text transcripts
Generation timestamps… See the full description on the dataset page: https://huggingface.co/datasets/humairawan/Urdu-aud0.Urdu-ONYX-WAV-kanade-V2
Urdu-ONYX-WAV-kanade-Annotated-V2
Version 2.0 - Artifact-Free Edition 🎉
Overview
This is an improved version of the Urdu-ONYX-WAV dataset, tokenized with the Kanade neural codec and optimized for artifact-free audio decoding. This dataset contains 143,627 samples of high-quality Urdu speech with comprehensive linguistic and acoustic annotations, totaling ~244 hours (~10 days) of continuous audio.
Key Features
🎯 Large-Scale: 143K+ samples, 244+ hours of… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-V2.Urdu-LjSpeech
Urdu-LjSpeech Dataset
Dataset Description
Urdu-LjSpeech is a high-quality Urdu speech dataset designed for Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) tasks. The dataset contains Urdu audio recordings paired with their corresponding text transcriptions.
Dataset Summary
Language: Urdu (اردو)
Format: Audio files with text transcriptions
Audio Specifications:
Sampling Rate: 22,050 Hz
Format: PCM 16-bit
Channels: Mono
Use Cases: Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/humairawan/Urdu-LjSpeech.
