zuhri025/urdu-eng-merged-dataset
--- language: - ur - en license: cc-by-4.0 task_categories: - automatic-speech-recognition tags: - urdu - english - speech - audio - asr - tts size_categories: - 10K<n<100K --- # Urdu + English merged speech dataset A merged dataset with: - **Urdu rows**: duration filter + normalization + cleaning - **English rows**: duration filter only, no text normalization ## Dataset Summary | Field | Value | |---|---| | **Samples** | 241,149 | | **Total audio** | 488.9 hours | |… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/urdu-eng-merged-dataset.
language:
- ur
- en license: cc-by-4.0 task_categories:
- automatic-speech-recognition tags:
- urdu
- english
- speech
- audio
- asr
- tts size_categories:
- 10K<n<100K ---
# Urdu + English merged speech dataset
A merged dataset with:
- Urdu rows: duration filter + normalization + cleaning
- English rows: duration filter only, no text normalization
## Dataset Summary
## Sources
- Urdu: zuhri025/munch-audio-NEW-parquet
- English: zuhri025/deep-orpheus-fixed
## Urdu text processing
Urdu rows were cleaned with:
- digit rejection
- Arabic/Urdu character normalization
- non-allowed characters replaced with spaces
- multiple spaces collapsed
- empty / too-short samples dropped
## English text processing
English rows were only filtered by duration. No text normalization was applied.
## Top 40 Characters
## License
Released under CC BY 4.0. Derived from source datasets; please cite them as well.
