datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.everyayah-wav
everyayah-wav — Quranic recitation audio mirror
Full-mushaf Quranic recitation audio at 16 kHz mono 16-bit WAV, re-encoded
from everyayah.com for ML / ASR research.
This dataset is intentionally audio-only — no transcription text and no
alignment timings. The canonical Quranic text is widely available from
Tanzil and other public sources; pair this audio with
whatever text edition fits your use case.
Schema
Column
Type
Notes
audio
Audio(16000)
16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/dev-ahmedhany/everyayah-wav.Urdu-ONYX-WAV-kanade-V2
Urdu-ONYX-WAV-kanade-Annotated-V2
Version 2.0 - Artifact-Free Edition 🎉
Overview
This is an improved version of the Urdu-ONYX-WAV dataset, tokenized with the Kanade neural codec and optimized for artifact-free audio decoding. This dataset contains 143,627 samples of high-quality Urdu speech with comprehensive linguistic and acoustic annotations, totaling ~244 hours (~10 days) of continuous audio.
Key Features
🎯 Large-Scale: 143K+ samples, 244+ hours of… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-V2.Urdu-ONYX-WAV-Annoted
Urdu ONYX WAV Annotated Dataset
This dataset is an enhanced version of humair025/Urdu-ONYX-WAV-real with added phoneme annotations using the urdu-g2p library.
Features
Column
Description
id
Sample ID
transcript
Original Urdu transcript
voice
Voice type (onyx)
text
Text content
timestamp
Recording timestamp
audio
Audio file (WAV format)
phonemes
Space-separated IPA phonemes with stress markers
phonemes_no_stress
Space-separated IPA phonemes without… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-Annoted.
