theusamaaslam/urdu-organic-collection
Urdu Organic Speech Collection Consolidated organic Urdu speech dataset containing 337,874 clips (~337.2 hours) in one unified repository: training (326,923 clips / ~323.5h), validation (5,609 clips / ~6.6h), and a locked test split (5,342 clips / ~7.1h). The corpus combines the previously published organic Urdu collection with the audited 208h organic Urdu corpus — a large, translator-independent organic dataset that was speaker-labeled and quality-audited before inclusion.… See the full description on the dataset page: https://huggingface.co/datasets/theusamaaslam/urdu-organic-collection.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face