CoolFace
Datasetpublicgated

theusamaaslam/urdu-organic-collection

Urdu Organic Speech Collection Consolidated organic Urdu speech dataset containing 337,874 clips (~337.2 hours) in one unified repository: training (326,923 clips / ~323.5h), validation (5,609 clips / ~6.6h), and a locked test split (5,342 clips / ~7.1h). The corpus combines the previously published organic Urdu collection with the audited 208h organic Urdu corpus — a large, translator-independent organic dataset that was speaker-labeled and quality-audited before inclusion.… See the full description on the dataset page: https://huggingface.co/datasets/theusamaaslam/urdu-organic-collection.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes15downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.