CoolFace
Datasetpublic

Shramadeepd/uyghur-ASR-dataset

Uyghur ASR Corpus (Latin Transliteration) A speech corpus for Uyghur automatic speech recognition, with transcriptions in a case-sensitive Latin transliteration scheme. Approximately 23 hours of audio across 9,468 clips. Uyghur is a Turkic language spoken by roughly 10–12 million people. It is severely under-represented in open speech datasets, and this corpus is intended to support ASR research for the language. Dataset summary Language Uyghur (ug)… See the full description on the dataset page: https://huggingface.co/datasets/Shramadeepd/uyghur-ASR-dataset.

sourceHugging Facecc-by-nc-4.0updated 8d agoView on Hugging Face
0likes99downloads

Shramadeepd/uyghur-ASR-dataset · main · files are served by the source, never re-hosted here