toiar/Khasi_ASR_Dataset
Khasi ASR Dataset The Khasi ASR Dataset is a large-scale speech recognition dataset for the Khasi language, an Indigenous language spoken primarily in Meghalaya, India. The dataset contains paired audio recordings and transcriptions designed for training and evaluating Automatic Speech Recognition (ASR) systems. This dataset consists of 73,900 audio-transcription pairs with a total duration of approximately 101 hours, 19 minutes, and 54.36 seconds of speech data. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi_ASR_Dataset.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face