CoolFace
Datasetpublicgated

toiar/Khasi_ASR_Dataset

Khasi ASR Dataset The Khasi ASR Dataset is a large-scale speech recognition dataset for the Khasi language, an Indigenous language spoken primarily in Meghalaya, India. The dataset contains paired audio recordings and transcriptions designed for training and evaluating Automatic Speech Recognition (ASR) systems. This dataset consists of 73,900 audio-transcription pairs with a total duration of approximately 101 hours, 19 minutes, and 54.36 seconds of speech data. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi_ASR_Dataset.

sourceHugging Facecc-by-nc-sa-4.0updated 3mo agoView on Hugging Face
0likes2downloads

No commit history came back for main. The revision may not exist, or the source declined the request.