datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-sample-v1.bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/1infinity0/bhasaflow-khasi-english-parallel-sample-v1.Khasi-OmniVoice-TTS-Data
Khasi Omni Voice TTS Dataset
The Khasi Omni Voice dataset is a comprehensive, high-quality audio collection designed specifically for Text-to-Speech (TTS) research and model training in the Khasi language. It features nearly 50 hours of speech data targeting realistic, modern Khasi speech patterns, including natural code-switching.
Key Statistics
Total Duration: 49 hours, 53 minutes, 19.98 seconds
Total Samples: 18,874 distinct audio utterances
Language: Khasi… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi-OmniVoice-TTS-Data.Khasi_ASR_Dataset
Khasi ASR Dataset
The Khasi ASR Dataset is a large-scale speech recognition dataset for the Khasi language, an Indigenous language spoken primarily in Meghalaya, India. The dataset contains paired audio recordings and transcriptions designed for training and evaluating Automatic Speech Recognition (ASR) systems.
This dataset consists of 73,900 audio-transcription pairs with a total duration of approximately 101 hours, 19 minutes, and 54.36 seconds of speech data.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi_ASR_Dataset.
