datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
english-khasi-parallel-corpusThe monolingual data used to generate this dataset was sourced from wikipedia, news websites and open source corpora and then parallelized to generate these synthetic parallel corpora.
Sources
news_and_wikipedia_corpus: News websites and Simple English Wikipedia
news_corpus: Khasi News Websites
small sentences: Sentences with 5 words or less sourced from the Samanantar English dataset
large sentences: Sentences with more than 5 words sourced from the Samanantar English dataset
Khasi-OCR-36K
Khasi-OCR-36K
Khasi-OCR-36K is a Vision-Language dataset designed for OCR, document understanding, and handwriting recognition in the Khasi language, with a smaller subset of English samples.
This specific version of the dataset has been pre-filtered and formatted strictly for Vision training (e.g., DeepSeek-VL/OCR). It contains only the Free OCR task, with conversations mapped to the strict <|User|> and <|Assistant|> token format. Images are natively embedded.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi-OCR-36K.bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-sample-v1.bhasaflow-khasi-monolingual-corpus-v1
BhasaFlow Khasi Monolingual Corpus v1
By Medharvix Systems Private Limited
Overview
A curated monolingual Khasi text corpus for language modeling, NLP research, and linguistic analysis, with a focus on preserving and digitizing low-resource languages of Northeast India.
Dataset Structure
Column
Description
khasi_sentence
Khasi language sentence
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-monolingual-corpus-v1.bhasaflow-khasi-english-parallel-corpus-v1
BhasaFlow Khasi-English Parallel Corpus v1
By Medharvix Systems Private Limited
Overview
A curated parallel corpus of Khasi-English sentence pairs designed for machine translation research and development, with a focus on low-resource language technology for Northeast India.
Dataset Structure
Column
Description
sentence_id
Unique sentence identifier
english_text
English sentence
khasi_text
Khasi translation
Usage
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-corpus-v1.FEM-Khasi-News-Monolingual-Corpus
Khasi Monolingual News Corpus (740K)
Project Attribution & Collaboration
This dataset was collected and curated as part of the research project titled "Financial Empowerment in Meghalaya: AI-Powered Multilingual E-Marketplace for Tribes."
This project is a collaborative research initiative conducted by:
National Law University (NLU) Meghalaya
Indian Institute of Information Technology (IIIT) Guwahati
Contributors:
This dataset is the result of a joint effort by the… See the full description on the dataset page: https://huggingface.co/datasets/Bapynshngain/FEM-Khasi-News-Monolingual-Corpus.khasi-more-rawEnglish-Khasi-Parallel-Corpus-v1data_source:
Web-scraped data manually vetted by me.
NIT Silchar’s "EnKhCorp1.0: An English–Khasi Corpus."
Tatoeba project.
samanantar : the largest publicly available parallel corpora collection for 11 indic languages
acknowledgments:
Special thanks to Ahlad, NIT Silchar for their "EnKhCorp1.0" dataset, IIT Madras and other institutions involved in the creation of Samanantar and to the contributors of the Tatoeba project.
bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/1infinity0/bhasaflow-khasi-english-parallel-sample-v1.khasi-mapped-vectorskhasi-essayskhasi-datasets
What is Khasi Language?
Location:
Primarily spoken in the northeastern Indian state of Meghalaya.
Also spoken in parts of Assam, Tripura, and Bangladesh.
Language Family:
Khasi is a member of the Austroasiatic language family.
Script:
Traditionally written using the Khasi script, which is a script created specifically for the Khasi language.
Culture and Identity:
The Khasi language is an integral part of the cultural identity of the… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/khasi-datasets.khasi-ner
Khasi Named Entity Recognition Corpus
A named-entity-annotated corpus for the Khasi language, an Austroasiatic
language spoken by approximately 1.6 million people, primarily in the
state of Meghalaya in north-east India. To the best of the authors'
knowledge, this is the first publicly described NER resource for Khasi.
The corpus was constructed as part of the PhD thesis "Machine Learning
based Named Entity Recognition for Khasi Language" by Ransly Hoojon
(North-Eastern Hill… See the full description on the dataset page: https://huggingface.co/datasets/RansMairoi/khasi-ner.khasi-raw-datamore-khasinew-khasi-23Khasi_Synthetic_ASR_Questions_FinalKhasi-OCR-21K
Khasi-OCR-21K
Khasi-OCR-21K is a curated Vision-Language dataset totaling 21,319 samples, specifically designed to train robust OCR models for the Khasi language. This version introduces a significant amount of high-quality real book data alongside synthetic samples to handle diverse document conditions.
Dataset Split
To ensure reliable model evaluation, the dataset is split into:
Training Set: ~20,000 samples.
Validation Set: 1,319 samples (Randomized with a 40/30/30… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi-OCR-21K.new-raw-khasiKhasi-OmniVoice-TTS-Data
Khasi Omni Voice TTS Dataset
The Khasi Omni Voice dataset is a comprehensive, high-quality audio collection designed specifically for Text-to-Speech (TTS) research and model training in the Khasi language. It features nearly 50 hours of speech data targeting realistic, modern Khasi speech patterns, including natural code-switching.
Key Statistics
Total Duration: 49 hours, 53 minutes, 19.98 seconds
Total Samples: 18,874 distinct audio utterances
Language: Khasi… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi-OmniVoice-TTS-Data.Khasi_Synthetic_ASR_Norm_Part_1Khasi_Questions_Part_1Khasi_Questions_Part_2Khasi_Questions_Part_3Khasi_ASR_Final_Combinedkhasi-instruction-response-v2
Khasi Instruction Response v2
The Khasi Instruction Response v2 dataset is a high-quality, curated collection of 77,810 instruction-response pairs designed to fine-tune Large Language Models (LLMs) for the Khasi language. This is an improved, expanded version of my previous v1 release, offering significantly higher data integrity and broader linguistic coverage.
It combines extensive cultural, literary, and translation-based Khasi data with high-reasoning capabilities from… See the full description on the dataset page: https://huggingface.co/datasets/toiar/khasi-instruction-response-v2.english-khasi-parallel-corpus-43kKhasi_Synthetic_ASR_Norm_Part_2khasi-instruction-response-v1
Dataset Card for Khasi Instruction-Response Dataset v1
Dataset Summary
Khasi Instruction-Response Dataset v1 is a curated collection of prompts and responses in the Khasi language, designed to support the fine-tuning of instruction-following language models. It includes tasks such as Q&A, translation, summarization, and culturally-grounded dialogue. The dataset reflects indigenous knowledge systems, cultural expressions, and educational content unique to the Khasi context… See the full description on the dataset page: https://huggingface.co/datasets/toiar/khasi-instruction-response-v1.khasi-english-Translation-corpus
