khasi
Datasets
All datasets matching “khasi”khasi-ocr-benchmark-2kenglish-khasi-parallel-corpusThe monolingual data used to generate this dataset was sourced from wikipedia, news websites and open source corpora and then parallelized to generate these synthetic parallel corpora.
Sources
news_and_wikipedia_corpus: News websites and Simple English Wikipedia
news_corpus: Khasi News Websites
small sentences: Sentences with 5 words or less sourced from the Samanantar English dataset
large sentences: Sentences with more than 5 words sourced from the Samanantar English dataset
Khasi-OCR-36K
Khasi-OCR-36K
Khasi-OCR-36K is a Vision-Language dataset designed for OCR, document understanding, and handwriting recognition in the Khasi language, with a smaller subset of English samples.
This specific version of the dataset has been pre-filtered and formatted strictly for Vision training (e.g., DeepSeek-VL/OCR). It contains only the Free OCR task, with conversations mapped to the strict <|User|> and <|Assistant|> token format. Images are natively embedded.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi-OCR-36K.bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-sample-v1.bhasaflow-khasi-monolingual-corpus-v1
BhasaFlow Khasi Monolingual Corpus v1
By Medharvix Systems Private Limited
Overview
A curated monolingual Khasi text corpus for language modeling, NLP research, and linguistic analysis, with a focus on preserving and digitizing low-resource languages of Northeast India.
Dataset Structure
Column
Description
khasi_sentence
Khasi language sentence
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-monolingual-corpus-v1.bhasaflow-khasi-english-parallel-corpus-v1
BhasaFlow Khasi-English Parallel Corpus v1
By Medharvix Systems Private Limited
Overview
A curated parallel corpus of Khasi-English sentence pairs designed for machine translation research and development, with a focus on low-resource language technology for Northeast India.
Dataset Structure
Column
Description
sentence_id
Unique sentence identifier
english_text
English sentence
khasi_text
Khasi translation
Usage
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-corpus-v1.
