datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bhasaflow-khasi-monolingual-corpus-v1
BhasaFlow Khasi Monolingual Corpus v1
By Medharvix Systems Private Limited
Overview
A curated monolingual Khasi text corpus for language modeling, NLP research, and linguistic analysis, with a focus on preserving and digitizing low-resource languages of Northeast India.
Dataset Structure
Column
Description
khasi_sentence
Khasi language sentence
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-monolingual-corpus-v1.FEM-Khasi-News-Monolingual-Corpus
Khasi Monolingual News Corpus (740K)
Project Attribution & Collaboration
This dataset was collected and curated as part of the research project titled "Financial Empowerment in Meghalaya: AI-Powered Multilingual E-Marketplace for Tribes."
This project is a collaborative research initiative conducted by:
National Law University (NLU) Meghalaya
Indian Institute of Information Technology (IIIT) Guwahati
Contributors:
This dataset is the result of a joint effort by the… See the full description on the dataset page: https://huggingface.co/datasets/Bapynshngain/FEM-Khasi-News-Monolingual-Corpus.khasi-datasets
What is Khasi Language?
Location:
Primarily spoken in the northeastern Indian state of Meghalaya.
Also spoken in parts of Assam, Tripura, and Bangladesh.
Language Family:
Khasi is a member of the Austroasiatic language family.
Script:
Traditionally written using the Khasi script, which is a script created specifically for the Khasi language.
Culture and Identity:
The Khasi language is an integral part of the cultural identity of the… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/khasi-datasets.khasi-instruction-response-v2
Khasi Instruction Response v2
The Khasi Instruction Response v2 dataset is a high-quality, curated collection of 77,810 instruction-response pairs designed to fine-tune Large Language Models (LLMs) for the Khasi language. This is an improved, expanded version of my previous v1 release, offering significantly higher data integrity and broader linguistic coverage.
It combines extensive cultural, literary, and translation-based Khasi data with high-reasoning capabilities from… See the full description on the dataset page: https://huggingface.co/datasets/toiar/khasi-instruction-response-v2.
