datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bhasaflow-khasi-monolingual-corpus-v1
BhasaFlow Khasi Monolingual Corpus v1
By Medharvix Systems Private Limited
Overview
A curated monolingual Khasi text corpus for language modeling, NLP research, and linguistic analysis, with a focus on preserving and digitizing low-resource languages of Northeast India.
Dataset Structure
Column
Description
khasi_sentence
Khasi language sentence
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-monolingual-corpus-v1.FEM-Khasi-News-Monolingual-Corpus
Khasi Monolingual News Corpus (740K)
Project Attribution & Collaboration
This dataset was collected and curated as part of the research project titled "Financial Empowerment in Meghalaya: AI-Powered Multilingual E-Marketplace for Tribes."
This project is a collaborative research initiative conducted by:
National Law University (NLU) Meghalaya
Indian Institute of Information Technology (IIIT) Guwahati
Contributors:
This dataset is the result of a joint effort by the… See the full description on the dataset page: https://huggingface.co/datasets/Bapynshngain/FEM-Khasi-News-Monolingual-Corpus.khasi-datasets
What is Khasi Language?
Location:
Primarily spoken in the northeastern Indian state of Meghalaya.
Also spoken in parts of Assam, Tripura, and Bangladesh.
Language Family:
Khasi is a member of the Austroasiatic language family.
Script:
Traditionally written using the Khasi script, which is a script created specifically for the Khasi language.
Culture and Identity:
The Khasi language is an integral part of the cultural identity of the… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/khasi-datasets.
