datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zomi-monolingual-corpus
Zomi Monolingual Corpus v1.0
The Zomi Monolingual Corpus v1.0 contains 363,401 cleaned, deduplicated,
reviewed, and permission-approved Zomi sentences. Zomi is represented with the
ISO 639-3 language code ctd (Tedim Chin).
Quick start
from datasets import load_dataset
dataset = load_dataset("LianHong/zomi-monolingual-corpus", split="train")
print(dataset.num_rows) # 363401
print(dataset[0]["zomi_text"])
Data fields
Field
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/LianHong/zomi-monolingual-corpus.assamese-monolingual-corpus
Assamese Monolingual Corpus (2025)
A high-quality, sentence-level Assamese monolingual dataset containing 1.61 million cleaned, segmented, and deduplicated sentences in Bengali script. This corpus supports Assamese NLP development, language modeling, and public deployment for Northeast India.
Dataset Summary
Language: Assamese (Bengali script)
Size: 1,613,879 sentences
Format: Plain text CSV (text column)
Total tokens: 77,427,585 (using IndicBERTv2 tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/assamese-monolingual-corpus.bhasaflow-khasi-monolingual-corpus-v1
BhasaFlow Khasi Monolingual Corpus v1
By Medharvix Systems Private Limited
Overview
A curated monolingual Khasi text corpus for language modeling, NLP research, and linguistic analysis, with a focus on preserving and digitizing low-resource languages of Northeast India.
Dataset Structure
Column
Description
khasi_sentence
Khasi language sentence
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-monolingual-corpus-v1.assamese-monolingual-corpus
Assamese Monolingual Corpus (2025)
A high-quality, sentence-level Assamese monolingual dataset containing 1.61 million cleaned, segmented, and deduplicated sentences in Bengali script. This corpus supports Assamese NLP development, language modeling, and public deployment for Northeast India.
Dataset Summary
Language: Assamese (Bengali script)
Size: 1,613,879 sentences
Format: Plain text CSV (text column)
Total tokens: 77,427,585 (using IndicBERTv2 tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/jintz0/assamese-monolingual-corpus.FEM-Khasi-News-Monolingual-Corpus
Khasi Monolingual News Corpus (740K)
Project Attribution & Collaboration
This dataset was collected and curated as part of the research project titled "Financial Empowerment in Meghalaya: AI-Powered Multilingual E-Marketplace for Tribes."
This project is a collaborative research initiative conducted by:
National Law University (NLU) Meghalaya
Indian Institute of Information Technology (IIIT) Guwahati
Contributors:
This dataset is the result of a joint effort by the… See the full description on the dataset page: https://huggingface.co/datasets/Bapynshngain/FEM-Khasi-News-Monolingual-Corpus.Emakhuwa-MonolingualBibTeX:
The dataset paper was published in EMNLP 2024.
Please cite as:
@inproceedings{ali-etal-2024-building,
title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks",
author = "Ali, Felermino D. M. A. and
Lopes Cardoso, Henrique and
Sousa-Silva, Rui",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Monolingual.monolingual_benmonolingual_eusmonolingual_ckbmonolingual_myamonolingual_cebmonolingual_danmonolingual_amh
