CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LianHong /zomi-monolingual-corpus Zomi Monolingual Corpus v1.0 The Zomi Monolingual Corpus v1.0 contains 363,401 cleaned, deduplicated, reviewed, and permission-approved Zomi sentences. Zomi is represented with the ISO 639-3 language code ctd (Tedim Chin). Quick start from datasets import load_dataset dataset = load_dataset("LianHong/zomi-monolingual-corpus", split="train") print(dataset.num_rows) # 363401 print(dataset[0]["zomi_text"]) Data fields Field Type Description… See the full description on the dataset page: https://huggingface.co/datasets/LianHong/zomi-monolingual-corpus.texttext-generation100K<n<1M0 likes426 downloads27d agoHugging Face02MWirelabs /assamese-monolingual-corpus Assamese Monolingual Corpus (2025) A high-quality, sentence-level Assamese monolingual dataset containing 1.61 million cleaned, segmented, and deduplicated sentences in Bengali script. This corpus supports Assamese NLP development, language modeling, and public deployment for Northeast India. Dataset Summary Language: Assamese (Bengali script) Size: 1,613,879 sentences Format: Plain text CSV (text column) Total tokens: 77,427,585 (using IndicBERTv2 tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/assamese-monolingual-corpus.texttext-generation1M<n<10M1 likes28 downloads10mo agoHugging Face03MEDHARVIX-SYSTEMS /bhasaflow-khasi-monolingual-corpus-v1 BhasaFlow Khasi Monolingual Corpus v1 By Medharvix Systems Private Limited Overview A curated monolingual Khasi text corpus for language modeling, NLP research, and linguistic analysis, with a focus on preserving and digitizing low-resource languages of Northeast India. Dataset Structure Column Description khasi_sentence Khasi language sentence Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-monolingual-corpus-v1.texttext-generationn<1K22 likes24 downloads5mo agoHugging Face04jintz0 /assamese-monolingual-corpus Assamese Monolingual Corpus (2025) A high-quality, sentence-level Assamese monolingual dataset containing 1.61 million cleaned, segmented, and deduplicated sentences in Bengali script. This corpus supports Assamese NLP development, language modeling, and public deployment for Northeast India. Dataset Summary Language: Assamese (Bengali script) Size: 1,613,879 sentences Format: Plain text CSV (text column) Total tokens: 77,427,585 (using IndicBERTv2 tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/jintz0/assamese-monolingual-corpus.texttext-generation1M<n<10M0 likes24 downloads4mo agoHugging Face05Bapynshngain /FEM-Khasi-News-Monolingual-Corpusgated Khasi Monolingual News Corpus (740K) Project Attribution & Collaboration This dataset was collected and curated as part of the research project titled "Financial Empowerment in Meghalaya: AI-Powered Multilingual E-Marketplace for Tribes." This project is a collaborative research initiative conducted by: National Law University (NLU) Meghalaya Indian Institute of Information Technology (IIIT) Guwahati Contributors: This dataset is the result of a joint effort by the… See the full description on the dataset page: https://huggingface.co/datasets/Bapynshngain/FEM-Khasi-News-Monolingual-Corpus.texttext-generation100K<n<1M0 likes22 downloads5mo agoHugging Face06LIACC /Emakhuwa-MonolingualBibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-building, title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks", author = "Ali, Felermino D. M. A. and Lopes Cardoso, Henrique and Sousa-Silva, Rui", editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen, Yun-Nung", booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Monolingual.tabulartranslation10K<n<100K0 likes11 downloads2y agoHugging Face07lecslab /monolingual_bentext1K<n<10K0 likes9 downloads5mo agoHugging Face08lecslab /monolingual_eustext1K<n<10K0 likes7 downloads5mo agoHugging Face09lecslab /monolingual_ckbtextn<1K0 likes6 downloads5mo agoHugging Face10lecslab /monolingual_myatextn<1K0 likes5 downloads5mo agoHugging Face11lecslab /monolingual_cebtext1K<n<10K0 likes4 downloads5mo agoHugging Face12lecslab /monolingual_dantextn<1K0 likes2 downloads5mo agoHugging Face13lecslab /monolingual_amhtext1K<n<10K0 likes1 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.