datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
myanmar-fineweb-2-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Fineweb2 Dataset
A preprocessed subset of the Fineweb2 dataset containing only Myanmar language text, with consistent Unicode encoding.
Dataset Description
This dataset is derived from the Fineweb2 created by HuggingFaceFW. It contains only the Myanmar language portion of the original Fineweb2 dataset, with additional preprocessing to standardize text encoding.
Filtered and Removed… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-fineweb-2-dataset.huggingface_myanmar_english_translation
Cleaned & Sorted Myanmar-English Translation Dataset
This dataset is a cleaned, Unicode-normalized, and sorted version of the Myanmar (Burmese) subset from the massive FineTranslations dataset.
While the original dataset is excellent, Myanmar text on the web is often a mix of standard Unicode and the non-standard Zawgyi encoding. This repository fixes those encoding issues to provide a high-quality dataset for NLP tasks.
Key Improvements in this Version
Zawgyi… See the full description on the dataset page: https://huggingface.co/datasets/freococo/huggingface_myanmar_english_translation.Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.myanmar-written-corpus
Myanmar Written Corpus
The Myanmar Written Corpus is a comprehensive collection of high-quality, but not fully CLEAN, written Myanmar text, designed to address the lack of large-scale, openly accessible resources for Myanmar Natural Language Processing (NLP). It is tailored to support various tasks such as text-to-speech (TTS), automatic speech recognition (ASR), translation, text generation, and more.
This dataset serves as a critical resource for researchers and developers aiming… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-written-corpus.myanmar-c4-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar C4 Dataset
A preprocessed subset of the C4 dataset containing only Myanmar language text, with consistent Unicode encoding.
Dataset Description
This dataset is derived from the Colossal Clean Crawled Corpus (C4) created by AllenAI. It contains only the Myanmar language portion of the original C4 dataset, with additional preprocessing to standardize text encoding.
Preprocessing… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-c4-dataset.myanmar-culturax-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar CulturaX Dataset
A preprocessed subset of the CulturaX dataset containing only Myanmar language text, with consistent Unicode encoding.
Dataset Description
This dataset is derived from the uonlp/CulturaX created by "The University of Oregon NLP Group". It contains only the Myanmar language portion of the original CulturaX dataset, with additional preprocessing to standardize text encoding.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-culturax-dataset.myanmar_spoken_corpus
Myanmar Spoken Corpus (Version 1.0)
Overview
Myanmar Spoken Corpus is a high-quality, but not fully CLEAN, open dataset of spoken Myanmar sentences designed to support NLP and ASR applications. The dataset focuses on providing clean and structured spoken language data for advancing Myanmar language technology.
Dataset Statistics
Number of Rows:
Local Parquet file: 16,020,011 rows
Hugging Face Dataset Viewer: 15,728,640 rows
File Size: 1.78 GB (Parquet… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_spoken_corpus.myanmar_quran_parallel_dataset_human_vs_ai
Myanmar Quran Parallel Dataset: Human vs AI
This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses.
It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context.
Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.myanmar-cc100-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar CC100 Dataset
A preprocessed subset of the CC100 dataset containing only Myanmar language text, with consistent Unicode encoding.
Dataset Description
This dataset is derived from the statmt/cc100 created by "Statistical and Neural Machine Translation". It contains only the Myanmar language portion of the original CC100 dataset, with additional preprocessing to standardize text encoding.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-cc100-dataset.myanmar_qna_dataset
Myanmar QnA Dataset v7
Language: Burmese (Myanmar)Total Entries: 22,783 QnA pairsTotal Sentences: ~ 466,330(Counted using the Myanmar sentence-ending symbol "။")License: CC0 1.0 (Public Domain)
Description
This dataset contains Myanmar-language question-answer pairs (QnA) generated with the assistance of ChatGPT-5 for question crafting with English and Gemini 3.0 Pro for Myanmar QnA generation. It is intended for research, AI training, and educational purposes.
Each entry… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_qna_dataset.tipitaka_myanmar_translation_books
Myanmar Tipitaka Translation (60 Books)
This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga.
The texts have been converted into a clean, structured JSONL format, suitable for:
Natural Language Processing (NLP)
LLM Training & Fine-tuning
Digital Humanities Research
Dhamma Study Applications
📊 Dataset Statistics
Total Books: 60
Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.myawady-news-title-generation-dataset
Myawady News Title Generation Dataset 🇲🇲
This dataset contains over 67,000 cleaned article titles extracted from the Myanmar state-run media outlet Myawady News Portal, intended for use in news title generation, text classification, and Myanmar NLP research.
The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ).
🗂️ Dataset Overview
Name:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-news-title-generation-dataset.myanmar-english-pali-dictionary
Myanmar–English–Pali Dictionary
Dataset Summary
This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein).
It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary.
The dataset is intended for research and educational purposes, including but not limited to:
Natural Language Processing (NLP)
Machine Translation (MT)
Lexicography
Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.Myanmar-Written-Spoken-Parallel-Corpus
Myanmar Written-Spoken Parallel Corpus (MWSPC)
Dataset Description
Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language.
Curated by: Khant Sint Heinn (Kalix Louis)
Organization: DatarrX | ဒေတာ-အက်စ်
Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.myawady-raw-dataset
Myawady Raw News Corpus 🇲🇲
This dataset contains over 59,000 full-text Burmese news articles scraped from the Myawady News Portal, the official media outlet of the Myanmar military government.
Unlike the title-only version, this dataset includes complete article content, with metadata fields such as category, publication date, and image URLs. It is intended for use in Myanmar NLP and AI research, including:
🧠 Language modeling
📰 Text summarization
🏷️ Named entity… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-raw-dataset.moi-myanmar-articles-lines
MOI Myanmar Articles Dataset - Lines (DatarrX/moi-myanmar-articles-lines)
Dataset Description
The MOI Myanmar Articles - Lines dataset is a derivative corpus created from the official articles published on the Ministry of Information (MOI) website of the Republic of the Union of Myanmar.
Unlike the main dataset (moi-myanmar-articles), which contains full-length article texts, this dataset has been systematically split line-by-line (sentence-by-sentence). This… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/moi-myanmar-articles-lines.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.myanmar-Wikipedia
Myanmar Wikipedia Dataset (20260501)
This dataset contains a cleaned, processed, and high-quality collection of Burmese Wikipedia articles, curated to serve as a robust foundation for Natural Language Processing (NLP) and Artificial Intelligence development in the Burmese language.
🏛️ About DatarrX
DatarrX (Burmese: ဒေတာအက်စ်) is a non-profit open-source foundation dedicated to building a robust digital foundation for the Burmese language in the AI era. We believe that… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-Wikipedia.moi-myanmar-articles
MOI Myanmar Articles Dataset (DatarrX/moi-myanmar-articles)
Dataset Description
The MOI Myanmar Articles dataset is a collection of official articles extracted directly from the Ministry of Information (MOI) website of the Republic of the Union of Myanmar. This dataset is curated exclusively to foster the growth, research, and development of the Myanmar (Burmese) language within the fields of Natural Language Processing (NLP) and Machine Learning (ML).… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/moi-myanmar-articles.Mpox-Myanmar
Mpox-Myanmar
Data Resources about Mpox(MonkeyPox) in Myanmar
Mpox-Myanmar is a dataset about Mpox(MonkeyPox virus) in Burmese Language.
Mpox(MonkeyPox) is becoming a wide alert virus. Thus, the information dataset about mpox will be built to build applications for knowledge and educate the public about mpox.
The dataset is gathered from the following web pages.
https://www.who.int/myanmar/emergencies/mpox
https://www.moi.gov.mm/article/60588
Questions are annotated by Min Si Thu.… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Mpox-Myanmar.jw_myanmar_bible_dataset
📖 JW Myanmar Bible Dataset (New World Translation)
A richly structured, fully aligned dataset of the Myanmar (Burmese) Bible, translated by Jehovah's Witnesses from the New World Translation. This dataset includes 66 books, 1,189 chapters, and 31,078 verses, each with chapter-level URLs and verse-level breakdowns.
✨ Highlights
- 📚 66 Canonical Books (Genesis to Revelation)
- 🧩 1,189 chapters, 31,078 verses (as parsed from the JW.org Myanmar edition)
- 🔗 Includes… See the full description on the dataset page: https://huggingface.co/datasets/freococo/jw_myanmar_bible_dataset.myX-myanmar-to-myanglish-corpus
📝 Myanglish (မြန်းဂလိ) corpus
A parallel text corpus containing high-quality, human-curated pairs of native Burmese (Myanmar Unicode) text and its corresponding Myanglish (Burmese Romanization) transliterations.
The initial high-quality release contains 2,121 fully approved rows designed to bridge the gap between formal script and the informal romanized phonetic typing structures widely used across social media, chat applications, and digital communication in Myanmar.… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-myanmar-to-myanglish-corpus.Quran_English_Myanmar_Parrelel_Corpus
Quran English-Myanmar Parallel Corpus
Description
This dataset is a parallel corpus of the Quran, containing translations in English and Myanmar. It includes 6,237 verses (ayahs) from all chapters (surahs), aligned by their respective Surah and Ayah numbers.
English Translation: Provided by Dr. Muhsin Khan and Dr. Hilali.
Myanmar Translation: Translated by the Myanmar Quran Translation Committee, comprising religious and non-religious scholars, and later published by… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/Quran_English_Myanmar_Parrelel_Corpus.myanmar-wikipedia-dataset
Myanmar Wikipedia Dataset (Last Crawl Date: 25/03/2025)
A collection of scraped Myanmar Wikipedia pages organized by category paths.
Overview
This dataset contains Myanmar Wikipedia articles scraped based on categorical organization. Unlike the official Wikimedia dataset (subset: 20231101.my), this repository provides an alternative approach to Myanmar Wikipedia content by following the categorical structure starting from the main entry page.
Figure 1: The initial… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-wikipedia-dataset.myanmar-aya-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Aya Dataset
A preprocessed subset of the Aya dataset containing only Myanmar language text.
Dataset Description
This dataset is derived from the Aya Dataset created by Cohere Labs. It contains only the Myanmar language portion of the original aya_dataset.
Dataset Structure
The dataset keep the same fields as the original aya_dataset dataset.
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-aya-dataset.myanmar-written-corpus
Myanmar Written Corpus
The Myanmar Written Corpus is a comprehensive collection of high-quality, but not fully CLEAN, written Myanmar text, designed to address the lack of large-scale, openly accessible resources for Myanmar Natural Language Processing (NLP). It is tailored to support various tasks such as text-to-speech (TTS), automatic speech recognition (ASR), translation, text generation, and more.
This dataset serves as a critical resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/SThandarTint/myanmar-written-corpus.myanmar-general-numerals-corpus
Myanmar General Numerals Corpus
This dataset is a collection of Myanmar (Burmese) sentences specifically curated to include various numerals, numerical classifiers, and units of measurement. It covers a wide range of linguistic styles, from daily conversations to formal and royal usage.
Dataset Details
Creator: Kalix Louis (Khant Sint Heinn)
Language: Myanmar (Burmese)
Format: Plain Text (.txt)
License: Apache-2.0
Source: Manually authored and curated by Kalix Louis.… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/myanmar-general-numerals-corpus.SO-Python_QA-API_Usage-tanh_score
Stack Overflow Python Q&A Dataset
Description
Filtered Python Q&A with API_Usage subcategory without:
Images
Links
Blocks of code
Scores in Q1-Q3 scaled with MaxAbsScaler. Tanh function applyed to joint Scores.
myanmar-cities-qa
Myanmar Cites Questions & Answers Dataset
This dataset is an ongoing project dedicated to compiling comprehensive information about various cities in Myanmar. It converts geographical and cultural data—including locations, brief histories, local products, and notable landmarks—into a conversational Question & Answering (Q&A) format.
Dataset Overview
Content: Information about cities in Myanmar (e.g., location, history, local economy, and culture).
Format: Q&A… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-cities-qa.1_pattern_10Kplus_myanmar_sentences
🧠 1_pattern_10Kplus_myanmar_sentences
A structured dataset of 11,452 Myanmar sentences generated from a single, powerful grammar pattern:
📌 Pattern:
Verb လည်း Verb တယ်။
A natural way to express repetition, emphasis, or causal connection in Myanmar.
💡 About the Dataset
This dataset demonstrates how applying just one syntactic pattern to a curated verb list — combined with syllable-aware rules — can produce a high-quality corpus of over 10,000 valid… See the full description on the dataset page: https://huggingface.co/datasets/freococo/1_pattern_10Kplus_myanmar_sentences.
