datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
myanmar-ocr-dataset-for-vlm
Myanmar OCR Dataset
A synthetic OCR dataset for fine-tuning Vision Language Models (VLMs) on Myanmar (Burmese) text recognition. It contains page images paired with their ground-truth text, sourced from chuuhtetnaing/mm-lib-book-dataset and rendered into page images using various Myanmar fonts.
Subsets
Subset
Description
Details
single_font
Rendered with Pyidaungsu font only
437 books
multi_font
Rendered with 76 Myanmar fonts
3 books (ပဋ္ဌာန်းမြတ်ဒေသနာ၊… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset-for-vlm.myanmar-ocr-dataset
Myanmar OCR Dataset
A synthetic dataset for training and fine-tuning Optical Character Recognition (OCR) models specifically for the Myanmar language.
Dataset Description
This dataset contains synthetically generated OCR images created specifically for Myanmar text recognition tasks. The images were generated using myanmar-ocr-data-generator, a fork of TextRecognitionDataGenerator with fixes for proper Myanmar character splitting.
Direct Download
Available… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset.nug_myanmar_asr
366 Hours NUG Myanmar ASR Dataset
The NUG Myanmar ASR Dataset is the first large-scale open Burmese speech dataset — now expanded to over 521,476 audio-text pairs, totaling ~366 hours of clean, segmented audio. All data was collected from public-service educational broadcasts by the National Unity Government (NUG) of Myanmar and the FOEIM Academy.
This dataset is released under a CC0 1.0 Universal license — fully open and public domain. No attribution required.
🕊️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/nug_myanmar_asr.ipfs_myanmar_laws_ir
Myanmar legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_myanmar_laws (revision b744df8f35b0f9eff35df4eeb57e3def96e80607) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Myanmar prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_myanmar_laws_ir.myanmar-fineweb-2-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Fineweb2 Dataset
A preprocessed subset of the Fineweb2 dataset containing only Myanmar language text, with consistent Unicode encoding.
Dataset Description
This dataset is derived from the Fineweb2 created by HuggingFaceFW. It contains only the Myanmar language portion of the original Fineweb2 dataset, with additional preprocessing to standardize text encoding.
Filtered and Removed… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-fineweb-2-dataset.voa_myanmar_voices
VOA Myanmar Voices
Burmese (Myanmar) speech corpus chunked into 20-second FLAC clips with transcripts. Derived from VOA Burmese radio broadcasts (public domain, U.S. 17 U.S.C. § 105).
Contents
File
Size
Description
voa-00000000.tar … voa-00000238.tar
498 GB
149 WebDataset shards
voa_transcripts.parquet
404 MB
1,424,257 (key, text) pairs
voa_transcripts.jsonl
1.5 GB
Same data, line-oriented
Audio: 16 kHz mono FLAC, exactly 20.00 s per chunk… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_voices.voa_myanmar_asr_audio_1
📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology.
Overview
This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.myanmar-speech-dataset-for-asrPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset for ASR
This dataset is a comprehensive collection of Myanmar language speech data specifically curated for Automatic Speech Recognition (ASR) task. It combines following datasets:
Myanmar Speech Dataset (Google Fleurs)
Myanmar Speech Dataset (OpenSLR-80)
Ko-Yin-Maung/mig-burmese-audio-transcription
By merging these complementary resources, this dataset provides a more robust… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-for-asr.myanmar_news
Dataset Card for Myanmar_News
Dataset Summary
The Myanmar news dataset contains article snippets in four categories:
Business, Entertainment, Politics, and Sport.
These were collected in October 2017 by Aye Hninn Khine
Languages
Myanmar/Burmese language
Dataset Structure
Data Fields
text - text from article
category - a topic: Business, Entertainment, Politic, or Sport (note spellings)
Data Splits
One training set (8,116 total… See the full description on the dataset page: https://huggingface.co/datasets/ayehninnkhine/myanmar_news.myanmar-shopvoice
Myanmar Shopvoice Dataset (300 Sample)
This is a randomly sampled subset of 300 Myanmar shop voice audio chunks for Whisper fine-tuning and evaluation.
Dataset Details
Total Audio Chunks: 300
Audio Format: 16 kHz WAV, mono
Language: Myanmar (Burmese)
Features
audio: Audio feature (16 kHz WAV audio player)
transcription: Myanmar sentence transcription text
source_audio: Source continuous recording file name
source_line: Index line of transcript… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/myanmar-shopvoice.115hours_pvtv_myanmar_asr
115 Hours PVTV Myanmar ASR
This dataset contains 156,262 audio-transcript pairs of spoken Burmese, totaling approximately 115.31 hours. The audio segments were extracted from publicly available YouTube videos published by PVTV and aligned using subtitle timestamps.
Dedication
This dataset would not exist without the persistent voices of PVTV editors, journalists, narrators, and production teams, who continue to speak to the people under difficult conditions. PVTV is the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/115hours_pvtv_myanmar_asr.huggingface_myanmar_english_translation
Cleaned & Sorted Myanmar-English Translation Dataset
This dataset is a cleaned, Unicode-normalized, and sorted version of the Myanmar (Burmese) subset from the massive FineTranslations dataset.
While the original dataset is excellent, Myanmar text on the web is often a mix of standard Unicode and the non-standard Zawgyi encoding. This repository fixes those encoding issues to provide a high-quality dataset for NLP tasks.
Key Improvements in this Version
Zawgyi… See the full description on the dataset page: https://huggingface.co/datasets/freococo/huggingface_myanmar_english_translation.myanmar-native-text
Dataset Card for myanmar-native-text
Dataset Summary
ဒီ dataset က myanmar-native-text အတွက် ဖန်တီးထားတာပါ။
Languages
Myanmar (my) / English (en)
Dataset Structure
Data Instances
{ "text": "နမူနာ စာသား", "label": "အညွှန်း" }
Data Fields
text: main content, label: optional.
Data Splits
Split
Files
train
data/native_00000.jsonl, data/native_00001.jsonl, data/native_00002.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/myanmar-native-text.thura-myanmar-news-mya-classificationref: https://huggingface.co/datasets/ThuraAung1601/myanmar_news
ipfs_myanmar_laws
Myanmar Laws
Legal corpus for Myanmar from the ipfs_datasets_py legal scrapers (endomorphosis).
Instruments: 725
Last updated: 2026-09-16
Source: official gazettes and legal portals; see source_url per record.
Schema: id, title, text, source_url, source_type, jurisdiction, country, language, eli, date, retrieved_at, license, law_status, identifiers.
myanmar-books-longform-1024
Dataset Card for myanmar-books-longform-1024
Dataset Summary
ဒီ dataset က myanmar-books-longform-1024 အတွက် ဖန်တီးထားတာပါ။
Languages
Myanmar (my) / English (en)
Dataset Structure
Data Instances
{ "text": "နမူနာ စာသား", "label": "အညွှန်း" }
Data Fields
text: main content, label: optional.
Data Splits
Split
Files
train
books_chunk_0000.jsonl, books_chunk_0001.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/myanmar-books-longform-1024.Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.myanmar-xnli
Dataset Card for myXNLI
Dataset Summary
The myXNLI corpus extends XNLI corpus with Myanmar (Burmese) language.
For myXNLI, we human-translated all 7,500 sentence pairs from XNLI English dev and test sets into Myanmar. The NLI and Genre labels from English dev and test sets are also reused for the Myanmar datasets.
The dataset also includes the NLI training data in Myanmar which is created by machine-translating the MultiNLI training data from English into Myanmar. Similar… See the full description on the dataset page: https://huggingface.co/datasets/akhtet/myanmar-xnli.myanmar_spoken_corpus
Credits and Acknowledgments
This dataset is built upon the foundational work of the Myanmar Spoken Corpus by freococo (Wynn).
Original Dataset Source: freococo/myanmar_spoken_corpus
Modifications: * Curated and filtered for specific training needs of the ShweYon model.
Integrated with 3% English subset for bilingual proficiency maintenance.
Re-formatted into 36 shards for optimized Continued Pre-training (CPT).
We are deeply grateful to freococo for their contribution to the… See the full description on the dataset page: https://huggingface.co/datasets/URajinda/myanmar_spoken_corpus.myanmar-speech-dataset-openslr-80Please visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset (OpenSLR-80)
This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual OpenSLR dataset.
For the complete multilingual dataset and additional information, please visit the original dataset repository
of OpenSLR HuggingFace page.
Original Source
OpenSLR is a site devoted to hosting speech and language resources, such as training… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-openslr-80.myanmar-burmese-speech-datasetMyanmar-English-general-text-translation
🇲🇲-🇬🇧 Myanmar-English General Text Translation Dataset
📚 Dataset Overview
This dataset is a high-quality, parallel corpus designed for training robust and accurate Myanmar-English Machine Translation (MT) models. It focuses on General Domain texts, covering a wide range of everyday scenarios, literature, conversations, and descriptive narratives.
Our primary goal in creating this dataset is to provide a clean, reliable resource to enhance the performance of… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/Myanmar-English-general-text-translation.myanmarnews-mya-clustering
myanmarnews-mya-clustering
Deduplicated copy of kornwtp/myanmarnews-mya-clustering,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/myanmarnews-mya-clustering
Deduplicated on: 2026-09-04
Task type: clustering
Splits: train
What changed
A text appearing under more than one gold cluster is collapsed to ONE row carrying the competing cluster ids in a new conflict_label column (NULL elsewhere). labels on that row holds the first-seen value as a… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/myanmarnews-mya-clustering.myanmar_typeset_dictionary_OCR
Myanmar Typeset Dictionary OCR Dataset
This is a synthetically generated, realistically formatted dataset modeling a Myanmar-Myanmar dictionary. It is designed for training and validating OCR models, Document Layout Analysis (DLA) pipelines, and structural key-value extraction models.
The dataset contains a highly diverse set of pages containing multiple column flows, tabular glossaries, running headers/footers, realistic backgrounds, and dynamic typography (four fonts paired… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_typeset_dictionary_OCR.myanmar-ner-datasetPlease visit the GitHub repository for other Myanmar Language datasets.
Myanmar NER Dataset
A token classification dataset for Myanmar (Burmese) Named Entity Recognition (NER), formatted for sequence labeling tasks using BIO tagging scheme.
📓 Data Preparation Notebook: dataset-preparation.ipynb
📓 Fine-Tuning Notebook: myanmar-ner-fine-tuning.ipynb (based on the HuggingFace Token Classification Guide)
Dataset Description
This dataset contains Myanmar language text… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ner-dataset.myanmar-speech-dataset-google-fleursPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset (Google Fleurs)
This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual Google Fleurs dataset.
For the complete multilingual dataset and additional information, please visit the original dataset repository
of Google Fleurs HuggingFace page.
Original Source
Fleurs is the speech version of the FLoRes machine translation benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-google-fleurs.khine-myanmarnews-mya-classification
KhineMyanmarNews_mya_Classification
Deduplicated copy of kornwtp/khine-myanmarnews-mya-classification.
Splits
split
rows
train
7,972
Myanmar-Motivation-Translation-Dataset
Motivational Speech Translations (English–Myanmar)
Dataset Description
This dataset contains sentence-level English–Myanmar translation pairs extracted from
transcribed motivational speech / video content. Each example is a short segment
(originally aligned to a timestamped subtitle line) paired with its Myanmar translation.
Languages: English (eng_Latn), Myanmar (mya_Mymr)
Total examples: 38,665
Source: Transcripts of motivational videos, segmented and… See the full description on the dataset page: https://huggingface.co/datasets/KyawSu/Myanmar-Motivation-Translation-Dataset.myanmar-written-corpus
Myanmar Written Corpus
The Myanmar Written Corpus is a comprehensive collection of high-quality, but not fully CLEAN, written Myanmar text, designed to address the lack of large-scale, openly accessible resources for Myanmar Natural Language Processing (NLP). It is tailored to support various tasks such as text-to-speech (TTS), automatic speech recognition (ASR), translation, text generation, and more.
This dataset serves as a critical resource for researchers and developers aiming… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-written-corpus.wikipedia-myanmar-qa
Dataset Card for wikipedia-myanmar-qa
Dataset Summary
ဒီ dataset က wikipedia-myanmar-qa အတွက် ဖန်တီးထားတာပါ။
Languages
Myanmar (my) / English (en)
Dataset Structure
Data Instances
{ "text": "နမူနာ စာသား", "label": "အညွှန်း" }
Data Fields
text: main content, label: optional.
Data Splits
Split
Files
train
qa_chunk_0000.jsonl, qa_chunk_0001.jsonl, qa_chunk_0002.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/wikipedia-myanmar-qa.
