datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.myanmar_qna_dataset
Myanmar QnA Dataset v7
Language: Burmese (Myanmar)Total Entries: 22,783 QnA pairsTotal Sentences: ~ 466,330(Counted using the Myanmar sentence-ending symbol "။")License: CC0 1.0 (Public Domain)
Description
This dataset contains Myanmar-language question-answer pairs (QnA) generated with the assistance of ChatGPT-5 for question crafting with English and Gemini 3.0 Pro for Myanmar QnA generation. It is intended for research, AI training, and educational purposes.
Each entry… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_qna_dataset.mig-english-myanmar-translation
👨💻 English <--> Myanmar Translation Dataset
English and Myanmar (Burmese) နှစ်ဘာသာကို စုဆောင်းပေးထားသော dataset ဖြစ်ပါတယ်။
Machine Translation လုပ်ငန်းစဉ်မှာ တစ်ထောင့်တစ်နေရာက အထောက်အကူပြုနိုင်လိမ့်မယ်လို့ မျှော်လင့်မိပါတယ်။
Samples ပေါင်း 13911 ရှိပါတယ်။
How to use
from datasets import load_dataset
dataset = load_dataset("Ko-Yin-Maung/Eng2Mm-Translation")
output
DatasetDict({
train: Dataset({
features: ['input', 'output'],
num_rows: 12645
})… See the full description on the dataset page: https://huggingface.co/datasets/Ko-Yin-Maung/mig-english-myanmar-translation.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.Mpox-Myanmar
Mpox-Myanmar
Data Resources about Mpox(MonkeyPox) in Myanmar
Mpox-Myanmar is a dataset about Mpox(MonkeyPox virus) in Burmese Language.
Mpox(MonkeyPox) is becoming a wide alert virus. Thus, the information dataset about mpox will be built to build applications for knowledge and educate the public about mpox.
The dataset is gathered from the following web pages.
https://www.who.int/myanmar/emergencies/mpox
https://www.moi.gov.mm/article/60588
Questions are annotated by Min Si Thu.… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Mpox-Myanmar.myanmar_qna_dataset
Myanmar QnA Dataset v7
Language: Burmese (Myanmar)Total Entries: 22,783 QnA pairsTotal Sentences: ~ 466,330(Counted using the Myanmar sentence-ending symbol "။")License: CC0 1.0 (Public Domain)
Description
This dataset contains Myanmar-language question-answer pairs (QnA) generated with the assistance of ChatGPT-5 for question crafting with English and Gemini 3.0 Pro for Myanmar QnA generation. It is intended for research, AI training, and educational purposes.
Each entry… See the full description on the dataset page: https://huggingface.co/datasets/EISETWYNE/myanmar_qna_dataset.
