datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
myanmar_qna_dataset
Myanmar QnA Dataset v7
Language: Burmese (Myanmar)Total Entries: 22,783 QnA pairsTotal Sentences: ~ 466,330(Counted using the Myanmar sentence-ending symbol "။")License: CC0 1.0 (Public Domain)
Description
This dataset contains Myanmar-language question-answer pairs (QnA) generated with the assistance of ChatGPT-5 for question crafting with English and Gemini 3.0 Pro for Myanmar QnA generation. It is intended for research, AI training, and educational purposes.
Each entry… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_qna_dataset.jw_myanmar_bible_dataset
📖 JW Myanmar Bible Dataset (New World Translation)
A richly structured, fully aligned dataset of the Myanmar (Burmese) Bible, translated by Jehovah's Witnesses from the New World Translation. This dataset includes 66 books, 1,189 chapters, and 31,078 verses, each with chapter-level URLs and verse-level breakdowns.
✨ Highlights
- 📚 66 Canonical Books (Genesis to Revelation)
- 🧩 1,189 chapters, 31,078 verses (as parsed from the JW.org Myanmar edition)
- 🔗 Includes… See the full description on the dataset page: https://huggingface.co/datasets/freococo/jw_myanmar_bible_dataset.Quran_English_Myanmar_Parrelel_Corpus
Quran English-Myanmar Parallel Corpus
Description
This dataset is a parallel corpus of the Quran, containing translations in English and Myanmar. It includes 6,237 verses (ayahs) from all chapters (surahs), aligned by their respective Surah and Ayah numbers.
English Translation: Provided by Dr. Muhsin Khan and Dr. Hilali.
Myanmar Translation: Translated by the Myanmar Quran Translation Committee, comprising religious and non-religious scholars, and later published by… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/Quran_English_Myanmar_Parrelel_Corpus.myanmar-linguistic-ambiguitie-001
🇲🇲 Myanmar Linguistic Ambiguities - Dataset 001
A comprehensive corpus designed for the disambiguation and grammatical error correction of homophones and confusing particles in the Burmese language.
📚 Overview: Myanmar Linguistic Ambiguities Series
The Myanmar Linguistic Ambiguities (MLA) project is an ongoing effort to build highly precise and contextually rich datasets targeting specific common grammatical and orthographic errors in Burmese (Myanmar Language)… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/myanmar-linguistic-ambiguitie-001.myanmar_qna_dataset
Myanmar QnA Dataset v7
Language: Burmese (Myanmar)Total Entries: 22,783 QnA pairsTotal Sentences: ~ 466,330(Counted using the Myanmar sentence-ending symbol "။")License: CC0 1.0 (Public Domain)
Description
This dataset contains Myanmar-language question-answer pairs (QnA) generated with the assistance of ChatGPT-5 for question crafting with English and Gemini 3.0 Pro for Myanmar QnA generation. It is intended for research, AI training, and educational purposes.
Each entry… See the full description on the dataset page: https://huggingface.co/datasets/EISETWYNE/myanmar_qna_dataset.
