datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
myanmar-written-corpus
Myanmar Written Corpus
The Myanmar Written Corpus is a comprehensive collection of high-quality, but not fully CLEAN, written Myanmar text, designed to address the lack of large-scale, openly accessible resources for Myanmar Natural Language Processing (NLP). It is tailored to support various tasks such as text-to-speech (TTS), automatic speech recognition (ASR), translation, text generation, and more.
This dataset serves as a critical resource for researchers and developers aiming… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-written-corpus.myanmar_spoken_corpus
Myanmar Spoken Corpus (Version 1.0)
Overview
Myanmar Spoken Corpus is a high-quality, but not fully CLEAN, open dataset of spoken Myanmar sentences designed to support NLP and ASR applications. The dataset focuses on providing clean and structured spoken language data for advancing Myanmar language technology.
Dataset Statistics
Number of Rows:
Local Parquet file: 16,020,011 rows
Hugging Face Dataset Viewer: 15,728,640 rows
File Size: 1.78 GB (Parquet… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_spoken_corpus.Myanmar-Written-Spoken-Parallel-Corpus
Myanmar Written-Spoken Parallel Corpus (MWSPC)
Dataset Description
Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language.
Curated by: Khant Sint Heinn (Kalix Louis)
Organization: DatarrX | ဒေတာ-အက်စ်
Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.myX-myanmar-to-myanglish-corpus
📝 Myanglish (မြန်းဂလိ) corpus
A parallel text corpus containing high-quality, human-curated pairs of native Burmese (Myanmar Unicode) text and its corresponding Myanglish (Burmese Romanization) transliterations.
The initial high-quality release contains 2,121 fully approved rows designed to bridge the gap between formal script and the informal romanized phonetic typing structures widely used across social media, chat applications, and digital communication in Myanmar.… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-myanmar-to-myanglish-corpus.Quran_English_Myanmar_Parrelel_Corpus
Quran English-Myanmar Parallel Corpus
Description
This dataset is a parallel corpus of the Quran, containing translations in English and Myanmar. It includes 6,237 verses (ayahs) from all chapters (surahs), aligned by their respective Surah and Ayah numbers.
English Translation: Provided by Dr. Muhsin Khan and Dr. Hilali.
Myanmar Translation: Translated by the Myanmar Quran Translation Committee, comprising religious and non-religious scholars, and later published by… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/Quran_English_Myanmar_Parrelel_Corpus.myanmar-written-corpus
Myanmar Written Corpus
The Myanmar Written Corpus is a comprehensive collection of high-quality, but not fully CLEAN, written Myanmar text, designed to address the lack of large-scale, openly accessible resources for Myanmar Natural Language Processing (NLP). It is tailored to support various tasks such as text-to-speech (TTS), automatic speech recognition (ASR), translation, text generation, and more.
This dataset serves as a critical resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/SThandarTint/myanmar-written-corpus.myanmar-general-numerals-corpus
Myanmar General Numerals Corpus
This dataset is a collection of Myanmar (Burmese) sentences specifically curated to include various numerals, numerical classifiers, and units of measurement. It covers a wide range of linguistic styles, from daily conversations to formal and royal usage.
Dataset Details
Creator: Kalix Louis (Khant Sint Heinn)
Language: Myanmar (Burmese)
Format: Plain Text (.txt)
License: Apache-2.0
Source: Manually authored and curated by Kalix Louis.… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/myanmar-general-numerals-corpus.
