myanmar-corpus
myanmar_spoken_corpus
Credits and Acknowledgments
This dataset is built upon the foundational work of the Myanmar Spoken Corpus by freococo (Wynn).
Original Dataset Source: freococo/myanmar_spoken_corpus
Modifications: * Curated and filtered for specific training needs of the ShweYon model.
Integrated with 3% English subset for bilingual proficiency maintenance.
Re-formatted into 36 shards for optimized Continued Pre-training (CPT).
We are deeply grateful to freococo for their contribution to the… See the full description on the dataset page: https://huggingface.co/datasets/URajinda/myanmar_spoken_corpus.myanmar-written-corpus
Myanmar Written Corpus
The Myanmar Written Corpus is a comprehensive collection of high-quality, but not fully CLEAN, written Myanmar text, designed to address the lack of large-scale, openly accessible resources for Myanmar Natural Language Processing (NLP). It is tailored to support various tasks such as text-to-speech (TTS), automatic speech recognition (ASR), translation, text generation, and more.
This dataset serves as a critical resource for researchers and developers aiming… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-written-corpus.rakhine-myanmar-parallel-corpus
🌐 Rakhine – Myanmar Parallel Corpus
This repository contains a parallel dataset for Rakhine ↔ Myanmar machine translation and Natural Language Processing (NLP) research.
🎯 Purpose
This dataset is designed for:
Machine Translation (MT)
Language Modeling
NLP Research
Dialect Analysis (Rakhine and Standard Burmese)
Language Preservation
📌 Project Overview
The Rakhine language is a major language variety spoken in Rakhine State, Myanmar. However… See the full description on the dataset page: https://huggingface.co/datasets/rakhine-nlp/rakhine-myanmar-parallel-corpus.myanmar_spoken_corpus
Myanmar Spoken Corpus (Version 1.0)
Overview
Myanmar Spoken Corpus is a high-quality, but not fully CLEAN, open dataset of spoken Myanmar sentences designed to support NLP and ASR applications. The dataset focuses on providing clean and structured spoken language data for advancing Myanmar language technology.
Dataset Statistics
Number of Rows:
Local Parquet file: 16,020,011 rows
Hugging Face Dataset Viewer: 15,728,640 rows
File Size: 1.78 GB (Parquet… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_spoken_corpus.myanmar-literature-corpus
မြန်မာ စာပေ Corpus (MYLT Corpus) 🇲🇲📚
✨ ခြုံငုံဖော်ပြချက် (Overview)
ဤ မြန်မာ စာပေ Corpus (MYLT Corpus) သည် မြန်မာဘာသာစကားဖြင့် ရေးသားထားသော စာအုပ်နှင့် ဝတ္ထုရှည်များစွာမှ စာသားများကို စုစည်းသန့်စင်ထားသော အထွေထွေ Corpus ဖြစ်ပါသည်။ အဓိကအားဖြင့် မြန်မာ NLP (Natural Language Processing) နယ်ပယ် တိုးတက်ရေးအတွက်၊ အထူးသဖြင့် Language Model များ၊ Text Generation နှင့် Text Classification စသော လုပ်ငန်းများတွင် သုတေသနပြုရန် ရည်ရွယ်ပါသည်။
🎯 ရည်ရွယ်ချက်နှင့်… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/myanmar-literature-corpus.Myanmar-Written-Spoken-Parallel-Corpus
Myanmar Written-Spoken Parallel Corpus (MWSPC)
Dataset Description
Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language.
Curated by: Khant Sint Heinn (Kalix Louis)
Organization: DatarrX | ဒေတာ-အက်စ်
Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.
