kalixlouiis/myanmar-linguistic-ambiguitie-001
๐ฒ๐ฒ Myanmar Linguistic Ambiguities - Dataset 001 A comprehensive corpus designed for the disambiguation and grammatical error correction of homophones and confusing particles in the Burmese language. ๐ Overview: Myanmar Linguistic Ambiguities Series The Myanmar Linguistic Ambiguities (MLA) project is an ongoing effort to build highly precise and contextually rich datasets targeting specific common grammatical and orthographic errors in Burmese (Myanmar Language)โฆ See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/myanmar-linguistic-ambiguitie-001.
๐ฒ๐ฒ Myanmar Linguistic Ambiguities - Dataset 001
A comprehensive corpus designed for the disambiguation and grammatical error correction of homophones and confusing particles in the Burmese language.
๐ Overview: Myanmar Linguistic Ambiguities Series
The Myanmar Linguistic Ambiguities (MLA) project is an ongoing effort to build highly precise and contextually rich datasets targeting specific common grammatical and orthographic errors in Burmese (Myanmar Language) writing. These errors often involve words that sound identical (homophones) but carry distinct grammatical functions, leading to significant challenges for downstream NLP tasks and Large Language Models (LLMs).
MLA-Dataset-001 focuses on the primary distinction between the two most frequently confused particles: `แแฒ` and `แแแบแธ`.
๐ฏ Dataset 001: The แแฒ vs. แแแบแธ Distinction
Dataset 001 provides 1600 unique data points meticulously labeled for contextual correctness, helping models discern the subtle grammatical roles of these two characters.
Key Distinction:
- `แแฒ`: Primarily used as a Question Particle at the end of a sentence (e.g., "Where are you going?"). It is also used in words meaning Verb (to fall, to exchange, to change) or as part of Fixed Noun/Adverbial phrases (e.g., Pearl, Repeatedly).
- `แแแบแธ`: Primarily used as an Inclusion Particle (Conjunction) meaning "also," "too," or "as well" (e.g., "I am going also."). It is also essential in fixed formal conjunctions (e.g.,
แแฑแฌแบแแแบแธแแฑแฌแแบแธ- "whether X or Y").
๐ Data Statistics
๐ ๏ธ Data Structure and Labeling
The dataset is provided in a straightforward CSV format, making it instantly usable for classification, sequence labeling, or fine-tuning Burmese LLMs.
Example Records (Conceptual)
โ๏ธ Data Creation Methodology
The MLA-Dataset-001 was manually constructed and labeled based on the official definitions provided in the Myanmar Language Commission's dictionary and standardized grammar rules.
- Rule Identification: All primary and secondary grammatical functions of
แแฒandแแแบแธwere identified and categorized into distinct rule types. - Synthetic Sentence Generation: Contextually rich Burmese sentences were synthetically generated by the project maintainers, ensuring wide linguistic coverage and variety in subjects, tenses, and sentence structures.
- Error Injection & Labeling: For the "Incorrect Use" categories, sentences were specifically created by injecting the incorrect homophone (
แแฒinstead ofแแแบแธ, or vice versa), and the truecorrect_labelwas applied. - Verification: All 1600 data points were double-checked to ensure consistency between the
rule_typeand thecorrect_label.
๐ Future Datasets in the MLA Series
This dataset is the first installment in the series, targeting common Burmese orthographic challenges. Future releases will include:
- MLA-Dataset-002 (Coming Soon): Focusing on the distinction between the numeral particle `แ` and the numeral `แแ แบ`.
- MLA-Dataset-003 (Coming Soon): Focusing on the verb and particle distinction between `แแป` and `แแผ`.
๐ค Creator & Attribution
This dataset was independently designed, manually labeled, and published by Khant Sint Heinn (Kalix Louis).
All annotation design, rule structuring, dataset organization, and quality review were carried out by the creator with the goal of supporting Burmese (Myanmar) Natural Language Processing research and improving language resources for the community.
If you use this dataset in research, applications, or derivative works, please respectfully credit:
Khant Sint Heinn (Kalix Louis) Creator of Myanmar Linguistic Ambiguities - Dataset 001
Suggested citation:
Khant Sint Heinn (Kalix Louis). Myanmar Linguistic Ambiguities - Dataset 001. Myanmar Linguistic Ambiguities (MLA) Dataset.
๐ License
This dataset is released under the Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) license.
Usage Rights:
You are free to:
- Share โ copy and redistribute the material in any medium or format.
- Adapt โ remix, transform, and build upon the material for any purpose, even commercially.
Terms:
- Attribution โ You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.
- ShareAlike โ If you remix, transform, or build upon the material, you must distribute your contributions under the same license as the original.
Please cite this repository (kalixlouiis/myanmar-linguistic-ambiguities-001) if you use this data in your research or applications.
About the Author
Khant Sint Heinn, working under the name Kalix Louis, is a Machine Learning Engineer focused on Natural Language Processing (NLP), data foundations, and open-source AI development. His work is centered on improving support for the Burmese (Myanmar) language in modern AI systems by building high-quality datasets, practical tools, and scalable infrastructure for language technology.
He is currently the Lead Developer at DatarrX, where he develops data pipelines, manages large-scale data collection workflows, and helps create open-source resources for researchers, developers, and organizations. His experience includes data engineering, web scripting, dataset curation, and building systems that support real-world machine learning applications.
Khant Sint Heinn is especially interested in advancing low-resource languages and making AI more accessible to underrepresented communities. Through his open-source contributions, he works to strengthen the Burmese (Myanmar) tech ecosystem and provide reliable building blocks for future language models, search systems, and intelligent applications.
His goal is simple: to turn limited language resources into practical opportunities through clean data, useful tools, and community-driven innovation.
Connect with the Author: GitHub | Hugging Face | Kaggle
