CoolFace
Datasetpublic

kalixlouiis/myanmar-linguistic-ambiguitie-001

๐Ÿ‡ฒ๐Ÿ‡ฒ Myanmar Linguistic Ambiguities - Dataset 001 A comprehensive corpus designed for the disambiguation and grammatical error correction of homophones and confusing particles in the Burmese language. ๐Ÿ“š Overview: Myanmar Linguistic Ambiguities Series The Myanmar Linguistic Ambiguities (MLA) project is an ongoing effort to build highly precise and contextually rich datasets targeting specific common grammatical and orthographic errors in Burmese (Myanmar Language)โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/myanmar-linguistic-ambiguitie-001.

sourceHugging Facecc-by-sa-4.0updated 5mo agoView on Hugging Face
6likes10downloads
Dataset Card

๐Ÿ‡ฒ๐Ÿ‡ฒ Myanmar Linguistic Ambiguities - Dataset 001

A comprehensive corpus designed for the disambiguation and grammatical error correction of homophones and confusing particles in the Burmese language.

๐Ÿ“š Overview: Myanmar Linguistic Ambiguities Series

The Myanmar Linguistic Ambiguities (MLA) project is an ongoing effort to build highly precise and contextually rich datasets targeting specific common grammatical and orthographic errors in Burmese (Myanmar Language) writing. These errors often involve words that sound identical (homophones) but carry distinct grammatical functions, leading to significant challenges for downstream NLP tasks and Large Language Models (LLMs).

MLA-Dataset-001 focuses on the primary distinction between the two most frequently confused particles: `แ€œแ€ฒ` and `แ€œแ€Šแ€บแ€ธ`.

๐ŸŽฏ Dataset 001: The แ€œแ€ฒ vs. แ€œแ€Šแ€บแ€ธ Distinction

Dataset 001 provides 1600 unique data points meticulously labeled for contextual correctness, helping models discern the subtle grammatical roles of these two characters.

Key Distinction:

  • โ€”`แ€œแ€ฒ`: Primarily used as a Question Particle at the end of a sentence (e.g., "Where are you going?"). It is also used in words meaning Verb (to fall, to exchange, to change) or as part of Fixed Noun/Adverbial phrases (e.g., Pearl, Repeatedly).
  • โ€”`แ€œแ€Šแ€บแ€ธ`: Primarily used as an Inclusion Particle (Conjunction) meaning "also," "too," or "as well" (e.g., "I am going also."). It is also essential in fixed formal conjunctions (e.g., แ€žแ€ฑแ€ฌแ€บแ€œแ€Šแ€บแ€ธแ€€แ€ฑแ€ฌแ€„แ€บแ€ธ - "whether X or Y").

๐Ÿ“Š Data Statistics

CategoryDescriptionData PointsCorrectness Ratio (0:1)
Correct Use (1)Contextually appropriate usage1000N/A
Incorrect Use (0)Grammatically incorrect usage600N/A
Total Corpus Size16005:3

๐Ÿ› ๏ธ Data Structure and Labeling

The dataset is provided in a straightforward CSV format, making it instantly usable for classification, sequence labeling, or fine-tuning Burmese LLMs.

Column NameData TypeDescription
sentence_idIntegerUnique identifier for the data point (1 to 1600).
textStringThe complete Burmese sentence containing the target_word. (This often includes grammatical errors for incorrect instances).
target_wordStringThe specific instance of 'แ€œแ€ฒ' or 'แ€œแ€Šแ€บแ€ธ' being evaluated in the sentence.
correct_labelStringThe grammatically correct form for the context (i.e., 'แ€œแ€ฒ' or 'แ€œแ€Šแ€บแ€ธ').
is_correct_in_contextBoolean (0 or 1)1 if target_word is correctly used; 0 if the usage is incorrect and should be corrected to correct_label.
rule_typeStringCategorization based on the grammatical function being tested (e.g., Question_Particle, Inclusion_Particle, Verb_Exchange, Conj(Incorrect).

Example Records (Conceptual)

`sentence_id``text``target_word``correct_label``is_correct_in_context``rule_type`
150แ€™แ€„แ€บแ€ธ แ€˜แ€šแ€บแ€žแ€ฝแ€ฌแ€ธแ€™แ€œแ€Šแ€บแ€ธแ‹แ€œแ€Šแ€บแ€ธแ€œแ€ฒ0Question_Particle (Incorrect)
528แ€™แ€„แ€บแ€ธแ€™แ€พแ€ฌ แ€•แ€…แ€นแ€…แ€Šแ€บแ€ธแ€›แ€พแ€ญแ€žแ€œแ€ญแ€ฏ แ€„แ€ซแ€ทแ€™แ€พแ€ฌแ€œแ€Šแ€บแ€ธ แ€•แ€…แ€นแ€…แ€Šแ€บแ€ธแ€›แ€พแ€ญแ€žแ€Šแ€บแ‹แ€œแ€Šแ€บแ€ธแ€œแ€Šแ€บแ€ธ1Inclusion_Particle
826แ€’แ€ฎแ€€แ€ฌแ€ธแ€›แ€ฒแ€ท แ€กแ€„แ€บแ€‚แ€ปแ€„แ€บแ€€แ€ญแ€ฏ แ€กแ€žแ€…แ€บแ€”แ€ฒแ€ท แ€œแ€ฒ แ€›แ€™แ€šแ€บแ‹แ€œแ€ฒแ€œแ€ฒ1Verb_Exchange
1401แ€•แ€Šแ€ฌแ€›แ€ฑแ€ธแ€แ€ฝแ€„แ€บแ€œแ€ฒแ€€แ€ฑแ€ฌแ€„แ€บแ€ธ แ€€แ€ปแ€”แ€บแ€ธแ€™แ€ฌแ€›แ€ฑแ€ธแ€แ€ฝแ€„แ€บแ€œแ€ฒแ€€แ€ฑแ€ฌแ€„แ€บแ€ธ แ€กแ€žแ€ฏแ€ถแ€ธแ€…แ€›แ€ญแ€แ€บแ€แ€ญแ€ฏแ€ธแ€™แ€ผแ€พแ€„แ€ทแ€บแ€›แ€™แ€Šแ€บแ‹แ€œแ€ฒแ€œแ€Šแ€บแ€ธ0Conj(Incorrect)

โœ๏ธ Data Creation Methodology

The MLA-Dataset-001 was manually constructed and labeled based on the official definitions provided in the Myanmar Language Commission's dictionary and standardized grammar rules.

  1. 1.Rule Identification: All primary and secondary grammatical functions of แ€œแ€ฒ and แ€œแ€Šแ€บแ€ธ were identified and categorized into distinct rule types.
  2. 2.Synthetic Sentence Generation: Contextually rich Burmese sentences were synthetically generated by the project maintainers, ensuring wide linguistic coverage and variety in subjects, tenses, and sentence structures.
  3. 3.Error Injection & Labeling: For the "Incorrect Use" categories, sentences were specifically created by injecting the incorrect homophone (แ€œแ€ฒ instead of แ€œแ€Šแ€บแ€ธ, or vice versa), and the true correct_label was applied.
  4. 4.Verification: All 1600 data points were double-checked to ensure consistency between the rule_type and the correct_label.

๐Ÿš€ Future Datasets in the MLA Series

This dataset is the first installment in the series, targeting common Burmese orthographic challenges. Future releases will include:

  • โ€”MLA-Dataset-002 (Coming Soon): Focusing on the distinction between the numeral particle `แ€` and the numeral `แ€แ€…แ€บ`.
  • โ€”MLA-Dataset-003 (Coming Soon): Focusing on the verb and particle distinction between `แ€€แ€ป` and `แ€€แ€ผ`.

๐Ÿ‘ค Creator & Attribution

This dataset was independently designed, manually labeled, and published by Khant Sint Heinn (Kalix Louis).

All annotation design, rule structuring, dataset organization, and quality review were carried out by the creator with the goal of supporting Burmese (Myanmar) Natural Language Processing research and improving language resources for the community.

If you use this dataset in research, applications, or derivative works, please respectfully credit:

Khant Sint Heinn (Kalix Louis) Creator of Myanmar Linguistic Ambiguities - Dataset 001

Suggested citation:

Khant Sint Heinn (Kalix Louis). Myanmar Linguistic Ambiguities - Dataset 001. Myanmar Linguistic Ambiguities (MLA) Dataset.

๐Ÿ“œ License

This dataset is released under the Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) license.

Usage Rights:

You are free to:

  • โ€”Share โ€” copy and redistribute the material in any medium or format.
  • โ€”Adapt โ€” remix, transform, and build upon the material for any purpose, even commercially.

Terms:

  • โ€”Attribution โ€” You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.
  • โ€”ShareAlike โ€” If you remix, transform, or build upon the material, you must distribute your contributions under the same license as the original.

Please cite this repository (kalixlouiis/myanmar-linguistic-ambiguities-001) if you use this data in your research or applications.

About the Author

Khant Sint Heinn, working under the name Kalix Louis, is a Machine Learning Engineer focused on Natural Language Processing (NLP), data foundations, and open-source AI development. His work is centered on improving support for the Burmese (Myanmar) language in modern AI systems by building high-quality datasets, practical tools, and scalable infrastructure for language technology.

He is currently the Lead Developer at DatarrX, where he develops data pipelines, manages large-scale data collection workflows, and helps create open-source resources for researchers, developers, and organizations. His experience includes data engineering, web scripting, dataset curation, and building systems that support real-world machine learning applications.

Khant Sint Heinn is especially interested in advancing low-resource languages and making AI more accessible to underrepresented communities. Through his open-source contributions, he works to strengthen the Burmese (Myanmar) tech ecosystem and provide reliable building blocks for future language models, search systems, and intelligent applications.

His goal is simple: to turn limited language resources into practical opportunities through clean data, useful tools, and community-driven innovation.

Connect with the Author: GitHub | Hugging Face | Kaggle