CoolFace
Datasetpublic

DatarrX/myX-myanmar-orthography-corpus

📝 myX-myanmar-orthography-corpus: Myanmar Orthography Error Correction Dataset The myX-Myanmar-Orthography-Corpus is an open-source initiative dedicated to improving the accuracy of Myanmar (Burmese) language digital processing. This project focuses specifically on Orthographic Accuracy (သတ်ပုံ), addressing common spelling errors, phonetic confusions, and keyboard typos. This project focuses specifically on orthographic errors such as phonetic confusions, visual similarities… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-myanmar-orthography-corpus.

sourceHugging Facecc-by-sa-4.0updated 4mo agoView on Hugging Face
7likes26downloads
Dataset Card

📝 myX-myanmar-orthography-corpus: Myanmar Orthography Error Correction Dataset

The myX-Myanmar-Orthography-Corpus is an open-source initiative dedicated to improving the accuracy of Myanmar (Burmese) language digital processing. This project focuses specifically on Orthographic Accuracy (သတ်ပုံ), addressing common spelling errors, phonetic confusions, and keyboard typos.

This project focuses specifically on orthographic errors such as phonetic confusions, visual similarities, and keyboard typos rather than grammatical or syntactic changes. It aims to provide a high-quality "Noisy-to-Clean" parallel corpus for training and benchmarking NLP models.

🚧 Status: Work in Progress (Ongoing)

This dataset is currently a work-in-progress. We are actively expanding the corpus to cover a wider range of common Burmese spelling mistakes.

  • Current Milestone: 500+ high-quality pairs for the "Phet vs. Bet" (ဖက် နှင့် ဘက်) phonetic confusion have been curated and uploaded.
  • Future Updates: We plan to integrate more confusion pairs (e.g., မီ vs မှီ, ကျ vs ကြ) and common digital typing errors.

Dataset Structure

The data is provided in CSV format and follows a structured parallel schema:

Column NameDescription
noisy_sentenceThe sentence containing the orthographic error.
correct_sentenceThe ground truth (correct spelling) version of the sentence.
error_categoryThe type of error (e.g., Phonetic Confusion, Keyboard Typo).
target_wordsThe specific words being corrected in that row.

Origin & Contribution

This dataset is developed and maintained by DatarrX, a Myanmar-based open-source data organization dedicated to building robust linguistic resources for the Myanmar (Burmese) language.

📜 Licensing

This work is licensed under the Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0). You are free to share and adapt the material, provided you give appropriate credit and distribute your contributions under the same license.


🤝 Contribution

As this is an ongoing project, we welcome contributions, suggestions, and bug reports. Stay tuned for further updates as we expand the corpus size and diversity.


For inquiries or contributions to the dataset, please reach out via the DatarrX organization page.