datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OSCAR-2019-Burmese-fix
Dataset Card for OSCAR-2019-Burmese-fix
Dataset Description
This dataset is a cleand version of Myanmar language in OSCAR 2019 dataset.
Contributions
Swan Htet Aung
eps-burmese-qa
EPS Burmese Legal QA
Burmese-language question–answer pairs about Korean labour and immigration law, grounded in
the statutes themselves, for Myanmar workers on E-9/EPS visas in South Korea.
Korean employment law governs the daily life of hundreds of thousands of migrant workers.
The statutes exist in Korean and in official English translation. Almost none of it exists
in Burmese. This dataset was built to change that, and to make it possible to measure
whether a model answers… See the full description on the dataset page: https://huggingface.co/datasets/MYOTHANTZIN/eps-burmese-qa.Burmese-Microbiology-1K
Burmese-Microbiology-1K
Min Si Thu, min@globalmagicko.com
Microbiology 1K QA pairs in Burmese Language
Purpose
Before this Burmese Clinical Microbiology 1K dataset, the open-source resources to train the Burmese Large Language Model in Medical fields were rare.
Thus, the high-quality dataset needs to be curated to cover medical knowledge for the development of LLM in the Burmese language
Motivation
I found an old notebook in my box. The book was… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Burmese-Microbiology-1K.burmese-text-corpus
Burmese Text Corpus For Natural Language Processing
🎫 Choose your language: 🌏 English Version | 🇲🇲 မြန်မာဗားရှင်း
🌏 English Version
This dataset is a specifically curated text corpus for the Burmese language. It is intended to support Natural Language Processing (NLP) tasks, language model training, and research related to the Burmese language.
1. About the Dataset
The primary goal of creating this burmese-text-corpus dataset is to address the scarcity of… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/burmese-text-corpus.burmese-mbpp
Burmese MBPP: A Large-Scale Programming Dataset for Burmese Coding Assistants
Dataset Summary
The Burmese MBPP dataset is a translated and augmented version of the Google Mostly Basic Python Problems (MBPP) benchmark. It is designed to facilitate the training and evaluation of Large Language Models (LLMs) in generating Python code from Burmese natural language instructions.
This dataset contains 974 programming tasks, each featuring:
Burmese Instructions: Formal and… See the full description on the dataset page: https://huggingface.co/datasets/WYNN747/burmese-mbpp.mini-burmesequran-burmese-word-alignment
Quran Burmese Word Alignment Dataset
Creator: freococoLicense: CC BY-NC 4.0Language: Burmese (Myanmar), ArabicFormat: JSONL (one word per line)Current Version: v10 (Surah 1–114)
📖 Overview
This dataset provides a word-by-word alignment between a Burmese (Myanmar) translation of the Quran and the original Arabic Quranic text.
Each Burmese word is represented as a single JSON object and is optionally linked to one or more corresponding Arabic word(s), with explicit… See the full description on the dataset page: https://huggingface.co/datasets/freococo/quran-burmese-word-alignment.burmese-only-wiki
Burmese Only Wiki Dataset
This dataset is a refined, high-precision monolingual collection derived from the DatarrX/myanmar-Wikipedia repository. It has been strictly filtered to ensure that every sentence contains exclusively Burmese characters, making it an ideal resource for language modeling, sequence-to-sequence tasks, and linguistic research where cross-lingual noise must be eliminated.
🏛️ About DatarrX
DatarrX (Burmese: ဒေတာအက်စ်) is a non-profit open-source… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/burmese-only-wiki.burmese-human-eval
Burmese HumanEval: A Benchmark for Evaluating Burmese Coding Assistants
Dataset Summary
The Burmese HumanEval dataset is an expansion and translation of the original OpenAI HumanEval benchmark, tailored specifically for evaluating the coding abilities of Large Language Models (LLMs) in the Burmese language. It consists of 100 programming problems designed to test functional correctness, language understanding, and problem-solving skills in Burmese contexts.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/WYNN747/burmese-human-eval.Burmese-English-Code-Mixed-Corpus
🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱
A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research.
Dataset Details
Organization: DatarrX
Creator: Khant Sint Heinn (Kalix Louis)
Number of Rows: 1,111
Language: Burmese (Unicode) & English Mix
Dataset Format: .txt
License: Apache 2.0
Description
The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-English-Code-Mixed-Corpus.general-burmese-sentences
🇲🇲 General Burmese Sentences
Dataset Summary
General Burmese Sentences is an open-source Burmese (Myanmar language) text corpus containing general-domain sentences written in natural Burmese. The dataset is intended to support research and development in Burmese Natural Language Processing (NLP), machine learning, and language model training.
This dataset is created and maintained by Khant Sint Heinn (Kalix Louis) under the Hugging Face repository:… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/general-burmese-sentences.myX-Burmese-Synthetic-Pseudo-Syllables
📝 Burmese Synthetic Pseudo-Syllables Dataset
This dataset contains 5,814,699 computer-generated (synthetic) Myanmar pseudo-syllables structured systematically based on specific complex linguistic and orthographic patterns.
Developed as part of the foundational research for low-resource language processing, this dataset serves as a rigorous baseline and stress-testing environment for Burmese Natural Language Processing (NLP), tokenization, font rendering, and spell-checking… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-Burmese-Synthetic-Pseudo-Syllables.Licensify-QA-Burmese
Licensify-QA-Burmese: An AI-Ready Instruction Dataset for Licensing & Compliance in Myanmar
Licensify-QA-Burmese is a specialized instruction-tuning dataset designed to help Large Language Models (LLMs) understand, compare, and explain various software, dataset, and content licenses.
Overview
As AI developers, navigating the legal complexities of open-source and proprietary licenses is a daily challenge. This project aims to bridge that gap by providing… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Licensify-QA-Burmese.Burmese-English-Code-Mixed-Corpus
🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱
A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research.
Dataset Details
Organization: DatarrX
Creator: Khant Sint Heinn (Kalix Louis)
Number of Rows: 1,111
Language: Burmese (Unicode) & English Mix
Dataset Format: .txt
License: Apache 2.0
Description
The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/hksamm/Burmese-English-Code-Mixed-Corpus.burmese-VOA
Dataset Card for Burmese VOA News Dataset
This dataset is a comprehensive collection of Burmese news articles crawled from Voice of America (VOA) Burmese. It is specifically curated and processed for Natural Language Processing (NLP) tasks, focusing on high-quality news content, including the "Science and Technology" category.
Dataset Summary
The Burmese VOA Dataset contains 270,546 rows of news articles. The data has been meticulously scraped and structured into a… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/burmese-VOA.Roleplay-Burmese
RolePlay-Burmese
Roleplay-Burmese Dataset is a dataset for roleplaying in Burmese language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other languages datasets for roleplay, it can be found at this github repo.… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Burmese.burmese-text-corpus
Burmese Text Corpus For Natural Language Processing
🎫 Choose your language: 🌏 English Version | 🇲🇲 မြန်မာဗားရှင်း
🌏 English Version
This dataset is a specifically curated text corpus for the Burmese language. It is intended to support Natural Language Processing (NLP) tasks, language model training, and research related to the Burmese language.
1. About the Dataset
The primary goal of creating this burmese-text-corpus dataset is to address the scarcity of… See the full description on the dataset page: https://huggingface.co/datasets/EISETWYNE/burmese-text-corpus.
