datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eps-burmese-qa
EPS Burmese Legal QA
Burmese-language question–answer pairs about Korean labour and immigration law, grounded in
the statutes themselves, for Myanmar workers on E-9/EPS visas in South Korea.
Korean employment law governs the daily life of hundreds of thousands of migrant workers.
The statutes exist in Korean and in official English translation. Almost none of it exists
in Burmese. This dataset was built to change that, and to make it possible to measure
whether a model answers… See the full description on the dataset page: https://huggingface.co/datasets/MYOTHANTZIN/eps-burmese-qa.burmese-corpusburmese-text-corpus
Burmese Text Corpus For Natural Language Processing
🎫 Choose your language: 🌏 English Version | 🇲🇲 မြန်မာဗားရှင်း
🌏 English Version
This dataset is a specifically curated text corpus for the Burmese language. It is intended to support Natural Language Processing (NLP) tasks, language model training, and research related to the Burmese language.
1. About the Dataset
The primary goal of creating this burmese-text-corpus dataset is to address the scarcity of… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/burmese-text-corpus.burmese-mbpp
Burmese MBPP: A Large-Scale Programming Dataset for Burmese Coding Assistants
Dataset Summary
The Burmese MBPP dataset is a translated and augmented version of the Google Mostly Basic Python Problems (MBPP) benchmark. It is designed to facilitate the training and evaluation of Large Language Models (LLMs) in generating Python code from Burmese natural language instructions.
This dataset contains 974 programming tasks, each featuring:
Burmese Instructions: Formal and… See the full description on the dataset page: https://huggingface.co/datasets/WYNN747/burmese-mbpp.quran-burmese-word-alignment
Quran Burmese Word Alignment Dataset
Creator: freococoLicense: CC BY-NC 4.0Language: Burmese (Myanmar), ArabicFormat: JSONL (one word per line)Current Version: v10 (Surah 1–114)
📖 Overview
This dataset provides a word-by-word alignment between a Burmese (Myanmar) translation of the Quran and the original Arabic Quranic text.
Each Burmese word is represented as a single JSON object and is optionally linked to one or more corresponding Arabic word(s), with explicit… See the full description on the dataset page: https://huggingface.co/datasets/freococo/quran-burmese-word-alignment.myX-Burmese-Morpho-Synthetic
myX-Burmese-Morpho-Synthetic
myX-Burmese-Morpho-Synthetic is a high-volume, synthetically augmented dataset consisting of over 37.8 million rows of Burmese word formations. Developed by Khant Sint Heinn (Kalix Louis) under the DatarrX organization, this resource is designed to advance the structural understanding of the Burmese language in the field of Natural Language Processing (NLP).
📌 Purpose
The primary goal of this dataset is to improve Burmese NLP by providing… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-Burmese-Morpho-Synthetic.burmese-contextual-pragmatics
Burmese Contextual Pragmatics Dataset
Created by freococo.
This dataset is a high-quality sociolinguistic resource for the Burmese (Myanmar) language. It provides a multi-dimensional mapping of 22 core conversational intents, showing how they transform across different social hierarchies, registers, and emotional contexts.
1. Overview & Licensing
Unlike simple phrasebooks, this dataset focuses on Pragmatics—how meaning changes based on social context, power dynamics, and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/burmese-contextual-pragmatics.Licensify-QA-Burmese
Licensify-QA-Burmese: An AI-Ready Instruction Dataset for Licensing & Compliance in Myanmar
Licensify-QA-Burmese is a specialized instruction-tuning dataset designed to help Large Language Models (LLMs) understand, compare, and explain various software, dataset, and content licenses.
Overview
As AI developers, navigating the legal complexities of open-source and proprietary licenses is a daily challenge. This project aims to bridge that gap by providing… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Licensify-QA-Burmese.burmese-VOA
Dataset Card for Burmese VOA News Dataset
This dataset is a comprehensive collection of Burmese news articles crawled from Voice of America (VOA) Burmese. It is specifically curated and processed for Natural Language Processing (NLP) tasks, focusing on high-quality news content, including the "Science and Technology" category.
Dataset Summary
The Burmese VOA Dataset contains 270,546 rows of news articles. The data has been meticulously scraped and structured into a… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/burmese-VOA.Burmese-MultiTurn-Chat-Corpusburmese-text-corpus
Burmese Text Corpus For Natural Language Processing
🎫 Choose your language: 🌏 English Version | 🇲🇲 မြန်မာဗားရှင်း
🌏 English Version
This dataset is a specifically curated text corpus for the Burmese language. It is intended to support Natural Language Processing (NLP) tasks, language model training, and research related to the Burmese language.
1. About the Dataset
The primary goal of creating this burmese-text-corpus dataset is to address the scarcity of… See the full description on the dataset page: https://huggingface.co/datasets/EISETWYNE/burmese-text-corpus.supportive-burmese-boyfriend-conversations
Burmese Supportive Boyfriend AI Dataset
Unique burmese chat dialogues for fine-tuning a supportive romantic partner (Koko) persona.
Files
File
Format
Rows
supportive_boyfriend_conversations_dataset.csv
CSV (prompt, response, category)
1,518
supportive_boyfriend_conversations_dataset.jsonl
JSONL (chat messages)
1,518
Schema
CSV (supportive_boyfriend_conversations_dataset.csv)
Column
Description
prompt
User message (partner… See the full description on the dataset page: https://huggingface.co/datasets/khaingmyel/supportive-burmese-boyfriend-conversations.
