datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
myanmar_quran_parallel_dataset_human_vs_ai
Myanmar Quran Parallel Dataset: Human vs AI
This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses.
It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context.
Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.myanimelist-embeddings
myanimelist-embeddings
This dataset is every non-empty anime synopsis from MyAnimeList.net ran
through the embed-multilingual-v2.0 embedding model from Cohere AI.
Sample code for searching for anime
Install some dependencies
pip install cohere==4.4.1 datasets==2.12.0 torch==2.0.1
Code heavily inspired by the Cohere Wikipedia embeddings sample
import os
import cohere
import torch
from datasets import load_dataset
co = cohere.Client(
os.environ["COHERE_API_KEY"]
) #… See the full description on the dataset page: https://huggingface.co/datasets/abatilo/myanimelist-embeddings.myanmar_idioms_lexicon
Myanmar Idioms Lexicon (မြန်မာဆိုရိုးစကား)
Dataset Description
The Myanmar Idioms Lexicon is a high-quality, linguistically enriched collection of traditional Burmese idioms (ဆိုရိုးစကား). While proverbs (စကားပုံ) often function as metaphorical moral allegories, Myanmar idioms (ဆိုရိုးစကား) are traditional sayings that describe social norms, biological observations, technical craftsmanship, and historical wisdom.
This dataset provides a comprehensive resource for NLP… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_idioms_lexicon.tipitaka_myanmar_translation_books
Myanmar Tipitaka Translation (60 Books)
This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga.
The texts have been converted into a clean, structured JSONL format, suitable for:
Natural Language Processing (NLP)
LLM Training & Fine-tuning
Digital Humanities Research
Dhamma Study Applications
📊 Dataset Statistics
Total Books: 60
Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.myanimelist-recommendations
myanimelist-recommendations
This is a scraped dataset taken from myanimelist.net's "Recommendations" feature. The top ~4,000 anime by popularity are included.
myanmar-english-pali-dictionary
Myanmar–English–Pali Dictionary
Dataset Summary
This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein).
It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary.
The dataset is intended for research and educational purposes, including but not limited to:
Natural Language Processing (NLP)
Machine Translation (MT)
Lexicography
Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.CryptoNewspali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.arabic_names_with_myanmar_and_english_transliteration
Arabic Names with Myanmar and English Transliteration
This dataset provides a collection of 23,074 Arabic names (primarily Hadith narrators and Islamic figures) with their corresponding English and Myanmar (Burmese) transliterations.
It is designed to help with NLP tasks involving name entity recognition, transliteration, and translation between Arabic, English, and Myanmar.
Dataset Structure
The dataset contains the following fields:
name_id (int64): Unique… See the full description on the dataset page: https://huggingface.co/datasets/freococo/arabic_names_with_myanmar_and_english_transliteration.myanmar-sft-datasetStack-3.0-examples-50KSO-Python_QA-filtered-2023-tanh_score-after_2023_02SO dataset of pythontag data
Question filters:
images
links
code blocks
Q_Score > 0
Answer_count > 0
CreationDate > 2023-02-01
Answers filters:
images
links
code blocks
Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores
stack-2-9-tool-examplesmyanmar-disaster-social-media
Myanmar Disaster Social Media Dataset
A dataset of ~1,000 Burmese (Myanmar) disaster-related social media posts,
labeled into four actionable categories. Intended for training and evaluating
text-classification models for disaster-response triage and monitoring.
Data fields
text (string): the Burmese social media post
category (string): the label
Labels
Label
Meaning
Immediate_Rescue_Needed
Posts requesting urgent rescue / help… See the full description on the dataset page: https://huggingface.co/datasets/MinThu11/myanmar-disaster-social-media.vinaya-pitaka-pali-myanmar-parallel
Vinaya Pitaka: Pali-Myanmar Parallel Dataset
Description
This dataset provides a professionally aligned, paragraph-level parallel corpus of the Vinaya Pitaka (The Code of Monastic Discipline). It features the original Pali text (presented in Myanmar script) alongside its modern Myanmar translation.
The dataset covers all five major volumes of the Vinaya:
Pārājika (ပါရာဇိကပါဠိ / ပါရာဇိကဏ်)
Pācittiya (ပါစိတ္တိယပါဠိ / ပါစိတ်)
Mahāvagga (မဟာဝဂ္ဂပါဠိ / မဟာဝါ)
Cūḷavagga… See the full description on the dataset page: https://huggingface.co/datasets/freococo/vinaya-pitaka-pali-myanmar-parallel.myanmar_proverbs_lexicon
Myanmar Proverbs Lexicon - V5
Dataset Description
The Myanmar Proverbs Lexicon is a comprehensive, carefully curated collection of 867 traditional Burmese proverbs, enriched with deep linguistic, cultural, and narrative context. This dataset is designed to preserve Myanmar's idiomatic wisdom while providing a high-quality resource for language learners, cultural researchers, and NLP practitioners.
A unique feature of this dataset is its dual-register approach: every… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_proverbs_lexicon.SO-Python_basics_QA-filtered-2023-T5_paraphrased-tanh_scoreSO-Python_QA-filtered-2023-tanh_scoreSO dataset of pythontag data
Question filters:
images
links
Q_Score > 0
Answer_count > 0
Answers filters:
images
links
code blocks
Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores
SO-Python_QA-filtered-2023-no_code-tanh_scoreSO dataset of pythontag data
Question filters:
images
links
code blocks
Q_Score > 0
Answer_count > 0
Answers filters:
images
links
code blocks
Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores
SO-Python_basics_QA-filtered-2023-tanh_scoreSO dataset of python tag data and "Python basics and Envirinment" subcategory
Question filters:
images
links
code blocks
Q_Score > 0
Answer_count > 0
Answers filters:
images
links
code blocks
Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores
SO_Python_basics_QA_human_prefContrastive dataset for Stack Overflow python basics QA with augmentations:
SO-SO comparisons: 6166
Par-SO comparisons: 0
SO-Par comparisons: 36366
Gen-SO comparisons: 0
SO-Gen comparisons: 87114
Gen-Par comparisons: 0
Par-Gen comparisons: 0
Gen-Gen comparisons: 0
Par-Par comparisons: 55494
Paraphrasing model: humarin/chatgpt_paraphraser_on_T5_base
my_arxivpali-words-myanmar-script
Pali Words in Myanmar Script (Master Index)
This dataset is a master index of 220,252 unique Pali words written in Myanmar (Burmese) Unicode script, intended for reuse across linguistic, religious, and computational workflows.
Data Fields
Each record contains:
word_id: A stable, sequential integer identifier.
pali_word: A Pali lexical item rendered in Myanmar Unicode script.
Data Processing Methodology
The dataset was constructed using the following steps:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/pali-words-myanmar-script.CryptoNews_50_50myanmar-llm-data
Myanmar LLM Data - chat-skill.md
A comprehensive Burmese (Myanmar) conversational dataset for training language models. This dataset contains multi-turn conversations covering various domains.
Skill Type: Chat/Skill
This dataset is part of the combined Myanmar LLM dataset collection:
chat-skill.md - Myanmar conversational data, translations, Q&A
agent-skill.md - amkyawdev/mm-llm-coder-agent-dataset
code-skill.md - amkyawdev/mm-llm-coder-dataset
Overview… See the full description on the dataset page: https://huggingface.co/datasets/amkyawdev/myanmar-llm-data.mon-eng-mya-pali-lexicon
🇲🇲 🇬🇧 Mon-Eng-Mya-Pali Lexicon Dataset for AI
This dataset is a comprehensive, multilingual lexicon of the Mon language (ဘာသာမန်), paired with English, Burmese (Myanmar), and Pali equivalents. It is structured specifically for Natural Language Processing (NLP), Large Language Model (LLM) fine-tuning, Machine Translation, and Retrieval-Augmented Generation (RAG) applications.
ဖိုင်ဒေတာဝေါဟာရ ကေုာံ အဘိဓာန်ဘာသာမန် (၄ ဘာသာ) သွက်ဂွံဗ္တောန် AI (Training) ကေုာံ စကာပ္ဍဲ AI… See the full description on the dataset page: https://huggingface.co/datasets/Nenemin95/mon-eng-mya-pali-lexicon.SO_Python_basics_QA_human_preferences_no_genmyanmar_11k_v8
Dataset Card for myanmar_11k_v8
Dataset Summary
ဒီ dataset က myanmar_11k_v8 အတွက် ဖန်တီးထားတာပါ။
Languages
Myanmar (my) / English (en)
Dataset Structure
Data Instances
{ "text": "နမူနာ စာသား", "label": "အညွှန်း" }
Data Fields
text: main content, label: optional.
Data Splits
Split
Files
train
myanmar_11k_v8.jsonl
Licensing Information
ဒီ dataset က CC BY-NC… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/myanmar_11k_v8.myanmar-literature-corpusstack-2-9-tool-20k-examples
