datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
myanmar_quran_parallel_dataset_human_vs_ai
Myanmar Quran Parallel Dataset: Human vs AI
This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses.
It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context.
Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.tipitaka_myanmar_translation_books
Myanmar Tipitaka Translation (60 Books)
This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga.
The texts have been converted into a clean, structured JSONL format, suitable for:
Natural Language Processing (NLP)
LLM Training & Fine-tuning
Digital Humanities Research
Dhamma Study Applications
📊 Dataset Statistics
Total Books: 60
Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.myanmar-english-pali-dictionary
Myanmar–English–Pali Dictionary
Dataset Summary
This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein).
It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary.
The dataset is intended for research and educational purposes, including but not limited to:
Natural Language Processing (NLP)
Machine Translation (MT)
Lexicography
Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.myanmar-cities-qa
Myanmar Cites Questions & Answers Dataset
This dataset is an ongoing project dedicated to compiling comprehensive information about various cities in Myanmar. It converts geographical and cultural data—including locations, brief histories, local products, and notable landmarks—into a conversational Question & Answering (Q&A) format.
Dataset Overview
Content: Information about cities in Myanmar (e.g., location, history, local economy, and culture).
Format: Q&A… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-cities-qa.myanmar_yes_affirmation_spoken_dataset
Myanmar Yes Affirmation Spoken Dataset
Creator: freococoLicense: CC0 1.0 (Public Domain)Recommended for: Hugging Face, LLM fine-tuning, NLP researchTested with: Gemini Pro 3.0, ChatGPT 5.0
📖 Dataset Description
This dataset contains spoken Burmese expressions that all convey the meaning of "yes" / affirmation. It includes variations across:
Formality levels (casual, polite, formal)
Speaker gender (male, female, unisex)
Contextual usage (friends, shopkeepers… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_yes_affirmation_spoken_dataset.
