freococo/tipitaka_myanmar_translation_books
Myanmar Tipitaka Translation (60 Books) This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga. The texts have been converted into a clean, structured JSONL format, suitable for: Natural Language Processing (NLP) LLM Training & Fine-tuning Digital Humanities Research Dhamma Study Applications ๐ Dataset Statistics Total Books: 60 Total Content Lines:โฆ See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.
Myanmar Tipitaka Translation (60 Books)
This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga.
The texts have been converted into a clean, structured JSONL format, suitable for:
- Natural Language Processing (NLP)
- LLM Training & Fine-tuning
- Digital Humanities Research
- Dhamma Study Applications
๐ Dataset Statistics
- Total Books: 60
- Total Content Lines: 194,730
- Language: Myanmar (Unicode)
- Format: JSONL
- Composition:
- 46 Books: Tipitaka (Canon)
- 14 Books: Commentaries & Treatises
โ๏ธ License & Philosophy
License: CC0 1.0 Universal (Public Domain)
I have chosen this universal license because the Dhamma taught by the Lord Buddha is a universal message for all human beings, regardless of race, nation, or creed. The Truth belongs to no one and everyone.
"The gift of Dhamma excels all other gifts." โ (Dhammapada, 354)
You are free to use this dataset for any purpose, including commercial use, without restriction.
๐ Credits, Acknowledgements & Dedication
Namo Tassa Bhagavato Arahato Sammasambuddhassa
This dataset is a digital tribute to the profound wisdom preserved by the Noble Order of the Sangha. All merit and credit for this work belong to the Great Teachers and the devoted community who have protected these teachings through the ages.
We offer our deepest homage and gratitude to:
- The Venerable Sayadaws and Mahatheras: The distinguished monks and learned scholars who, with great wisdom and compassion, translated the profound Pali Canon into the Myanmar language.
- The Custodians of the Texts: The dedicated monks and lay devotees who facilitated the Sixth Buddhist Council (Chattha Sangayana) and other historic councils to purify and preserve the texts.
- The Digital Dhammaduta (Volunteers): The countless unsung heroesโscanners, typists, proofreaders, and editorsโwho tirelessly converted physical palm-leaf manuscripts and printed books into digital text. Without their immense effort in typing and scanning, this dataset would not exist.
- The Open Source Maintainers: Special gratitude to [pndaza](https://github.com/pndaza) and contributors. Their technical dedication to maintaining the digital Tipitaka files provided the foundation for this JSONL conversion.
- Source Repository: https://github.com/pndaza
Dataset Creator:
- freococo: I am merely a humble language enthusiast who fell in love with the beauty of these texts. My only role was to convert these sacred teachings into a machine-readable format to assist in preservation and study. I claim no ownership over this wisdom; I am simply a bridge between the ancient texts and modern technology.
May this work contribute to the longevity of the Sasana.
๐ป How to Use & Filter
You can load the dataset using the Hugging Face datasets library. Since the Canon and Commentaries are labeled in the category field, you can easily filter them.
from datasets import load_dataset
# 1. Load the full dataset
dataset = load_dataset("freococo/tipitaka_myanmar_translation_books", split="train")
# 2. Filter for Canon only (The 46 Tipitaka Books)
# We exclude items where the category contains "แกแแนแแแแฌ" (Commentary)
canon_only = dataset.filter(lambda x: "แกแแนแแแแฌ" not in x['category'])
# 3. Filter for Commentaries only (The 14 Atthakatha Books)
commentaries = dataset.filter(lambda x: "แกแแนแแแแฌ" in x['category'])
# Print stats
print(f"Total rows: {len(dataset)}")
print(f"Canon rows: {len(canon_only)}")
print(f"Commentary rows: {len(commentaries)}")๐ Data Structure
Each row represents a paragraph or a heading.
{
"id": "01_vinaya_01_2674",
"book_id": "01_vinaya_01",
"book_name": "แแซแแฌแแญแแแบ",
"category": "แแญแแแแญแแ (แแฏแแนแแแแญแแฌแ)",
"chapter": "แ-แแถแแฌแแญแแญแแบ แกแแแบแธ",
"text": "แแญแฏ (แ
แฑแฌแแแแแแบแธ) แแแบ ''แแซแแฌแแญแแกแฌแแแบแแญแฏแท แแฑแฌแแบแแฒแแผแ
แบแแฑแฌ แแแแบแธแแญแฏ แแผแแบแกแแบแ..."
}๐ Full Book List (60 Items)
This dataset includes the following 60 books, processed and verified:
- 01_vinaya_01 - แแซแแฌแแญแแแบ (Parajika)
- 01_vinaya_02 - แแซแ แญแแบ (Pacittiya)
- 01_vinaya_03 - แแแฌแแซ (Mahavagga)
- 01_vinaya_04 - แ แฐแ แแซ (Cullavagga)
- 01_vinaya_05 - แแแญแแซ (Parivara)
- 02_digha_01 - แแฎแแแนแแแบ (Silakkhandha Vagga)
- 02_digha_02 - แแฏแแบแแแฌแแซ (Maha Vagga)
- 02_digha_03 - แแซแแญแ (Pathika Vagga)
- 03_majjhima_01 - แแฐแแแแนแแฌแ (Mulapannasa)
- 03_majjhima_02 - แแแนแแญแแแแนแแฌแ (Majjhimapannasa)
- 03_majjhima_03 - แฅแแแญแแแนแแฌแ (Uparipannasa)
- 04_sanyutta_01 - แแแซแแฌแแแนแแแถแแฏแแบ (Sagatha Vagga)
- 04_sanyutta_02 - แแญแแซแแแแนแแแถแแฏแแบ (Nidana Vagga)
- 04_sanyutta_03 - แแแนแแแแนแแแถแแฏแแบ (Khandha Vagga)
- 04_sanyutta_04 - แแ แฌแแแแแแนแแแถแแฏแแบ (Salayatana Vagga)
- 04_sanyutta_05 - แแแฌแแแนแแแถแแฏแแบ (Maha Vagga)
- 05_anguttara_01 - แงแแแแญแแซแแบ (Ekaka Nipata)
- 05_anguttara_02 - แแฏแแแญแแซแแบ (Duka Nipata)
- 05_anguttara_03 - แแญแแแญแแซแแบ (Tika Nipata)
- 05_anguttara_04 - แ แแฏแแนแแแญแแซแแบ (Catukka Nipata)
- 05_anguttara_05 - แแแนแ แแแญแแซแแบ (Pancaka Nipata)
- 05_anguttara_06 - แแแนแแแญแแซแแบ (Chakka Nipata)
- 05_anguttara_07 - แแแนแแแแญแแซแแบ (Sattaka Nipata)
- 05_anguttara_08 - แกแแนแแแแญแแซแแบ (Atthaka Nipata)
- 05_anguttara_09 - แแแแแญแแซแแบ (Navaka Nipata)
- 05_anguttara_10 - แแแแแญแแซแแบ (Dasaka Nipata)
- 05_anguttara_11 - แงแแฌแแแแแญแแซแแบ (Ekadasaka Nipata)
- 06_khuddaka_01 - แแฏแแนแแแแซแ (Khuddakapatha)
- 06_khuddaka_02 - แแแนแแแ (Dhammapada)
- 06_khuddaka_03 - แฅแแซแแบแธ (Udana)
- 06_khuddaka_04 - แฃแแญแแฏแแบ (Itivuttaka)
- 06_khuddaka_05 - แแฏแแนแแแญแแซแแบ (Suttanipata)
- 06_khuddaka_06 - แแญแแฌแแแแนแแฏ (Vimanavatthu)
- 06_khuddaka_07 - แแฑแแแแนแแฏ (Petavatthu)
- 06_khuddaka_08 - แแฑแแแซแแฌ (Theragatha)
- 06_khuddaka_09 - แแฑแแฎแแซแแฌ (Therigatha)
- 06_khuddaka_10 - แกแแแซแแบ (แ) [แแฑแ] (Apadana 1)
- 06_khuddaka_11 - แกแแแซแแบ (แแฏ) [แแฑแ+แแฑแแฎ] (Apadana 2)
- 06_khuddaka_12 - แแฏแแนแแแถแ (Buddhavamsa)
- 06_khuddaka_13 - แ แแญแแฌแแญแแ (Cariyapitaka)
- 06_khuddaka_18 - แแแญแแแนแแญแแซแแแบ (Patisambhidamagga)
- 06_khuddaka_19 - แแญแแญแแนแแแแพแฌ (Milindapanha)
- 07_abhidhamma_01 - แแแนแแแแบแนแแแฎ (Dhammasangani)
- 07_abhidhamma_02 - แแญแแแบแธ (Vibhanga)
- 07_abhidhamma_03 - แแฏแแนแแแแแแบ (Puggalapannatti)
- 07_abhidhamma_05 - แแแฌแแแนแแฏ (Kathavatthu)
- 08_dhammapada_01 - แแแนแแแแกแแนแแแแฌ-แ (Dhammapada Commentary 1)
- 08_dhammapada_02 - แแแนแแแแกแแนแแแแฌ-แแฏ (Dhammapada Commentary 2)
- 08_jataka_01 - แแฌแแแกแแนแแแแฌ-แ (Jataka Commentary 1)
- 08_jataka_02 - แแฌแแแกแแนแแแแฌ-แแฏ (Jataka Commentary 2)
- 08_jataka_03 - แแฌแแแกแแนแแแแฌ-แ (Jataka Commentary 3)
- 08_jataka_04 - แแฌแแแกแแนแแแแฌ-แ (Jataka Commentary 4)
- 08_jataka_05 - แแฌแแแกแแนแแแแฌ-แแแนแ (Jataka Commentary 5)
- 08_jataka_06 - แแฌแแแกแแนแแแแฌ-แ (Jataka Commentary 6)
- 08_jataka_07 - แแฌแแแกแแนแแแแฌ-แแแนแ (Jataka Commentary 7)
- 08_visuddhimagga_00 - แแญแแฏแแนแแญแแแบ แแญแแซแแบแธ (Visuddhimagga Intro)
- 08_visuddhimagga_01 - แแญแแฏแแนแแญแแแบ แแผแแบแแฌแแผแแบ-แ (Visuddhimagga Vol 1)
- 08_visuddhimagga_02 - แแญแแฏแแนแแญแแแบ แแผแแบแแฌแแผแแบ-แแฏ (Visuddhimagga Vol 2)
- 08_visuddhimagga_03 - แแญแแฏแแนแแญแแแบ แแผแแบแแฌแแผแแบ-แ (Visuddhimagga Vol 3)
- 08_visuddhimagga_04 - แแญแแฏแแนแแญแแแบ แแผแแบแแฌแแผแแบ-แ (Visuddhimagga Vol 4)
Disclaimer: These texts are religious and philosophical in nature, reflecting traditional Theravada Buddhist scholarship. Users are encouraged to apply care, respect, and contextual awareness when using these texts.
