datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tipitaka_myanmar_translation_books
Myanmar Tipitaka Translation (60 Books)
This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga.
The texts have been converted into a clean, structured JSONL format, suitable for:
Natural Language Processing (NLP)
LLM Training & Fine-tuning
Digital Humanities Research
Dhamma Study Applications
📊 Dataset Statistics
Total Books: 60
Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.tipitaka-dataset
Myanmar Tipitaka Dataset (Pali Corpus)
Dataset Summary
The Myanmar Tipitaka Dataset is a high-quality, structured collection of the Buddhist Pali Canon, transcribed in the Myanmar (Burmese) script. This dataset contains 462,504 paragraphs, covering the entire "Triple Basket" (Tipitaka) of Theravada Buddhism, including the original Mula (Canonical texts), Atthakatha (Commentaries), and Tika (Sub-commentaries).
This project was initiated by DatarrX to provide a clean… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/tipitaka-dataset.tipitaka-storage
