datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PAlign-PAPI-personality_prompt.json-cleanedAdapted from
"Personality Alignment of Large Language Models" by Minjun Zhu and Linyi Yang and Yue Zhang
and the associated GitHub repository zhu-minjun/PAlign.
The contents of said repo were declared public domain; in that spirit, this Alpaca-formatted file has also been released as public domain.
myanmar-english-pali-dictionary
Myanmar–English–Pali Dictionary
Dataset Summary
This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein).
It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary.
The dataset is intended for research and educational purposes, including but not limited to:
Natural Language Processing (NLP)
Machine Translation (MT)
Lexicography
Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.P-ALIGNvinaya-pitaka-pali-myanmar-parallel
Vinaya Pitaka: Pali-Myanmar Parallel Dataset
Description
This dataset provides a professionally aligned, paragraph-level parallel corpus of the Vinaya Pitaka (The Code of Monastic Discipline). It features the original Pali text (presented in Myanmar script) alongside its modern Myanmar translation.
The dataset covers all five major volumes of the Vinaya:
Pārājika (ပါရာဇိကပါဠိ / ပါရာဇိကဏ်)
Pācittiya (ပါစိတ္တိယပါဠိ / ပါစိတ်)
Mahāvagga (မဟာဝဂ္ဂပါဠိ / မဟာဝါ)
Cūḷavagga… See the full description on the dataset page: https://huggingface.co/datasets/freococo/vinaya-pitaka-pali-myanmar-parallel.pali-vietpali-words-myanmar-script
Pali Words in Myanmar Script (Master Index)
This dataset is a master index of 220,252 unique Pali words written in Myanmar (Burmese) Unicode script, intended for reuse across linguistic, religious, and computational workflows.
Data Fields
Each record contains:
word_id: A stable, sequential integer identifier.
pali_word: A Pali lexical item rendered in Myanmar Unicode script.
Data Processing Methodology
The dataset was constructed using the following steps:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/pali-words-myanmar-script.mon-eng-mya-pali-lexicon
🇲🇲 🇬🇧 Mon-Eng-Mya-Pali Lexicon Dataset for AI
This dataset is a comprehensive, multilingual lexicon of the Mon language (ဘာသာမန်), paired with English, Burmese (Myanmar), and Pali equivalents. It is structured specifically for Natural Language Processing (NLP), Large Language Model (LLM) fine-tuning, Machine Translation, and Retrieval-Augmented Generation (RAG) applications.
ဖိုင်ဒေတာဝေါဟာရ ကေုာံ အဘိဓာန်ဘာသာမန် (၄ ဘာသာ) သွက်ဂွံဗ္တောန် AI (Training) ကေုာံ စကာပ္ဍဲ AI… See the full description on the dataset page: https://huggingface.co/datasets/Nenemin95/mon-eng-mya-pali-lexicon.palinkahtest-palindromepalioshoptest
