paodigitalhub/pao-sentences-dataset
Pa'O Sentences Dataset is a text corpus for Pa'O language (ပအိုဝ်ႏ) containing structured line-by-line sentences designed for NLP, LLM pre-training, and machine translation. 📝 Pa'O Sentences Dataset (ပအိုဝ်ႏ လိက်လာႏငေါဝ်းရဲဉ်ႏ ရွမ်ခြွဉ်းဗူႏ) 📌 Project Summary (ထာꩻမာꩻခြပ်ရဲဉ်ႏ နပ်ထွားရဲပ်အအဲဉ်ႏ) The Pa'O Sentences Dataset is an open-source textual corpus developed to support Natural Language Processing (NLP), Large Language Model (LLM) pre-training, Machine… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-sentences-dataset.
Update README.md
new update dataset
new update to Pa'O (blk) pao-sentences-dataset-002(txt/parquet)file
Open
blk-Mymr
Update README.md
Upload pao-sentences-001.parquet
Delete data/pao_sentences_001.parquet
Upload pao_sentences_001.parquet
txt to parquet changed
Update README.md
Update README.md
Update data/pao-sentences-000.txt
Update README.md
Update README.md
Upload pao-sentences-001.txt
Create data/pao-sentences-000.txt
initial commit
