CoolFace
Datasetpublic

paodigitalhub/blk-text-corpus

Verified Pa'O (Blk) Text Corpus - Pa'O Digital Hub Dataset Summary This is the official, verified parallel dataset for the Pa'O language (ISO 639-3: blk) and Burmese (Myanmar) translations, published by Pa'O Digital Hub. The corpus is systematically collected, reviewed, standardized, and verified through the established linguistic and editorial workflow of Pa'O Digital Hub. The Pa'O sentences are based on authentic language usage by Pa'O native speakers and are… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/blk-text-corpus.

sourceHugging Facecc-by-4.0updated 7d agoView on Hugging Face
1likes309downloads
Dataset Card

Verified Pa'O (Blk) Text Corpus - Pa'O Digital Hub

Dataset Summary

This is the official, verified parallel dataset for the Pa'O language (ISO 639-3: blk) and Burmese (Myanmar) translations, published by Pa'O Digital Hub.

The corpus is systematically collected, reviewed, standardized, and verified through the established linguistic and editorial workflow of Pa'O Digital Hub. The Pa'O sentences are based on authentic language usage by Pa'O native speakers and are reviewed to preserve genuine Pa'O linguistic forms, writing conventions, spelling, orthography, and grammar.

ဤသည် Pa'O Digital Hub ၏ တရားဝင်လုပ်ငန်းစဉ်များမှတစ်ဆင့် ပအိုဝ်းဒေသခံများ၏ အမှန်တကယ်အသုံးပြုသော ပအိုဝ်းစကားနှင့် စာပေကို အခြေခံ၍ စနစ်တကျ စုဆောင်း၊ စိစစ်၊ ပြင်ဆင်၊ အတည်ပြုပြီး အများပြည်သူ လေ့လာအသုံးပြုနိုင်ရန် ထုတ်ပြန်ထားသော တရားဝင် ပအိုဝ်း-မြန်မာ အပြိုင်ဝါကျဒေတာစု ဖြစ်ပါသည်။

Linguistic Verification & Quality Assurance

(ဘာသာစကားနှင့် အရည်အသွေး စိစစ်အတည်ပြုမှု)

The Pa'O text in this corpus undergoes a multi-stage linguistic and editorial verification process before publication.

  1. 1.Native Pa'O Language Verification
  2. 2.The corpus is based on authentic Pa'O language usage by Pa'O native speakers and local Pa'O linguistic contributors.
  3. 3.The text is reviewed to ensure that it represents genuine Pa'O language usage rather than artificially constructed or mechanically translated sentences.
  1. 1.Purity of Pa'O Language Usage
  2. 2.The text is reviewed for unnecessary mixing or borrowing from other languages.
  3. 3.Where the purpose is to represent genuine Pa'O vocabulary and expression, words or expressions that are identified as foreign-language insertions, unnecessary borrowings, or non-native forms are filtered or revised during the verification process.
  4. 4.This process is intended to preserve authentic Pa'O language usage in the corpus.
  1. 1.Orthography, Spelling, and Written Forms
  2. 2.Pa'O sentences are carefully reviewed for:
  3. 3.Spelling accuracy;
  4. 4.Correct written forms;
  5. 5.Correct character and character-sequence usage;
  6. 6.Consistent orthographic conventions; and
  7. 7.Consistency in Pa'O written language.
  1. 1.Pa'O Grammar Verification
  2. 2.The sentences are reviewed according to Pa'O grammatical structures and usage to ensure that the sentence construction, word order, particles, expressions, and overall linguistic forms are appropriate and meaningful in Pa'O.
  1. 1.Pa'O Literary Standard & Writing Conventions
  2. 2.The corpus strictly adheres to the Taunggyi Kham Kaung literary standard (တောင်ကြီး ပအိုဝ်းခမ်းကောင်စာမူ) based on the official "Pa'O Primer Book" (ပအိုဝ်ႏလိတ်မွူးစောင်ႏ) published by the Pa'O Central Literature and Culture Sangha Association (ပအိုဝ်ႏလိတ်လုဲင်ꩻစွစ်ꩻလီတာႏဗဟိုႏသံႏဃာႏစွိုꩻ).
  3. 3.It also follows established writing practices associated with Khun Ku Wae (ခွန်ကုဝေ), a prominent Pa'O literary scholar and contributor to Pa'O written-language standardization.
  1. 1.Guidance from Pa'O Literary Scholarship
  2. 2.The verification process also follows linguistic and literary guidance associated with Sa Htom Pay Khun Maung (သထွုံႏပေႏခွန်မောင်ႏ), particularly with regard to preserving authentic Pa'O language and genuine Pa'O linguistic usage.
  1. 1.Editorial Review Before Publication
  2. 2.Only after the relevant linguistic, spelling, grammatical, orthographic, and editorial checks are completed is the material prepared and published as part of the Pa'O Digital Hub open dataset.

Therefore, “Verified” in the title of this corpus refers to the linguistic and editorial verification process applied to the Pa'O text, rather than merely indicating that the data has been collected or uploaded.


Key Highlights & Verification

(ယုံကြည်စိတ်ချရမှုနှင့် အတည်ပြုချက်)

  • Native Sourced: ပအိုဝ်းဒေသခံများနှင့် ပအိုဝ်းဘာသာစကားကျွမ်းကျင်သူများ၏ အမှန်တကယ်အသုံးပြုသော ဘာသာစကားကို အခြေခံထားခြင်း။
  • Authentic Pa'O Usage: ပအိုဝ်းစကားစစ်စစ်နှင့် သဘာဝကျသော ပအိုဝ်းဘာသာစကားအသုံးအနှုန်းများကို ထိန်းသိမ်းရန် စိစစ်ထားခြင်း။
  • Language Purity Review: အခြားဘာသာစကားမှ မလိုအပ်ဘဲ ရောနှောအသုံးပြုထားသော စကားလုံးများ၊ မွေးစားအသုံးအနှုန်းများနှင့် မူရင်းပအိုဝ်းအသုံးအနှုန်းနှင့် မကိုက်ညီသော ပုံစံများကို စိစစ်ပြီး လိုအပ်ပါက ဖယ်ရှားခြင်း သို့မဟုတ် ပြင်ဆင်ခြင်းပြုထားခြင်း။
  • Orthographically Reviewed: ပအိုဝ်းစာလုံးပေါင်း၊ စာလုံးဆင့်၊ စာရေးသားပုံနှင့် အရေးအသားစံများကို စိစစ်ထားခြင်း။
  • Grammatically Reviewed: ပအိုဝ်းသဒ္ဒါနှင့် ဝါကျတည်ဆောက်ပုံအရ စိစစ်ထားခြင်း။
  • Literary Standard: "တောင်ကြီး ပအိုဝ်းခမ်းကောင်စာမူ" ကို အခြေခံထားပြီး "ပအိုဝ်ႏလိတ်လုဲင်ꩻစွစ်ꩻလီတာႏဗဟိုႏသံႏဃာႏစွိုꩻ" မှ ထုတ်ဝေသည့် "ပအိုဝ်ႏလိတ်မွူးစောင်ႏ" ပါ ရေးထုံးများနှင့် စာပေပညာရှင် ခွန်ကုဝေ ၏ စာရေးထုံး စံနှုန်းများကို စနစ်တကျ လိုက်နာရေးသားထားခြင်း။
  • Literary Guidance: စာပေပညာရှင် သထွုံႏပေႏခွန်မောင်ႏ ၏ လမ်းညွှန်ချက်များနှင့်အညီ စစ်မှန်သော ပအိုဝ်းစကားအသုံးအနှုန်းများကို ထိန်းသိမ်းထားခြင်း။
  • Officially Verified: အထက်ပါ စိစစ်ရေးအဆင့်များကို ဖြတ်သန်းပြီးမှ Pa'O Digital Hub ၏ တရားဝင် open dataset အဖြစ် ထုတ်ပြန်ထားခြင်း။
  • Standardized Format: ISO 639-3 language code (blk) ကို အသုံးပြု၍ dataset ID နှင့် data structure များကို စနစ်တကျ တည်ဆောက်ထားခြင်း။


Dataset Structure

(ဒေတာဖွဲ့စည်းပုံ)

The dataset is stored in CSV (Comma-Separated Values) format with UTF-8 BOM encoding.

Column NameTypeDescription
idstringUnique Sentence Identifier (e.g., blk_001)
paoh_sentencestringVerified Pa'O sentence
myanmar_translationstringBurmese (Myanmar) translation of the Pa'O sentence

Example Data

id,paohsentence,myanmartranslation blk001,မင်္ဂလာႏဒျာႏဩ,မင်္ဂလာပါ blk002,အုံဟောဝ်နေဟောင်း,နေကောင်းရဲ့လား blk_003,အောဝ်ႏမာꩻခိꩻတမုဲင်ꩻဟောင်း,ဘာတွေလုပ်နေလဲ


Data Quality Principles

(ဒေတာအရည်အသွေး ထိန်းသိမ်းရေးမူများ)

The corpus is intended to prioritize:

  1. 1.Authenticity — genuine Pa'O language usage.
  2. 2.Native linguistic validity — language forms used and recognized by Pa'O native speakers.
  3. 3.Orthographic accuracy — correct Pa'O written forms and spelling.
  4. 4.Grammatical accuracy — conformity with Pa'O grammatical structures.
  5. 5.Literary consistency — adherence to established Pa'O writing conventions (Taunggyi Kham Koung standard).
  6. 6.Translation accuracy — Burmese translations that accurately represent the intended meaning of the Pa'O source sentence.
  7. 7.Transparency — publication as an openly accessible dataset for research, education, language technology, and preservation.

Intended Uses

This corpus may be useful for:

  • Pa'O language research;
  • Pa'O–Myanmar machine translation;
  • Natural Language Processing (NLP);
  • Speech and language technology;
  • Pa'O language learning;
  • Linguistic research;
  • Educational applications;
  • Digital preservation of the Pa'O language;
  • Development and evaluation of Pa'O language models.

Researchers and developers should preserve the original `paoh_sentence` text when using the corpus for linguistic or computational research.


How to Use

Python / Hugging Face Datasets

# 1. Install and import required libraries for reading data # !pip install datasets pandas --quiet

import pandas as alldataset from datasets import load_dataset

print("🔄 Starting to download all data for 'paodigitalhub/blk-text-corpus' from Hugging Face...")

try: # 2. Load the dataset rawdata = loaddataset("paodigitalhub/blk-text-corpus")

# 3. Convert all retrieved data into a DataFrame splitname = list(rawdata.keys())[0] df = alldataset.DataFrame(rawdata[splitname])

# 4. Set display options to show all results without truncation alldataset.setoption('display.maxrows', None) alldataset.setoption('display.maxcolumns', None) alldataset.setoption('display.maxcolwidth', None) alldataset.set_option('display.width', 1000)

# 5. Display the total number of rows print(f"✅ Data loaded successfully. Total rows: {len(df)}\n")

# 6. Display all rows completely on the screen print("🖥️ --- All text in the dataset is as follows ---") print(df)

except Exception as e: print(f"❌ Connection failed: {e}")

Load the CSV File Explicitly

If you want to specify the CSV files directly:

from datasets import load_dataset

dataset = loaddataset( "csv", datafiles="data/paoh-sentences/*.csv" ) print(dataset["train"][0])

Read with Pandas

import pandas as pd

df = pd.read_csv( "data/paoh-sentences/blk-sentences-001.csv", encoding="utf-8-sig" ) print(df.head())


Directory Structure

blk-text-corpus/ ├── README.md └── data/ └── paoh-sentences/ └── blk-sentences-001.csv


Language Identification

  • Language: Pa'O
  • ISO 639-3: blk
  • Native Name: ပအိုဝ်ႏ

The blk language identifier is used consistently throughout the dataset to identify the Pa'O language.


Open Data and Transparency

Pa'O Digital Hub publishes this corpus as an openly accessible resource so that researchers, educators, developers, and Pa'O communities can study, use, evaluate, and further contribute to Pa'O language technology and preservation.

The corpus is intended not only as machine-readable data, but also as a contribution to the preservation, documentation, standardization, and digital development of the Pa'O language.


Attribution

When using this dataset in research, software, educational materials, or derivative datasets, please credit:

Pa'O Digital Hub — Verified Pa'O (Blk) Text Corpus

Please retain the original dataset attribution and license information when redistributing or creating derivative works.


License

This dataset is released under the CC-BY-4.0 license.

Users are permitted to share and adapt the material in accordance with the terms of the license, with appropriate attribution to Pa'O Digital Hub.


Citation

If you use this dataset in academic research, publications, software, or other public projects, please cite:

Pa'O Digital Hub. Verified Pa'O (Blk) Text Corpus. Pa'O Digital Hub, 2026. Language: Pa'O (ISO 639-3: blk). License: CC-BY-4.0.


Maintained by

Pa'O Digital Hub A community-oriented initiative for Pa'O language documentation, preservation, standardization, education, and digital language technology.