datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Burmese-Handwritten-Sentence-Dataset
Burmese Handwritten Sentence Dataset (BHSD)
BHSD is a sentence-level Burmese handwriting dataset developed for optical character recognition (OCR), handwritten text recognition (HTR), error analysis, robustness testing, and research on low-resource scripts.
The dataset was created by Ah Maung Oo and DatarrX through the voluntary contributions of 54 handwriting writers.
This dataset would not have been possible without its volunteers. Every handwritten image in BSHD exists… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-Handwritten-Sentence-Dataset.burmese_ocr_data_hfPlease visit the GitHub repository for other Myanmar Language datasets.
Burmese OCR Dataset (Hugging Face Format)
This is a reformatted version of alexbeatson/burmese_ocr_data converted into the native Hugging Face datasets format for easier loading and integration with modern OCR training pipelines.
Dataset Description
This dataset contains Burmese text images and their corresponding ground truth text extracted from real-life documents, suitable for training Optical… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/burmese_ocr_data_hf.burmese-ocr-1m
Burmese OCR 1M
A dataset of 897,883 rendered line images for Burmese/Pali text recognition
(OCR). Images are grayscale/bi-level line crops paired with their transcription.
General-purpose: usable with any OCR engine (Tesseract, Kraken, Calamari, or
custom HTR models).
Splits
split
rows
files
train
718,515
5 parquet shards
validation
89,625
1 parquet shard
test
89,743
1 parquet shard
total
897,883
7
Columns
column
type… See the full description on the dataset page: https://huggingface.co/datasets/pndaza/burmese-ocr-1m.burmese-pyu-character-recognition
Burmese Pyu Character Recognition Dataset
English | မြန်မာဘာသာ
Overview
This Dataset is an image collection created for the purpose of computer recognition (Character Recognition) and research of the "Pyu" alphabet, an ancient script of Myanmar.
Brief Historical Background
The Pyu people were one of the earliest major ethnic groups to inhabit Myanmar, settling in the region since the early AD periods. The Pyu culture is of great importance when studying the… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/burmese-pyu-character-recognition.burmese-ocr-5m
Burmese OCR 5M
A dataset of 5,104,476 rendered line images for Burmese/Pali text recognition
(OCR). General-purpose: usable with any OCR engine (Tesseract, Kraken, Calamari,
or custom HTR models).
Splits
split
rows
train
4,082,086
validation
511,435
test
510,955
total
5,104,476
Columns
column
type
description
image
image
PNG line image (struct: bytes + path)
text
string
Transcription (Burmese/Pali)
image_id… See the full description on the dataset page: https://huggingface.co/datasets/pndaza/burmese-ocr-5m.
