Burmese
Datasets
All datasets matching “Burmese”Burmese-Handwritten-Sentence-Dataset
Burmese Handwritten Sentence Dataset (BHSD)
BHSD is a sentence-level Burmese handwriting dataset developed for optical character recognition (OCR), handwritten text recognition (HTR), error analysis, robustness testing, and research on low-resource scripts.
The dataset was created by Ah Maung Oo and DatarrX through the voluntary contributions of 54 handwriting writers.
This dataset would not have been possible without its volunteers. Every handwritten image in BSHD exists… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-Handwritten-Sentence-Dataset.burmese-speech-refined-openslr-80
Burmese Speech Refined OpenSLR-80
Summary
This dataset is a speech dataset developed based on the original OpenSLR Dataset (SLR80), with the text and audio data carefully reviewed and further refined for Burmese language applications.
In the original OpenSLR Dataset, the Burmese text was transcribed based on how the words were pronounced in the corresponding audio recordings. In this dataset, the original audio and text data were used as a reference, and the text… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/burmese-speech-refined-openslr-80.Burmese-Classics-OCR-RAW
Burmese (Myanmar) Books Dataset – Burmese Classics OCR (Daily Rolling Project)
Overview
A daily rolling dataset of Burmese books, built for AI, OCR, and NLP research.Each entry includes OCR text with metadata: title, author, page index, and Burmese character ratio.
This project fills a critical gap in Burmese-language resources:
Scarcity of public-domain Burmese text.
High technical and financial barriers to corpus building.
Enables incremental, open access for… See the full description on the dataset page: https://huggingface.co/datasets/minthanthtoo-cs/Burmese-Classics-OCR-RAW.fine-burmesemyanmar-burmese-speech-datasetburmese-synthetic-speech-corpus
Burmese Synthetic Speech Corpus (DatarrX/burmese-synthetic-speech-corpus)
Overview
The Burmese Synthetic Speech Corpus is a high-fidelity, manually curated audio dataset specifically designed to advance Text-to-Speech (TTS) systems, speech recognition, and other audio-driven Machine Learning tasks for the Burmese (Myanmar) language.
Created by DatarrX, this dataset bridges the gap in low-resource speech technologies by providing highly natural, native-sounding… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/burmese-synthetic-speech-corpus.
