Myanmar
Datasets
All datasets matching “Myanmar”myanmar-ocr-dataset-for-vlm
Myanmar OCR Dataset
A synthetic OCR dataset for fine-tuning Vision Language Models (VLMs) on Myanmar (Burmese) text recognition. It contains page images paired with their ground-truth text, sourced from chuuhtetnaing/mm-lib-book-dataset and rendered into page images using various Myanmar fonts.
Subsets
Subset
Description
Details
single_font
Rendered with Pyidaungsu font only
437 books
multi_font
Rendered with 76 Myanmar fonts
3 books (ပဋ္ဌာန်းမြတ်ဒေသနာ၊… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset-for-vlm.myanmar-ocr-dataset
Myanmar OCR Dataset
A synthetic dataset for training and fine-tuning Optical Character Recognition (OCR) models specifically for the Myanmar language.
Dataset Description
This dataset contains synthetically generated OCR images created specifically for Myanmar text recognition tasks. The images were generated using myanmar-ocr-data-generator, a fork of TextRecognitionDataGenerator with fixes for proper Myanmar character splitting.
Direct Download
Available… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset.nug_myanmar_asr
366 Hours NUG Myanmar ASR Dataset
The NUG Myanmar ASR Dataset is the first large-scale open Burmese speech dataset — now expanded to over 521,476 audio-text pairs, totaling ~366 hours of clean, segmented audio. All data was collected from public-service educational broadcasts by the National Unity Government (NUG) of Myanmar and the FOEIM Academy.
This dataset is released under a CC0 1.0 Universal license — fully open and public domain. No attribution required.
🕊️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/nug_myanmar_asr.ipfs_myanmar_laws_ir
Myanmar legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_myanmar_laws (revision b744df8f35b0f9eff35df4eeb57e3def96e80607) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Myanmar prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_myanmar_laws_ir.myanmar-fineweb-2-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Fineweb2 Dataset
A preprocessed subset of the Fineweb2 dataset containing only Myanmar language text, with consistent Unicode encoding.
Dataset Description
This dataset is derived from the Fineweb2 created by HuggingFaceFW. It contains only the Myanmar language portion of the original Fineweb2 dataset, with additional preprocessing to standardize text encoding.
Filtered and Removed… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-fineweb-2-dataset.voa_myanmar_voices
VOA Myanmar Voices
Burmese (Myanmar) speech corpus chunked into 20-second FLAC clips with transcripts. Derived from VOA Burmese radio broadcasts (public domain, U.S. 17 U.S.C. § 105).
Contents
File
Size
Description
voa-00000000.tar … voa-00000238.tar
498 GB
149 WebDataset shards
voa_transcripts.parquet
404 MB
1,424,257 (key, text) pairs
voa_transcripts.jsonl
1.5 GB
Same data, line-oriented
Audio: 16 kHz mono FLAC, exactly 20.00 s per chunk… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_voices.
