CoolFace
20 results

Myanmar

chuuhtetnaing /myanmar-ocr-dataset-for-vlm Myanmar OCR Dataset A synthetic OCR dataset for fine-tuning Vision Language Models (VLMs) on Myanmar (Burmese) text recognition. It contains page images paired with their ground-truth text, sourced from chuuhtetnaing/mm-lib-book-dataset and rendered into page images using various Myanmar fonts. Subsets Subset Description Details single_font Rendered with Pyidaungsu font only 437 books multi_font Rendered with 76 Myanmar fonts 3 books (ပဋ္ဌာန်းမြတ်ဒေသနာ၊… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset-for-vlm.image100K<n<1M2 likes680 downloads5mo agoHugging Facechuuhtetnaing /myanmar-ocr-dataset Myanmar OCR Dataset A synthetic dataset for training and fine-tuning Optical Character Recognition (OCR) models specifically for the Myanmar language. Dataset Description This dataset contains synthetically generated OCR images created specifically for Myanmar text recognition tasks. The images were generated using myanmar-ocr-data-generator, a fork of TextRecognitionDataGenerator with fixes for proper Myanmar character splitting. Direct Download Available… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset.imageimage-to-text1M<n<10M10 likes348 downloads1y agoHugging Facefreococo /nug_myanmar_asr 366 Hours NUG Myanmar ASR Dataset The NUG Myanmar ASR Dataset is the first large-scale open Burmese speech dataset — now expanded to over 521,476 audio-text pairs, totaling ~366 hours of clean, segmented audio. All data was collected from public-service educational broadcasts by the National Unity Government (NUG) of Myanmar and the FOEIM Academy. This dataset is released under a CC0 1.0 Universal license — fully open and public domain. No attribution required. 🕊️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/nug_myanmar_asr.audioautomatic-speech-recognition100K<n<1M3 likes260 downloads1y agoHugging Facejusticedao /ipfs_myanmar_laws_ir Myanmar legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_myanmar_laws (revision b744df8f35b0f9eff35df4eeb57e3def96e80607) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Myanmar prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_myanmar_laws_ir.tabulartext-retrieval10K<n<100K0 likes246 downloads2h agoHugging Facechuuhtetnaing /myanmar-fineweb-2-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Fineweb2 Dataset A preprocessed subset of the Fineweb2 dataset containing only Myanmar language text, with consistent Unicode encoding. Dataset Description This dataset is derived from the Fineweb2 created by HuggingFaceFW. It contains only the Myanmar language portion of the original Fineweb2 dataset, with additional preprocessing to standardize text encoding. Filtered and Removed… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-fineweb-2-dataset.tabulartext-generation1M<n<10M0 likes238 downloads1y agoHugging Facefreococo /voa_myanmar_voices VOA Myanmar Voices Burmese (Myanmar) speech corpus chunked into 20-second FLAC clips with transcripts. Derived from VOA Burmese radio broadcasts (public domain, U.S. 17 U.S.C. § 105). Contents File Size Description voa-00000000.tar … voa-00000238.tar 498 GB 149 WebDataset shards voa_transcripts.parquet 404 MB 1,424,257 (key, text) pairs voa_transcripts.jsonl 1.5 GB Same data, line-oriented Audio: 16 kHz mono FLAC, exactly 20.00 s per chunk… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_voices.automatic-speech-recognition1M<n<10M0 likes228 downloads2d agoHugging Face