CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chuuhtetnaing /myanmar-ocr-dataset-for-vlm Myanmar OCR Dataset A synthetic OCR dataset for fine-tuning Vision Language Models (VLMs) on Myanmar (Burmese) text recognition. It contains page images paired with their ground-truth text, sourced from chuuhtetnaing/mm-lib-book-dataset and rendered into page images using various Myanmar fonts. Subsets Subset Description Details single_font Rendered with Pyidaungsu font only 437 books multi_font Rendered with 76 Myanmar fonts 3 books (ပဋ္ဌာန်းမြတ်ဒေသနာ၊… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset-for-vlm.image100K<n<1M2 likes680 downloads5mo agoHugging Face02chuuhtetnaing /myanmar-ocr-dataset Myanmar OCR Dataset A synthetic dataset for training and fine-tuning Optical Character Recognition (OCR) models specifically for the Myanmar language. Dataset Description This dataset contains synthetically generated OCR images created specifically for Myanmar text recognition tasks. The images were generated using myanmar-ocr-data-generator, a fork of TextRecognitionDataGenerator with fixes for proper Myanmar character splitting. Direct Download Available… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset.imageimage-to-text1M<n<10M10 likes348 downloads1y agoHugging Face03freococo /nug_myanmar_asr 366 Hours NUG Myanmar ASR Dataset The NUG Myanmar ASR Dataset is the first large-scale open Burmese speech dataset — now expanded to over 521,476 audio-text pairs, totaling ~366 hours of clean, segmented audio. All data was collected from public-service educational broadcasts by the National Unity Government (NUG) of Myanmar and the FOEIM Academy. This dataset is released under a CC0 1.0 Universal license — fully open and public domain. No attribution required. 🕊️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/nug_myanmar_asr.audioautomatic-speech-recognition100K<n<1M3 likes260 downloads1y agoHugging Face04justicedao /ipfs_myanmar_laws_ir Myanmar legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_myanmar_laws (revision b744df8f35b0f9eff35df4eeb57e3def96e80607) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Myanmar prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_myanmar_laws_ir.tabulartext-retrieval10K<n<100K0 likes246 downloads7h agoHugging Face05chuuhtetnaing /myanmar-fineweb-2-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Fineweb2 Dataset A preprocessed subset of the Fineweb2 dataset containing only Myanmar language text, with consistent Unicode encoding. Dataset Description This dataset is derived from the Fineweb2 created by HuggingFaceFW. It contains only the Myanmar language portion of the original Fineweb2 dataset, with additional preprocessing to standardize text encoding. Filtered and Removed… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-fineweb-2-dataset.tabulartext-generation1M<n<10M0 likes238 downloads1y agoHugging Face06freococo /voa_myanmar_voices VOA Myanmar Voices Burmese (Myanmar) speech corpus chunked into 20-second FLAC clips with transcripts. Derived from VOA Burmese radio broadcasts (public domain, U.S. 17 U.S.C. § 105). Contents File Size Description voa-00000000.tar … voa-00000238.tar 498 GB 149 WebDataset shards voa_transcripts.parquet 404 MB 1,424,257 (key, text) pairs voa_transcripts.jsonl 1.5 GB Same data, line-oriented Audio: 16 kHz mono FLAC, exactly 20.00 s per chunk… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_voices.automatic-speech-recognition1M<n<10M0 likes228 downloads2d agoHugging Face07freococo /voa_myanmar_asr_audio_1 📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology. Overview This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.audioautomatic-speech-recognition1M<n<10M1 likes214 downloads1y agoHugging Face08chuuhtetnaing /myanmar-speech-dataset-for-asrPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Speech Dataset for ASR This dataset is a comprehensive collection of Myanmar language speech data specifically curated for Automatic Speech Recognition (ASR) task. It combines following datasets: Myanmar Speech Dataset (Google Fleurs) Myanmar Speech Dataset (OpenSLR-80) Ko-Yin-Maung/mig-burmese-audio-transcription By merging these complementary resources, this dataset provides a more robust… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-for-asr.audioautomatic-speech-recognition1K<n<10K2 likes190 downloads9mo agoHugging Face09ayehninnkhine /myanmar_news Dataset Card for Myanmar_News Dataset Summary The Myanmar news dataset contains article snippets in four categories: Business, Entertainment, Politics, and Sport. These were collected in October 2017 by Aye Hninn Khine Languages Myanmar/Burmese language Dataset Structure Data Fields text - text from article category - a topic: Business, Entertainment, Politic, or Sport (note spellings) Data Splits One training set (8,116 total… See the full description on the dataset page: https://huggingface.co/datasets/ayehninnkhine/myanmar_news.texttext-classification1K<n<10K6 likes186 downloads2y agoHugging Face10thantzinphyo /myanmar-shopvoice Myanmar Shopvoice Dataset (300 Sample) This is a randomly sampled subset of 300 Myanmar shop voice audio chunks for Whisper fine-tuning and evaluation. Dataset Details Total Audio Chunks: 300 Audio Format: 16 kHz WAV, mono Language: Myanmar (Burmese) Features audio: Audio feature (16 kHz WAV audio player) transcription: Myanmar sentence transcription text source_audio: Source continuous recording file name source_line: Index line of transcript… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/myanmar-shopvoice.audioautomatic-speech-recognitionn<1K0 likes168 downloads1mo agoHugging Face11freococo /115hours_pvtv_myanmar_asr 115 Hours PVTV Myanmar ASR This dataset contains 156,262 audio-transcript pairs of spoken Burmese, totaling approximately 115.31 hours. The audio segments were extracted from publicly available YouTube videos published by PVTV and aligned using subtitle timestamps. Dedication This dataset would not exist without the persistent voices of PVTV editors, journalists, narrators, and production teams, who continue to speak to the people under difficult conditions. PVTV is the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/115hours_pvtv_myanmar_asr.audioautomatic-speech-recognition100K<n<1M2 likes159 downloads1y agoHugging Face12freococo /huggingface_myanmar_english_translation Cleaned & Sorted Myanmar-English Translation Dataset This dataset is a cleaned, Unicode-normalized, and sorted version of the Myanmar (Burmese) subset from the massive FineTranslations dataset. While the original dataset is excellent, Myanmar text on the web is often a mix of standard Unicode and the non-standard Zawgyi encoding. This repository fixes those encoding issues to provide a high-quality dataset for NLP tasks. Key Improvements in this Version Zawgyi… See the full description on the dataset page: https://huggingface.co/datasets/freococo/huggingface_myanmar_english_translation.texttranslation1M<n<10M1 likes153 downloads7mo agoHugging Face13kkomyoeminaung /myanmar-native-text Dataset Card for myanmar-native-text Dataset Summary ဒီ dataset က myanmar-native-text အတွက် ဖန်တီးထားတာပါ။ Languages Myanmar (my) / English (en) Dataset Structure Data Instances { "text": "နမူနာ စာသား", "label": "အညွှန်း" } Data Fields text: main content, label: optional. Data Splits Split Files train data/native_00000.jsonl, data/native_00001.jsonl, data/native_00002.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/myanmar-native-text.1 likes150 downloads2mo agoHugging Face14kornwtp /thura-myanmar-news-mya-classificationref: https://huggingface.co/datasets/ThuraAung1601/myanmar_news text10K<n<100K1 likes148 downloads2y agoHugging Face15endomorphosis /ipfs_myanmar_laws Myanmar Laws Legal corpus for Myanmar from the ipfs_datasets_py legal scrapers (endomorphosis). Instruments: 725 Last updated: 2026-09-16 Source: official gazettes and legal portals; see source_url per record. Schema: id, title, text, source_url, source_type, jurisdiction, country, language, eli, date, retrieved_at, license, law_status, identifiers. textn<1K0 likes147 downloads8d agoHugging Face16kkomyoeminaung /myanmar-books-longform-1024 Dataset Card for myanmar-books-longform-1024 Dataset Summary ဒီ dataset က myanmar-books-longform-1024 အတွက် ဖန်တီးထားတာပါ။ Languages Myanmar (my) / English (en) Dataset Structure Data Instances { "text": "နမူနာ စာသား", "label": "အညွှန်း" } Data Fields text: main content, label: optional. Data Splits Split Files train books_chunk_0000.jsonl, books_chunk_0001.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/myanmar-books-longform-1024.0 likes144 downloads2mo agoHugging Face17jojo-ai-mst /Myanmar-Tuberculosis-Guidelines-Instructions Myanmar Tuberculosis Guidelines Instructions A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages. Authors: Min Si Thu, Khin Myat Noe Abstract Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.imagequestion-answering1K<n<10K1 likes143 downloads5mo agoHugging Face18akhtet /myanmar-xnli Dataset Card for myXNLI Dataset Summary The myXNLI corpus extends XNLI corpus with Myanmar (Burmese) language. For myXNLI, we human-translated all 7,500 sentence pairs from XNLI English dev and test sets into Myanmar. The NLI and Genre labels from English dev and test sets are also reused for the Myanmar datasets. The dataset also includes the NLI training data in Myanmar which is created by machine-translating the MultiNLI training data from English into Myanmar. Similar… See the full description on the dataset page: https://huggingface.co/datasets/akhtet/myanmar-xnli.texttext-classification100K<n<1M7 likes140 downloads1y agoHugging Face19URajinda /myanmar_spoken_corpus Credits and Acknowledgments This dataset is built upon the foundational work of the Myanmar Spoken Corpus by freococo (Wynn). Original Dataset Source: freococo/myanmar_spoken_corpus Modifications: * Curated and filtered for specific training needs of the ShweYon model. Integrated with 3% English subset for bilingual proficiency maintenance. Re-formatted into 36 shards for optimized Continued Pre-training (CPT). We are deeply grateful to freococo for their contribution to the… See the full description on the dataset page: https://huggingface.co/datasets/URajinda/myanmar_spoken_corpus.text1M<n<10M0 likes119 downloads8mo agoHugging Face20chuuhtetnaing /myanmar-speech-dataset-openslr-80Please visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Speech Dataset (OpenSLR-80) This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual OpenSLR dataset. For the complete multilingual dataset and additional information, please visit the original dataset repository of OpenSLR HuggingFace page. Original Source OpenSLR is a site devoted to hosting speech and language resources, such as training… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-openslr-80.audiotext-to-speech1K<n<10K7 likes116 downloads1y agoHugging Face21shunyalabs /myanmar-burmese-speech-datasetaudio1K<n<10K1 likes112 downloads1y agoHugging Face22kalixlouiis /Myanmar-English-general-text-translation 🇲🇲-🇬🇧 Myanmar-English General Text Translation Dataset 📚 Dataset Overview This dataset is a high-quality, parallel corpus designed for training robust and accurate Myanmar-English Machine Translation (MT) models. It focuses on General Domain texts, covering a wide range of everyday scenarios, literature, conversations, and descriptive narratives. Our primary goal in creating this dataset is to provide a clean, reliable resource to enhance the performance of… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/Myanmar-English-general-text-translation.texttranslation10K<n<100K12 likes112 downloads5mo agoHugging Face23puttatidam /myanmarnews-mya-clustering myanmarnews-mya-clustering Deduplicated copy of kornwtp/myanmarnews-mya-clustering, part of the SEA-BED data-quality work. Source dataset: kornwtp/myanmarnews-mya-clustering Deduplicated on: 2026-09-04 Task type: clustering Splits: train What changed A text appearing under more than one gold cluster is collapsed to ONE row carrying the competing cluster ids in a new conflict_label column (NULL elsewhere). labels on that row holds the first-seen value as a… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/myanmarnews-mya-clustering.text1K<n<10K0 likes93 downloads8d agoHugging Face24freococo /myanmar_typeset_dictionary_OCR Myanmar Typeset Dictionary OCR Dataset This is a synthetically generated, realistically formatted dataset modeling a Myanmar-Myanmar dictionary. It is designed for training and validating OCR models, Document Layout Analysis (DLA) pipelines, and structural key-value extraction models. The dataset contains a highly diverse set of pages containing multiple column flows, tabular glossaries, running headers/footers, realistic backgrounds, and dynamic typography (four fonts paired… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_typeset_dictionary_OCR.imageobject-detectionn<1K2 likes92 downloads2mo agoHugging Face25chuuhtetnaing /myanmar-ner-datasetPlease visit the GitHub repository for other Myanmar Language datasets. Myanmar NER Dataset A token classification dataset for Myanmar (Burmese) Named Entity Recognition (NER), formatted for sequence labeling tasks using BIO tagging scheme. 📓 Data Preparation Notebook: dataset-preparation.ipynb 📓 Fine-Tuning Notebook: myanmar-ner-fine-tuning.ipynb (based on the HuggingFace Token Classification Guide) Dataset Description This dataset contains Myanmar language text… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ner-dataset.texttoken-classification10K<n<100K0 likes85 downloads9mo agoHugging Face26chuuhtetnaing /myanmar-speech-dataset-google-fleursPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Speech Dataset (Google Fleurs) This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual Google Fleurs dataset. For the complete multilingual dataset and additional information, please visit the original dataset repository of Google Fleurs HuggingFace page. Original Source Fleurs is the speech version of the FLoRes machine translation benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-google-fleurs.audiotext-to-speech1K<n<10K0 likes80 downloads1y agoHugging Face27suwaimyo /khine-myanmarnews-mya-classification KhineMyanmarNews_mya_Classification Deduplicated copy of kornwtp/khine-myanmarnews-mya-classification. Splits split rows train 7,972 text1K<n<10K0 likes78 downloads27d agoHugging Face28KyawSu /Myanmar-Motivation-Translation-Dataset Motivational Speech Translations (English–Myanmar) Dataset Description This dataset contains sentence-level English–Myanmar translation pairs extracted from transcribed motivational speech / video content. Each example is a short segment (originally aligned to a timestamped subtitle line) paired with its Myanmar translation. Languages: English (eng_Latn), Myanmar (mya_Mymr) Total examples: 38,665 Source: Transcripts of motivational videos, segmented and… See the full description on the dataset page: https://huggingface.co/datasets/KyawSu/Myanmar-Motivation-Translation-Dataset.texttranslation10K<n<100K0 likes78 downloads25d agoHugging Face29freococo /myanmar-written-corpus Myanmar Written Corpus The Myanmar Written Corpus is a comprehensive collection of high-quality, but not fully CLEAN, written Myanmar text, designed to address the lack of large-scale, openly accessible resources for Myanmar Natural Language Processing (NLP). It is tailored to support various tasks such as text-to-speech (TTS), automatic speech recognition (ASR), translation, text generation, and more. This dataset serves as a critical resource for researchers and developers aiming… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-written-corpus.tabulartext-classification10M<n<100M4 likes76 downloads1y agoHugging Face30kkomyoeminaung /wikipedia-myanmar-qa Dataset Card for wikipedia-myanmar-qa Dataset Summary ဒီ dataset က wikipedia-myanmar-qa အတွက် ဖန်တီးထားတာပါ။ Languages Myanmar (my) / English (en) Dataset Structure Data Instances { "text": "နမူနာ စာသား", "label": "အညွှန်း" } Data Fields text: main content, label: optional. Data Splits Split Files train qa_chunk_0000.jsonl, qa_chunk_0001.jsonl, qa_chunk_0002.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/wikipedia-myanmar-qa.0 likes69 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.