CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chuuhtetnaing /myanmar-fineweb-2-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Fineweb2 Dataset A preprocessed subset of the Fineweb2 dataset containing only Myanmar language text, with consistent Unicode encoding. Dataset Description This dataset is derived from the Fineweb2 created by HuggingFaceFW. It contains only the Myanmar language portion of the original Fineweb2 dataset, with additional preprocessing to standardize text encoding. Filtered and Removed… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-fineweb-2-dataset.tabulartext-generation1M<n<10M0 likes247 downloads1y agoHugging Face02freococo /huggingface_myanmar_english_translation Cleaned & Sorted Myanmar-English Translation Dataset This dataset is a cleaned, Unicode-normalized, and sorted version of the Myanmar (Burmese) subset from the massive FineTranslations dataset. While the original dataset is excellent, Myanmar text on the web is often a mix of standard Unicode and the non-standard Zawgyi encoding. This repository fixes those encoding issues to provide a high-quality dataset for NLP tasks. Key Improvements in this Version Zawgyi… See the full description on the dataset page: https://huggingface.co/datasets/freococo/huggingface_myanmar_english_translation.texttranslation1M<n<10M1 likes243 downloads8mo agoHugging Face03jojo-ai-mst /Myanmar-Tuberculosis-Guidelines-Instructions Myanmar Tuberculosis Guidelines Instructions A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages. Authors: Min Si Thu, Khin Myat Noe Abstract Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.imagequestion-answering1K<n<10K1 likes146 downloads5mo agoHugging Face04freococo /myanmar-written-corpus Myanmar Written Corpus The Myanmar Written Corpus is a comprehensive collection of high-quality, but not fully CLEAN, written Myanmar text, designed to address the lack of large-scale, openly accessible resources for Myanmar Natural Language Processing (NLP). It is tailored to support various tasks such as text-to-speech (TTS), automatic speech recognition (ASR), translation, text generation, and more. This dataset serves as a critical resource for researchers and developers aiming… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-written-corpus.tabulartext-classification10M<n<100M4 likes68 downloads1y agoHugging Face05chuuhtetnaing /myanmar-c4-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar C4 Dataset A preprocessed subset of the C4 dataset containing only Myanmar language text, with consistent Unicode encoding. Dataset Description This dataset is derived from the Colossal Clean Crawled Corpus (C4) created by AllenAI. It contains only the Myanmar language portion of the original C4 dataset, with additional preprocessing to standardize text encoding. Preprocessing… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-c4-dataset.texttext-generation100K<n<1M0 likes62 downloads1y agoHugging Face06chuuhtetnaing /myanmar-culturax-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar CulturaX Dataset A preprocessed subset of the CulturaX dataset containing only Myanmar language text, with consistent Unicode encoding. Dataset Description This dataset is derived from the uonlp/CulturaX created by "The University of Oregon NLP Group". It contains only the Myanmar language portion of the original CulturaX dataset, with additional preprocessing to standardize text encoding.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-culturax-dataset.texttext-generation100K<n<1M0 likes57 downloads1y agoHugging Face07freococo /myanmar_spoken_corpus Myanmar Spoken Corpus (Version 1.0) Overview Myanmar Spoken Corpus is a high-quality, but not fully CLEAN, open dataset of spoken Myanmar sentences designed to support NLP and ASR applications. The dataset focuses on providing clean and structured spoken language data for advancing Myanmar language technology. Dataset Statistics Number of Rows: Local Parquet file: 16,020,011 rows Hugging Face Dataset Viewer: 15,728,640 rows File Size: 1.78 GB (Parquet… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_spoken_corpus.texttext-classification10M<n<100M3 likes56 downloads1y agoHugging Face08freococo /myanmar_quran_parallel_dataset_human_vs_ai Myanmar Quran Parallel Dataset: Human vs AI This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses. It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context. Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.texttranslation1K<n<10K0 likes53 downloads8mo agoHugging Face09chuuhtetnaing /myanmar-cc100-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar CC100 Dataset A preprocessed subset of the CC100 dataset containing only Myanmar language text, with consistent Unicode encoding. Dataset Description This dataset is derived from the statmt/cc100 created by "Statistical and Neural Machine Translation". It contains only the Myanmar language portion of the original CC100 dataset, with additional preprocessing to standardize text encoding.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-cc100-dataset.texttext-generation10M<n<100M0 likes51 downloads1y agoHugging Face10freococo /myanmar_qna_dataset Myanmar QnA Dataset v7 Language: Burmese (Myanmar)Total Entries: 22,783 QnA pairsTotal Sentences: ~ 466,330(Counted using the Myanmar sentence-ending symbol "။")License: CC0 1.0 (Public Domain) Description This dataset contains Myanmar-language question-answer pairs (QnA) generated with the assistance of ChatGPT-5 for question crafting with English and Gemini 3.0 Pro for Myanmar QnA generation. It is intended for research, AI training, and educational purposes. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_qna_dataset.tabularquestion-answering10K<n<100K0 likes46 downloads9mo agoHugging Face11freococo /tipitaka_myanmar_translation_books Myanmar Tipitaka Translation (60 Books) This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga. The texts have been converted into a clean, structured JSONL format, suitable for: Natural Language Processing (NLP) LLM Training & Fine-tuning Digital Humanities Research Dhamma Study Applications 📊 Dataset Statistics Total Books: 60 Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.tabulartext-generation100K<n<1M0 likes45 downloads8mo agoHugging Face12freococo /myawady-news-title-generation-dataset Myawady News Title Generation Dataset 🇲🇲 This dataset contains over 67,000 cleaned article titles extracted from the Myanmar state-run media outlet Myawady News Portal, intended for use in news title generation, text classification, and Myanmar NLP research. The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ). 🗂️ Dataset Overview Name:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-news-title-generation-dataset.texttext-generation10K<n<100K0 likes42 downloads1y agoHugging Face13freococo /myanmar-english-pali-dictionary Myanmar–English–Pali Dictionary Dataset Summary This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein). It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary. The dataset is intended for research and educational purposes, including but not limited to: Natural Language Processing (NLP) Machine Translation (MT) Lexicography Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.texttranslation10K<n<100K1 likes42 downloads8mo agoHugging Face14DatarrX /Myanmar-Written-Spoken-Parallel-Corpus Myanmar Written-Spoken Parallel Corpus (MWSPC) Dataset Description Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language. Curated by: Khant Sint Heinn (Kalix Louis) Organization: DatarrX | ဒေတာ-အက်စ် Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.texttext-generation1K<n<10K6 likes41 downloads4mo agoHugging Face15freococo /myawady-raw-dataset Myawady Raw News Corpus 🇲🇲 This dataset contains over 59,000 full-text Burmese news articles scraped from the Myawady News Portal, the official media outlet of the Myanmar military government. Unlike the title-only version, this dataset includes complete article content, with metadata fields such as category, publication date, and image URLs. It is intended for use in Myanmar NLP and AI research, including: 🧠 Language modeling 📰 Text summarization 🏷️ Named entity… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-raw-dataset.imagetext-classification10K<n<100K0 likes40 downloads1y agoHugging Face16DatarrX /moi-myanmar-articles-lines MOI Myanmar Articles Dataset - Lines (DatarrX/moi-myanmar-articles-lines) Dataset Description The MOI Myanmar Articles - Lines dataset is a derivative corpus created from the official articles published on the Ministry of Information (MOI) website of the Republic of the Union of Myanmar. Unlike the main dataset (moi-myanmar-articles), which contains full-length article texts, this dataset has been systematically split line-by-line (sentence-by-sentence). This… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/moi-myanmar-articles-lines.texttext-generation100K<n<1M4 likes37 downloads4mo agoHugging Face17DatarrX /pali-myanmar-dictionary-corpus Pali-Myanmar Dictionary Corpus (Instruction-Ready) Dataset Summary The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning. Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.texttranslation100K<n<1M6 likes36 downloads5mo agoHugging Face18DatarrX /myanmar-Wikipedia Myanmar Wikipedia Dataset (20260501) This dataset contains a cleaned, processed, and high-quality collection of Burmese Wikipedia articles, curated to serve as a robust foundation for Natural Language Processing (NLP) and Artificial Intelligence development in the Burmese language. 🏛️ About DatarrX DatarrX (Burmese: ဒေတာအက်စ်) is a non-profit open-source foundation dedicated to building a robust digital foundation for the Burmese language in the AI era. We believe that… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-Wikipedia.texttext-generation100K<n<1M5 likes31 downloads4mo agoHugging Face19DatarrX /moi-myanmar-articles MOI Myanmar Articles Dataset (DatarrX/moi-myanmar-articles) Dataset Description The MOI Myanmar Articles dataset is a collection of official articles extracted directly from the Ministry of Information (MOI) website of the Republic of the Union of Myanmar. This dataset is curated exclusively to foster the growth, research, and development of the Myanmar (Burmese) language within the fields of Natural Language Processing (NLP) and Machine Learning (ML).… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/moi-myanmar-articles.texttext-generation1K<n<10K4 likes29 downloads4mo agoHugging Face20jojo-ai-mst /Mpox-Myanmar Mpox-Myanmar Data Resources about Mpox(MonkeyPox) in Myanmar Mpox-Myanmar is a dataset about Mpox(MonkeyPox virus) in Burmese Language. Mpox(MonkeyPox) is becoming a wide alert virus. Thus, the information dataset about mpox will be built to build applications for knowledge and educate the public about mpox. The dataset is gathered from the following web pages. https://www.who.int/myanmar/emergencies/mpox https://www.moi.gov.mm/article/60588 Questions are annotated by Min Si Thu.… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Mpox-Myanmar.textquestion-answeringn<1K2 likes26 downloads2y agoHugging Face21freococo /jw_myanmar_bible_dataset 📖 JW Myanmar Bible Dataset (New World Translation) A richly structured, fully aligned dataset of the Myanmar (Burmese) Bible, translated by Jehovah's Witnesses from the New World Translation. This dataset includes 66 books, 1,189 chapters, and 31,078 verses, each with chapter-level URLs and verse-level breakdowns. ✨ Highlights - 📚 66 Canonical Books (Genesis to Revelation) - 🧩 1,189 chapters, 31,078 verses (as parsed from the JW.org Myanmar edition) - 🔗 Includes… See the full description on the dataset page: https://huggingface.co/datasets/freococo/jw_myanmar_bible_dataset.tabulartext-generation10K<n<100K0 likes25 downloads1y agoHugging Face22DatarrX /myX-myanmar-to-myanglish-corpus 📝 Myanglish (မြန်းဂလိ) corpus A parallel text corpus containing high-quality, human-curated pairs of native Burmese (Myanmar Unicode) text and its corresponding Myanglish (Burmese Romanization) transliterations. The initial high-quality release contains 2,121 fully approved rows designed to bridge the gap between formal script and the informal romanized phonetic typing structures widely used across social media, chat applications, and digital communication in Myanmar.… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-myanmar-to-myanglish-corpus.texttranslation1K<n<10K5 likes25 downloads4mo agoHugging Face23kingkaung /Quran_English_Myanmar_Parrelel_Corpus Quran English-Myanmar Parallel Corpus Description This dataset is a parallel corpus of the Quran, containing translations in English and Myanmar. It includes 6,237 verses (ayahs) from all chapters (surahs), aligned by their respective Surah and Ayah numbers. English Translation: Provided by Dr. Muhsin Khan and Dr. Hilali. Myanmar Translation: Translated by the Myanmar Quran Translation Committee, comprising religious and non-religious scholars, and later published by… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/Quran_English_Myanmar_Parrelel_Corpus.tabulartranslation1K<n<10K0 likes22 downloads2y agoHugging Face24chuuhtetnaing /myanmar-wikipedia-dataset Myanmar Wikipedia Dataset (Last Crawl Date: 25/03/2025) A collection of scraped Myanmar Wikipedia pages organized by category paths. Overview This dataset contains Myanmar Wikipedia articles scraped based on categorical organization. Unlike the official Wikimedia dataset (subset: 20231101.my), this repository provides an alternative approach to Myanmar Wikipedia content by following the categorical structure starting from the main entry page. Figure 1: The initial… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-wikipedia-dataset.texttext-generation100K<n<1M1 likes18 downloads2y agoHugging Face25chuuhtetnaing /myanmar-aya-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Aya Dataset A preprocessed subset of the Aya dataset containing only Myanmar language text. Dataset Description This dataset is derived from the Aya Dataset created by Cohere Labs. It contains only the Myanmar language portion of the original aya_dataset. Dataset Structure The dataset keep the same fields as the original aya_dataset dataset. Usage from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-aya-dataset.texttext-generationn<1K1 likes17 downloads1y agoHugging Face26SThandarTint /myanmar-written-corpus Myanmar Written Corpus The Myanmar Written Corpus is a comprehensive collection of high-quality, but not fully CLEAN, written Myanmar text, designed to address the lack of large-scale, openly accessible resources for Myanmar Natural Language Processing (NLP). It is tailored to support various tasks such as text-to-speech (TTS), automatic speech recognition (ASR), translation, text generation, and more. This dataset serves as a critical resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/SThandarTint/myanmar-written-corpus.tabulartext-classification10M<n<100M0 likes17 downloads1mo agoHugging Face27kalixlouiis /myanmar-general-numerals-corpus Myanmar General Numerals Corpus This dataset is a collection of Myanmar (Burmese) sentences specifically curated to include various numerals, numerical classifiers, and units of measurement. It covers a wide range of linguistic styles, from daily conversations to formal and royal usage. Dataset Details Creator: Kalix Louis (Khant Sint Heinn) Language: Myanmar (Burmese) Format: Plain Text (.txt) License: Apache-2.0 Source: Manually authored and curated by Kalix Louis.… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/myanmar-general-numerals-corpus.texttext-generationn<1K5 likes16 downloads5mo agoHugging Face28Myashka /SO-Python_QA-API_Usage-tanh_score Stack Overflow Python Q&A Dataset Description Filtered Python Q&A with API_Usage subcategory without: Images Links Blocks of code Scores in Q1-Q3 scaled with MaxAbsScaler. Tanh function applyed to joint Scores. tabulartext-generation1K<n<10K0 likes13 downloads3y agoHugging Face29DatarrX /myanmar-cities-qa Myanmar Cites Questions & Answers Dataset This dataset is an ongoing project dedicated to compiling comprehensive information about various cities in Myanmar. It converts geographical and cultural data—including locations, brief histories, local products, and notable landmarks—into a conversational Question & Answering (Q&A) format. Dataset Overview Content: Information about cities in Myanmar (e.g., location, history, local economy, and culture). Format: Q&A… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-cities-qa.texttext-generationn<1K4 likes10 downloads4mo agoHugging Face30freococo /1_pattern_10Kplus_myanmar_sentences 🧠 1_pattern_10Kplus_myanmar_sentences A structured dataset of 11,452 Myanmar sentences generated from a single, powerful grammar pattern: 📌 Pattern: Verb လည်း Verb တယ်။ A natural way to express repetition, emphasis, or causal connection in Myanmar. 💡 About the Dataset This dataset demonstrates how applying just one syntactic pattern to a curated verb list — combined with syllable-aware rules — can produce a high-quality corpus of over 10,000 valid… See the full description on the dataset page: https://huggingface.co/datasets/freococo/1_pattern_10Kplus_myanmar_sentences.texttext-generation10K<n<100K0 likes9 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.