datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
my-audio-appmyanmar-ocr-dataset-for-vlm
Myanmar OCR Dataset
A synthetic OCR dataset for fine-tuning Vision Language Models (VLMs) on Myanmar (Burmese) text recognition. It contains page images paired with their ground-truth text, sourced from chuuhtetnaing/mm-lib-book-dataset and rendered into page images using various Myanmar fonts.
Subsets
Subset
Description
Details
single_font
Rendered with Pyidaungsu font only
437 books
multi_font
Rendered with 76 Myanmar fonts
3 books (ပဋ္ဌာန်းမြတ်ဒေသနာ၊… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset-for-vlm.myanmar-ocr-dataset
Myanmar OCR Dataset
A synthetic dataset for training and fine-tuning Optical Character Recognition (OCR) models specifically for the Myanmar language.
Dataset Description
This dataset contains synthetically generated OCR images created specifically for Myanmar text recognition tasks. The images were generated using myanmar-ocr-data-generator, a fork of TextRecognitionDataGenerator with fixes for proper Myanmar character splitting.
Direct Download
Available… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset.Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.my-AI-vision-resultBioVITAT2IRetrieval
BioVITAT2IRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Measures whether a taxon name retrieves photographs of that taxon. Each query is the name of one held-out species or genus, and the model ranks 100 candidate taxa -- the queried taxon plus 99 distractors -- over an index of 2,835 wildlife photographs, where a taxon is represented by every photograph of that taxon. A taxon scores its best-matching photograph and the 100 taxa are ranked by that score, so the reported… See the full description on the dataset page: https://huggingface.co/datasets/myang333/BioVITAT2IRetrieval.myanmar_typeset_dictionary_OCR
Myanmar Typeset Dictionary OCR Dataset
This is a synthetically generated, realistically formatted dataset modeling a Myanmar-Myanmar dictionary. It is designed for training and validating OCR models, Document Layout Analysis (DLA) pipelines, and structural key-value extraction models.
The dataset contains a highly diverse set of pages containing multiple column flows, tabular glossaries, running headers/footers, realistic backgrounds, and dynamic typography (four fonts paired… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_typeset_dictionary_OCR.myanmar_complex_document_layouts
🇲🇲 Myanmar Complex Document Layouts
A large-scale, high-quality synthetic dataset containing 17,632 images of complex document layouts, dashboards, and infographics entirely in the Myanmar (Burmese) language.
This dataset is specifically designed to train and benchmark modern Computer Vision and multimodal LLMs on complex Myanmar typography, structured data, and diverse graphical layouts.
📊 Dataset Overview
Total Images: 17,632 high-resolution pages.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_complex_document_layouts.BioVITAA2IRetrieval
BioVITAA2IRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Measures whether a model can connect an animal's call to its appearance without text as an intermediary. Each query is a field recording of a single animal, and the model ranks 100 candidate taxa -- the recorded taxon plus 99 distractors -- over an index of 2,835 wildlife photographs, where a taxon is represented by every photograph of that taxon. A taxon scores its best-matching photograph and the 100 taxa are… See the full description on the dataset page: https://huggingface.co/datasets/myang333/BioVITAA2IRetrieval.autotrain-myanmar-kathein-festival-cartoonmyartworkMyanmarOCR-ImageText
🇲🇲 MyanmarOCR-ImageText Dataset
A clean and diverse Burmese Image-to-Text dataset for OCR and multimodal AI research.
📌 Summary
Total images: 41,664
Unique Burmese text entries: 1,139
Styles per text: 32 variations each
Resolution: 512 × 512
File types: PNG/JPG images
Dataset split: train only
Use cases: OCR, I2T (image-to-text), VLM pretrain/fine-tune
All text is Burmese only.No English words and no punctuation like: ? , ' " -
🔡 Text… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/MyanmarOCR-ImageText.BioVITAI2ARetrieval
BioVITAI2ARetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Measures whether a photograph of an animal can retrieve that animal's call. Each query is a wildlife photograph, and the model ranks 100 candidate taxa -- the photographed taxon plus 99 distractors -- over an index of 1,024 field recordings, where a taxon is represented by every recording of that taxon. A taxon scores its best-matching recording and the 100 taxa are ranked by that score, so the reported… See the full description on the dataset page: https://huggingface.co/datasets/myang333/BioVITAI2ARetrieval.english-myanmar-dictionary-dataset-EngMyanDictionaryPlease visit the GitHub repository for other Myanmar Language datasets.
English-Myanmar Dictionary Dataset
An English-Myanmar (Burmese) dictionary dataset containing 21,984 word entries with definitions, synonyms, and images.
Source
This dataset is derived from the EngMyanDictionary Android application by Soe Minn Minn. The original dictionary database (dictionary.db) was extracted and converted to the Hugging Face Dataset format.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/english-myanmar-dictionary-dataset-EngMyanDictionary.autotrain-myanmar-ancient-malemy-ai-portraitsmyawady-raw-dataset
Myawady Raw News Corpus 🇲🇲
This dataset contains over 59,000 full-text Burmese news articles scraped from the Myawady News Portal, the official media outlet of the Myanmar military government.
Unlike the title-only version, this dataset includes complete article content, with metadata fields such as category, publication date, and image URLs. It is intended for use in Myanmar NLP and AI research, including:
🧠 Language modeling
📰 Text summarization
🏷️ Named entity… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-raw-dataset.BioVITAI2TRetrieval
BioVITAI2TRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Measures fine-grained visual species recognition posed as retrieval. Each query is a wildlife photograph, and the model ranks 100 candidate taxa -- the photographed taxon plus 99 distractors -- represented by their taxon names in a 325-entry text index. A taxon scores the highest similarity over its own index entries and the 100 taxa are ranked by that score, so the reported taxon_top_k_accuracy is taxon-level rather… See the full description on the dataset page: https://huggingface.co/datasets/myang333/BioVITAI2TRetrieval.myanmar-synthetic-syllable-glyphs
🇲🇲 Myanmar Synthetic Syllable Glyphs (MSSG)
The Myanmar Synthetic Syllable Glyphs (MSSG) is a massive-scale, high-fidelity synthetic image dataset containing 14,295,552 heavily augmented glyph images (128x64 pixels, grayscale) representing the structural combinatorial matrix of the Burmese script.
Developed and engineered by Khant Sint Heinn (Kalix Louis), this core foundational dataset is officially published and maintained under DatarrX (Myanmar Open Source Organization… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-synthetic-syllable-glyphs.LS_SSDD_small_singlemyanmar_love_letters_ocr
Myanmar Love Letters OCR Dataset
Dataset Summary
This synthetic OCR dataset was generated utilizing the GEMINI 3.1 Flash model. To ensure high quality, the generated outputs underwent a level of human editing and verification. While we did not do exhaustive microscopic detail checking, human oversight was applied to ensure the Burmese text displays correctly, spelling is accurate, and the formatting perfectly aligns with what is required for a robust OCR (Optical… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_love_letters_ocr.myanmar-word-glyphs
🇲🇲 Myanmar Word Glyphs (MWG)
The Myanmar Word Glyphs (MWG) is a curated vocabulary-based synthetic image dataset containing 49,800 high-quality word/phrase glyph images (256x64 pixels, grayscale).
Developed and engineered by Khant Sint Heinn, this dataset is officially published and distributed under DatarrX (Myanmar Open Source Organization, NPO). While our sibling project—MSSG—explores the absolute mathematical grid of theoretical syllables, MWG is designed to map out authentic… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-word-glyphs.MyanmarOCR-ImageText
🇲🇲 MyanmarOCR-ImageText Dataset
A clean and diverse Burmese Image-to-Text dataset for OCR and multimodal AI research.
📌 Summary
Total images: 41,664
Unique Burmese text entries: 1,139
Styles per text: 32 variations each
Resolution: 512 × 512
File types: PNG/JPG images
Dataset split: train only
Use cases: OCR, I2T (image-to-text), VLM pretrain/fine-tune
All text is Burmese only.No English words and no punctuation like: ? , ' " -
🔡 Text… See the full description on the dataset page: https://huggingface.co/datasets/kme819870/MyanmarOCR-ImageText.myanmar-numeral-glyphs
🇲🇲 Myanmar Numeral Glyphs (MNG)
The Myanmar Numeral Glyphs (MNG) dataset is a curated, high-quality hybrid image dataset designed for optical character recognition (OCR) and image classification tasks targeting native Burmese digits (၀ to ၉). Released under DatarrX, this dataset bridges the gap in low-resource language resources by combining clean, human-annotated handwritten data with robust computer-generated font variations.
📌 Dataset Overview
Total Images: 1… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-numeral-glyphs.breast_cancer_cell
About Dataset
In this dataset, there are 58 H&E stained histopathology images used in breast cancer cell detection with associated ground truth data available. Routine histology uses the stain combination of hematoxylin and eosin, commonly referred to as H&E. These images are stained since most cells are essentially transparent, with little or no intrinsic pigment. Certain special stains, which bind selectively to particular components, are be used to identify biological structures… See the full description on the dataset page: https://huggingface.co/datasets/myatmo/breast_cancer_cell.my_awesome_new_bikemyagedatasetMy_anime_data
Dataset Card for "My_anime_data"
More Information needed
my-animals-groupmy-anima-lora-dataset
