CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lakshya1234 /my-audio-appaudion<1K0 likes1.1k downloads3d agoHugging Face02chuuhtetnaing /myanmar-ocr-dataset-for-vlm Myanmar OCR Dataset A synthetic OCR dataset for fine-tuning Vision Language Models (VLMs) on Myanmar (Burmese) text recognition. It contains page images paired with their ground-truth text, sourced from chuuhtetnaing/mm-lib-book-dataset and rendered into page images using various Myanmar fonts. Subsets Subset Description Details single_font Rendered with Pyidaungsu font only 437 books multi_font Rendered with 76 Myanmar fonts 3 books (ပဋ္ဌာန်းမြတ်ဒေသနာ၊… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset-for-vlm.image100K<n<1M2 likes566 downloads5mo agoHugging Face03chuuhtetnaing /myanmar-ocr-dataset Myanmar OCR Dataset A synthetic dataset for training and fine-tuning Optical Character Recognition (OCR) models specifically for the Myanmar language. Dataset Description This dataset contains synthetically generated OCR images created specifically for Myanmar text recognition tasks. The images were generated using myanmar-ocr-data-generator, a fork of TextRecognitionDataGenerator with fixes for proper Myanmar character splitting. Direct Download Available… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset.imageimage-to-text1M<n<10M10 likes355 downloads1y agoHugging Face04jojo-ai-mst /Myanmar-Tuberculosis-Guidelines-Instructions Myanmar Tuberculosis Guidelines Instructions A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages. Authors: Min Si Thu, Khin Myat Noe Abstract Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.imagequestion-answering1K<n<10K1 likes146 downloads5mo agoHugging Face05cuonguyenphu /my-AI-vision-resultimagen<1K1 likes133 downloads27d agoHugging Face06myang333 /BioVITAT2IRetrieval BioVITAT2IRetrieval An MTEB dataset Massive Text Embedding Benchmark Measures whether a taxon name retrieves photographs of that taxon. Each query is the name of one held-out species or genus, and the model ranks 100 candidate taxa -- the queried taxon plus 99 distractors -- over an index of 2,835 wildlife photographs, where a taxon is represented by every photograph of that taxon. A taxon scores its best-matching photograph and the 100 taxa are ranked by that score, so the reported… See the full description on the dataset page: https://huggingface.co/datasets/myang333/BioVITAT2IRetrieval.imageother10K<n<100K0 likes91 downloads1mo agoHugging Face07freococo /myanmar_typeset_dictionary_OCR Myanmar Typeset Dictionary OCR Dataset This is a synthetically generated, realistically formatted dataset modeling a Myanmar-Myanmar dictionary. It is designed for training and validating OCR models, Document Layout Analysis (DLA) pipelines, and structural key-value extraction models. The dataset contains a highly diverse set of pages containing multiple column flows, tabular glossaries, running headers/footers, realistic backgrounds, and dynamic typography (four fonts paired… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_typeset_dictionary_OCR.imageobject-detectionn<1K2 likes69 downloads2mo agoHugging Face08freococo /myanmar_complex_document_layouts 🇲🇲 Myanmar Complex Document Layouts A large-scale, high-quality synthetic dataset containing 17,632 images of complex document layouts, dashboards, and infographics entirely in the Myanmar (Burmese) language. This dataset is specifically designed to train and benchmark modern Computer Vision and multimodal LLMs on complex Myanmar typography, structured data, and diverse graphical layouts. 📊 Dataset Overview Total Images: 17,632 high-resolution pages.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_complex_document_layouts.image10K<n<100K1 likes67 downloads2mo agoHugging Face09myang333 /BioVITAA2IRetrieval BioVITAA2IRetrieval An MTEB dataset Massive Text Embedding Benchmark Measures whether a model can connect an animal's call to its appearance without text as an intermediary. Each query is a field recording of a single animal, and the model ranks 100 candidate taxa -- the recorded taxon plus 99 distractors -- over an index of 2,835 wildlife photographs, where a taxon is represented by every photograph of that taxon. A taxon scores its best-matching photograph and the 100 taxa are… See the full description on the dataset page: https://huggingface.co/datasets/myang333/BioVITAA2IRetrieval.audioother10K<n<100K0 likes65 downloads1mo agoHugging Face10waloneai /autotrain-myanmar-kathein-festival-cartoonimagen<1K0 likes55 downloads2y agoHugging Face11tripletsahurr /myartworkimagen<1K1 likes50 downloads4mo agoHugging Face12kalixlouiis /MyanmarOCR-ImageText 🇲🇲 MyanmarOCR-ImageText Dataset A clean and diverse Burmese Image-to-Text dataset for OCR and multimodal AI research. 📌 Summary Total images: 41,664 Unique Burmese text entries: 1,139 Styles per text: 32 variations each Resolution: 512 × 512 File types: PNG/JPG images Dataset split: train only Use cases: OCR, I2T (image-to-text), VLM pretrain/fine-tune All text is Burmese only.No English words and no punctuation like: ? , ' " - 🔡 Text… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/MyanmarOCR-ImageText.imagefeature-extraction10K<n<100K9 likes49 downloads5mo agoHugging Face13myang333 /BioVITAI2ARetrieval BioVITAI2ARetrieval An MTEB dataset Massive Text Embedding Benchmark Measures whether a photograph of an animal can retrieve that animal's call. Each query is a wildlife photograph, and the model ranks 100 candidate taxa -- the photographed taxon plus 99 distractors -- over an index of 1,024 field recordings, where a taxon is represented by every recording of that taxon. A taxon scores its best-matching recording and the 100 taxa are ranked by that score, so the reported… See the full description on the dataset page: https://huggingface.co/datasets/myang333/BioVITAI2ARetrieval.audioother10K<n<100K0 likes47 downloads1mo agoHugging Face14chuuhtetnaing /english-myanmar-dictionary-dataset-EngMyanDictionaryPlease visit the GitHub repository for other Myanmar Language datasets. English-Myanmar Dictionary Dataset An English-Myanmar (Burmese) dictionary dataset containing 21,984 word entries with definitions, synonyms, and images. Source This dataset is derived from the EngMyanDictionary Android application by Soe Minn Minn. The original dictionary database (dictionary.db) was extracted and converted to the Hugging Face Dataset format. Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/english-myanmar-dictionary-dataset-EngMyanDictionary.image10K<n<100K0 likes43 downloads9mo agoHugging Face15waloneai /autotrain-myanmar-ancient-maleimagen<1K0 likes41 downloads2y agoHugging Face16IndividualGamer /my-ai-portraitsimagen<1K0 likes41 downloads2mo agoHugging Face17freococo /myawady-raw-dataset Myawady Raw News Corpus 🇲🇲 This dataset contains over 59,000 full-text Burmese news articles scraped from the Myawady News Portal, the official media outlet of the Myanmar military government. Unlike the title-only version, this dataset includes complete article content, with metadata fields such as category, publication date, and image URLs. It is intended for use in Myanmar NLP and AI research, including: 🧠 Language modeling 📰 Text summarization 🏷️ Named entity… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-raw-dataset.imagetext-classification10K<n<100K0 likes40 downloads1y agoHugging Face18myang333 /BioVITAI2TRetrieval BioVITAI2TRetrieval An MTEB dataset Massive Text Embedding Benchmark Measures fine-grained visual species recognition posed as retrieval. Each query is a wildlife photograph, and the model ranks 100 candidate taxa -- the photographed taxon plus 99 distractors -- represented by their taxon names in a 325-entry text index. A taxon scores the highest similarity over its own index entries and the 100 taxa are ranked by that score, so the reported taxon_top_k_accuracy is taxon-level rather… See the full description on the dataset page: https://huggingface.co/datasets/myang333/BioVITAI2TRetrieval.imageother1K<n<10K0 likes37 downloads1mo agoHugging Face19DatarrX /myanmar-synthetic-syllable-glyphs 🇲🇲 Myanmar Synthetic Syllable Glyphs (MSSG) The Myanmar Synthetic Syllable Glyphs (MSSG) is a massive-scale, high-fidelity synthetic image dataset containing 14,295,552 heavily augmented glyph images (128x64 pixels, grayscale) representing the structural combinatorial matrix of the Burmese script. Developed and engineered by Khant Sint Heinn (Kalix Louis), this core foundational dataset is officially published and maintained under DatarrX (Myanmar Open Source Organization… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-synthetic-syllable-glyphs.imageimage-classification10M<n<100M6 likes27 downloads4mo agoHugging Face20MyAuroralPlace /LS_SSDD_small_singleimage1K<n<10K0 likes25 downloads3y agoHugging Face21freococo /myanmar_love_letters_ocr Myanmar Love Letters OCR Dataset Dataset Summary This synthetic OCR dataset was generated utilizing the GEMINI 3.1 Flash model. To ensure high quality, the generated outputs underwent a level of human editing and verification. While we did not do exhaustive microscopic detail checking, human oversight was applied to ensure the Burmese text displays correctly, spelling is accurate, and the formatting perfectly aligns with what is required for a robust OCR (Optical… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_love_letters_ocr.imageobject-detectionn<1K0 likes19 downloads2mo agoHugging Face22DatarrX /myanmar-word-glyphs 🇲🇲 Myanmar Word Glyphs (MWG) The Myanmar Word Glyphs (MWG) is a curated vocabulary-based synthetic image dataset containing 49,800 high-quality word/phrase glyph images (256x64 pixels, grayscale). Developed and engineered by Khant Sint Heinn, this dataset is officially published and distributed under DatarrX (Myanmar Open Source Organization, NPO). While our sibling project—MSSG—explores the absolute mathematical grid of theoretical syllables, MWG is designed to map out authentic… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-word-glyphs.imageimage-classification10K<n<100K5 likes16 downloads4mo agoHugging Face23kme819870 /MyanmarOCR-ImageText 🇲🇲 MyanmarOCR-ImageText Dataset A clean and diverse Burmese Image-to-Text dataset for OCR and multimodal AI research. 📌 Summary Total images: 41,664 Unique Burmese text entries: 1,139 Styles per text: 32 variations each Resolution: 512 × 512 File types: PNG/JPG images Dataset split: train only Use cases: OCR, I2T (image-to-text), VLM pretrain/fine-tune All text is Burmese only.No English words and no punctuation like: ? , ' " - 🔡 Text… See the full description on the dataset page: https://huggingface.co/datasets/kme819870/MyanmarOCR-ImageText.imagefeature-extraction10K<n<100K1 likes15 downloads7mo agoHugging Face24DatarrX /myanmar-numeral-glyphs 🇲🇲 Myanmar Numeral Glyphs (MNG) The Myanmar Numeral Glyphs (MNG) dataset is a curated, high-quality hybrid image dataset designed for optical character recognition (OCR) and image classification tasks targeting native Burmese digits (၀ to ၉). Released under DatarrX, this dataset bridges the gap in low-resource language resources by combining clean, human-annotated handwritten data with robust computer-generated font variations. 📌 Dataset Overview Total Images: 1… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-numeral-glyphs.imageimage-classification1K<n<10K5 likes13 downloads4mo agoHugging Face25myatmo /breast_cancer_cell About Dataset In this dataset, there are 58 H&E stained histopathology images used in breast cancer cell detection with associated ground truth data available. Routine histology uses the stain combination of hematoxylin and eosin, commonly referred to as H&E. These images are stained since most cells are essentially transparent, with little or no intrinsic pigment. Certain special stains, which bind selectively to particular components, are be used to identify biological structures… See the full description on the dataset page: https://huggingface.co/datasets/myatmo/breast_cancer_cell.imagen<1K0 likes9 downloads2y agoHugging Face26RuudVelo /my_awesome_new_bikeimagen<1K0 likes8 downloads4y agoHugging Face27yudem /myagedatasetimagen<1K0 likes8 downloads4y agoHugging Face28Muhammad89 /My_anime_data Dataset Card for "My_anime_data" More Information needed image1K<n<10K0 likes8 downloads3y agoHugging Face29RDmango /my-animals-groupimagen<1K0 likes4 downloads3y agoHugging Face30YureDaaed /my-anima-lora-datasetimagen<1K0 likes4 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.