CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lakshya1234 /my-audio-appaudion<1K0 likes1.1k downloads2d agoHugging Face02chuuhtetnaing /myanmar-ocr-dataset-for-vlm Myanmar OCR Dataset A synthetic OCR dataset for fine-tuning Vision Language Models (VLMs) on Myanmar (Burmese) text recognition. It contains page images paired with their ground-truth text, sourced from chuuhtetnaing/mm-lib-book-dataset and rendered into page images using various Myanmar fonts. Subsets Subset Description Details single_font Rendered with Pyidaungsu font only 437 books multi_font Rendered with 76 Myanmar fonts 3 books (ပဋ္ဌာန်းမြတ်ဒေသနာ၊… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset-for-vlm.image100K<n<1M2 likes566 downloads5mo agoHugging Face03apockill /myarm-8-synthetic-cube-to-cup-largeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 874, "total_frames": 421190, "total_tasks": 1, "total_videos": 1748, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:874" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/apockill/myarm-8-synthetic-cube-to-cup-large.tabularrobotics100K<n<1M0 likes420 downloads2y agoHugging Face04chuuhtetnaing /myanmar-ocr-dataset Myanmar OCR Dataset A synthetic dataset for training and fine-tuning Optical Character Recognition (OCR) models specifically for the Myanmar language. Dataset Description This dataset contains synthetically generated OCR images created specifically for Myanmar text recognition tasks. The images were generated using myanmar-ocr-data-generator, a fork of TextRecognitionDataGenerator with fixes for proper Myanmar character splitting. Direct Download Available… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset.imageimage-to-text1M<n<10M10 likes355 downloads1y agoHugging Face05trinhtuyen201 /my-audio-dataset Dataset Card for "my-audio-dataset" More Information needed audio10K<n<100K0 likes299 downloads2y agoHugging Face06justicedao /ipfs_myanmar_laws_ir Myanmar legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_myanmar_laws (revision b744df8f35b0f9eff35df4eeb57e3def96e80607) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Myanmar prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_myanmar_laws_ir.tabulartext-retrieval10K<n<100K0 likes291 downloads2d agoHugging Face07CatLee123 /My_Assets0 likes287 downloads1mo agoHugging Face08apockill /myarm-7-synthetic-cube-to-cupThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 277, "total_frames": 131449, "total_tasks": 1, "total_videos": 554, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:277" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/apockill/myarm-7-synthetic-cube-to-cup.tabularrobotics100K<n<1M0 likes278 downloads2y agoHugging Face09freococo /nug_myanmar_asr 366 Hours NUG Myanmar ASR Dataset The NUG Myanmar ASR Dataset is the first large-scale open Burmese speech dataset — now expanded to over 521,476 audio-text pairs, totaling ~366 hours of clean, segmented audio. All data was collected from public-service educational broadcasts by the National Unity Government (NUG) of Myanmar and the FOEIM Academy. This dataset is released under a CC0 1.0 Universal license — fully open and public domain. No attribution required. 🕊️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/nug_myanmar_asr.audioautomatic-speech-recognition100K<n<1M3 likes265 downloads1y agoHugging Face10chuuhtetnaing /myanmar-fineweb-2-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Fineweb2 Dataset A preprocessed subset of the Fineweb2 dataset containing only Myanmar language text, with consistent Unicode encoding. Dataset Description This dataset is derived from the Fineweb2 created by HuggingFaceFW. It contains only the Myanmar language portion of the original Fineweb2 dataset, with additional preprocessing to standardize text encoding. Filtered and Removed… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-fineweb-2-dataset.tabulartext-generation1M<n<10M0 likes247 downloads1y agoHugging Face11freococo /huggingface_myanmar_english_translation Cleaned & Sorted Myanmar-English Translation Dataset This dataset is a cleaned, Unicode-normalized, and sorted version of the Myanmar (Burmese) subset from the massive FineTranslations dataset. While the original dataset is excellent, Myanmar text on the web is often a mix of standard Unicode and the non-standard Zawgyi encoding. This repository fixes those encoding issues to provide a high-quality dataset for NLP tasks. Key Improvements in this Version Zawgyi… See the full description on the dataset page: https://huggingface.co/datasets/freococo/huggingface_myanmar_english_translation.texttranslation1M<n<10M1 likes243 downloads8mo agoHugging Face12freococo /voa_myanmar_voices VOA Myanmar Voices Burmese (Myanmar) speech corpus chunked into 20-second FLAC clips with transcripts. Derived from VOA Burmese radio broadcasts (public domain, U.S. 17 U.S.C. § 105). Contents File Size Description voa-00000000.tar … voa-00000238.tar 498 GB 149 WebDataset shards voa_transcripts.parquet 404 MB 1,424,257 (key, text) pairs voa_transcripts.jsonl 1.5 GB Same data, line-oriented Audio: 16 kHz mono FLAC, exactly 20.00 s per chunk… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_voices.automatic-speech-recognition1M<n<10M0 likes234 downloads3d agoHugging Face13apockill /myarm-6-put-cube-to-the-side-30fps-lowresThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "myarm", "total_episodes": 253, "total_frames": 82078, "total_tasks": 1, "total_videos": 506, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:253" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/apockill/myarm-6-put-cube-to-the-side-30fps-lowres.tabularrobotics10K<n<100K0 likes229 downloads2y agoHugging Face14freococo /voa_myanmar_asr_audio_1 📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology. Overview This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.audioautomatic-speech-recognition1M<n<10M1 likes219 downloads1y agoHugging Face15apockill /myarm-4-put-cube-in-basket-highres-two-cameraThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "myarm", "total_episodes": 101, "total_frames": 22893, "total_tasks": 1, "total_videos": 202, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:101" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/apockill/myarm-4-put-cube-in-basket-highres-two-camera.tabularrobotics10K<n<100K0 likes208 downloads2y agoHugging Face16chuuhtetnaing /myanmar-speech-dataset-for-asrPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Speech Dataset for ASR This dataset is a comprehensive collection of Myanmar language speech data specifically curated for Automatic Speech Recognition (ASR) task. It combines following datasets: Myanmar Speech Dataset (Google Fleurs) Myanmar Speech Dataset (OpenSLR-80) Ko-Yin-Maung/mig-burmese-audio-transcription By merging these complementary resources, this dataset provides a more robust… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-for-asr.audioautomatic-speech-recognition1K<n<10K2 likes194 downloads9mo agoHugging Face17t7188409 /mya-tiktok-asr-120h Burmese TikTok ASR (121h) A weakly-supervised Burmese (Myanmar, my) speech corpus: 168,852 short audio clips / 121.1 hours, segmented from 3,902 public TikTok videos and paired with the Burmese subtitles TikTok generates for those videos. Intended for pre-training and fine-tuning Burmese ASR models (e.g. Whisper) in a language with very little open speech data. ⚠️ Read this first. The transcripts are machine-generated, not human-verified — see Labels are ASR output. Treat that… See the full description on the dataset page: https://huggingface.co/datasets/t7188409/mya-tiktok-asr-120h.audioautomatic-speech-recognition100K<n<1M0 likes184 downloads1mo agoHugging Face18ayehninnkhine /myanmar_news Dataset Card for Myanmar_News Dataset Summary The Myanmar news dataset contains article snippets in four categories: Business, Entertainment, Politics, and Sport. These were collected in October 2017 by Aye Hninn Khine Languages Myanmar/Burmese language Dataset Structure Data Fields text - text from article category - a topic: Business, Entertainment, Politic, or Sport (note spellings) Data Splits One training set (8,116 total… See the full description on the dataset page: https://huggingface.co/datasets/ayehninnkhine/myanmar_news.texttext-classification1K<n<10K6 likes170 downloads2y agoHugging Face19thantzinphyo /myanmar-shopvoice Myanmar Shopvoice Dataset (300 Sample) This is a randomly sampled subset of 300 Myanmar shop voice audio chunks for Whisper fine-tuning and evaluation. Dataset Details Total Audio Chunks: 300 Audio Format: 16 kHz WAV, mono Language: Myanmar (Burmese) Features audio: Audio feature (16 kHz WAV audio player) transcription: Myanmar sentence transcription text source_audio: Source continuous recording file name source_line: Index line of transcript… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/myanmar-shopvoice.audioautomatic-speech-recognitionn<1K0 likes169 downloads1mo agoHugging Face20apockill /myarm-5-put-cube-to-the-sideThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "myarm", "total_episodes": 52, "total_frames": 13157, "total_tasks": 1, "total_videos": 104, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:52" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/apockill/myarm-5-put-cube-to-the-side.tabularrobotics10K<n<100K0 likes163 downloads2y agoHugging Face21fsh317574518 /my-app-assetsaudion<1K0 likes160 downloads5mo agoHugging Face22freococo /115hours_pvtv_myanmar_asr 115 Hours PVTV Myanmar ASR This dataset contains 156,262 audio-transcript pairs of spoken Burmese, totaling approximately 115.31 hours. The audio segments were extracted from publicly available YouTube videos published by PVTV and aligned using subtitle timestamps. Dedication This dataset would not exist without the persistent voices of PVTV editors, journalists, narrators, and production teams, who continue to speak to the people under difficult conditions. PVTV is the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/115hours_pvtv_myanmar_asr.audioautomatic-speech-recognition100K<n<1M2 likes156 downloads1y agoHugging Face23kkomyoeminaung /myanmar-native-text Dataset Card for myanmar-native-text Dataset Summary ဒီ dataset က myanmar-native-text အတွက် ဖန်တီးထားတာပါ။ Languages Myanmar (my) / English (en) Dataset Structure Data Instances { "text": "နမူနာ စာသား", "label": "အညွှန်း" } Data Fields text: main content, label: optional. Data Splits Split Files train data/native_00000.jsonl, data/native_00001.jsonl, data/native_00002.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/myanmar-native-text.1 likes150 downloads2mo agoHugging Face24endomorphosis /ipfs_myanmar_laws Myanmar Laws Legal corpus for Myanmar from the ipfs_datasets_py legal scrapers (endomorphosis). Instruments: 725 Last updated: 2026-09-16 Source: official gazettes and legal portals; see source_url per record. Schema: id, title, text, source_url, source_type, jurisdiction, country, language, eli, date, retrieved_at, license, law_status, identifiers. textn<1K0 likes147 downloads9d agoHugging Face25jojo-ai-mst /Myanmar-Tuberculosis-Guidelines-Instructions Myanmar Tuberculosis Guidelines Instructions A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages. Authors: Min Si Thu, Khin Myat Noe Abstract Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.imagequestion-answering1K<n<10K1 likes146 downloads5mo agoHugging Face26kkomyoeminaung /myanmar-books-longform-1024 Dataset Card for myanmar-books-longform-1024 Dataset Summary ဒီ dataset က myanmar-books-longform-1024 အတွက် ဖန်တီးထားတာပါ။ Languages Myanmar (my) / English (en) Dataset Structure Data Instances { "text": "နမူနာ စာသား", "label": "အညွှန်း" } Data Fields text: main content, label: optional. Data Splits Split Files train books_chunk_0000.jsonl, books_chunk_0001.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/myanmar-books-longform-1024.0 likes146 downloads2mo agoHugging Face27kornwtp /thura-myanmar-news-mya-classificationref: https://huggingface.co/datasets/ThuraAung1601/myanmar_news text10K<n<100K1 likes141 downloads2y agoHugging Face28akhtet /myanmar-xnli Dataset Card for myXNLI Dataset Summary The myXNLI corpus extends XNLI corpus with Myanmar (Burmese) language. For myXNLI, we human-translated all 7,500 sentence pairs from XNLI English dev and test sets into Myanmar. The NLI and Genre labels from English dev and test sets are also reused for the Myanmar datasets. The dataset also includes the NLI training data in Myanmar which is created by machine-translating the MultiNLI training data from English into Myanmar. Similar… See the full description on the dataset page: https://huggingface.co/datasets/akhtet/myanmar-xnli.texttext-classification100K<n<1M7 likes139 downloads1y agoHugging Face29cuonguyenphu /my-AI-vision-resultimagen<1K1 likes133 downloads27d agoHugging Face30KhaMinh /my-acne-dataset0 likes128 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.