datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nug_myanmar_asr
366 Hours NUG Myanmar ASR Dataset
The NUG Myanmar ASR Dataset is the first large-scale open Burmese speech dataset — now expanded to over 521,476 audio-text pairs, totaling ~366 hours of clean, segmented audio. All data was collected from public-service educational broadcasts by the National Unity Government (NUG) of Myanmar and the FOEIM Academy.
This dataset is released under a CC0 1.0 Universal license — fully open and public domain. No attribution required.
🕊️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/nug_myanmar_asr.voa_myanmar_voices
VOA Myanmar Voices
Burmese (Myanmar) speech corpus chunked into 20-second FLAC clips with transcripts. Derived from VOA Burmese radio broadcasts (public domain, U.S. 17 U.S.C. § 105).
Contents
File
Size
Description
voa-00000000.tar … voa-00000238.tar
498 GB
149 WebDataset shards
voa_transcripts.parquet
404 MB
1,424,257 (key, text) pairs
voa_transcripts.jsonl
1.5 GB
Same data, line-oriented
Audio: 16 kHz mono FLAC, exactly 20.00 s per chunk… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_voices.voa_myanmar_asr_audio_1
📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology.
Overview
This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.myanmar-speech-dataset-for-asrPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset for ASR
This dataset is a comprehensive collection of Myanmar language speech data specifically curated for Automatic Speech Recognition (ASR) task. It combines following datasets:
Myanmar Speech Dataset (Google Fleurs)
Myanmar Speech Dataset (OpenSLR-80)
Ko-Yin-Maung/mig-burmese-audio-transcription
By merging these complementary resources, this dataset provides a more robust… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-for-asr.mya-tiktok-asr-120h
Burmese TikTok ASR (121h)
A weakly-supervised Burmese (Myanmar, my) speech corpus: 168,852 short audio clips / 121.1 hours, segmented from 3,902 public TikTok videos and paired with the Burmese subtitles TikTok generates for those videos.
Intended for pre-training and fine-tuning Burmese ASR models (e.g. Whisper) in a language with very little open speech data.
⚠️ Read this first. The transcripts are machine-generated, not human-verified — see Labels are ASR output. Treat that… See the full description on the dataset page: https://huggingface.co/datasets/t7188409/mya-tiktok-asr-120h.myanmar-shopvoice
Myanmar Shopvoice Dataset (300 Sample)
This is a randomly sampled subset of 300 Myanmar shop voice audio chunks for Whisper fine-tuning and evaluation.
Dataset Details
Total Audio Chunks: 300
Audio Format: 16 kHz WAV, mono
Language: Myanmar (Burmese)
Features
audio: Audio feature (16 kHz WAV audio player)
transcription: Myanmar sentence transcription text
source_audio: Source continuous recording file name
source_line: Index line of transcript… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/myanmar-shopvoice.115hours_pvtv_myanmar_asr
115 Hours PVTV Myanmar ASR
This dataset contains 156,262 audio-transcript pairs of spoken Burmese, totaling approximately 115.31 hours. The audio segments were extracted from publicly available YouTube videos published by PVTV and aligned using subtitle timestamps.
Dedication
This dataset would not exist without the persistent voices of PVTV editors, journalists, narrators, and production teams, who continue to speak to the people under difficult conditions. PVTV is the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/115hours_pvtv_myanmar_asr.myanmar-speech-dataset-openslr-80Please visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset (OpenSLR-80)
This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual OpenSLR dataset.
For the complete multilingual dataset and additional information, please visit the original dataset repository
of OpenSLR HuggingFace page.
Original Source
OpenSLR is a site devoted to hosting speech and language resources, such as training… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-openslr-80.myanmar-speech-dataset-google-fleursPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset (Google Fleurs)
This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual Google Fleurs dataset.
For the complete multilingual dataset and additional information, please visit the original dataset repository
of Google Fleurs HuggingFace page.
Original Source
Fleurs is the speech version of the FLoRes machine translation benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-google-fleurs.myanmar_cele_voices
Myanmar Celebrity Voices
A high-quality speech dataset extracted from the official TikTok channel of Myanmar Celebrity TV.
Myanmar Celebrity Voices is a collection of 69,781 short audio segments (≈46 hours total) derived from public TikTok videos by The Official TikTok Channel of Myanmar Celebrity TV — one of the most popular digital media platforms in Myanmar.
The source channel regularly publishes:
Interviews with Myanmar’s top movie actors and actresses
Behind-the-scenes… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_cele_voices.myanmar-english-accent-speech
Myanmar English Accent Speech (PVTV & FOEIM)
This dataset contains English speech by Myanmar speakers, collected from public videos published by PVTV and FOEIM — two media channels operating under the National Unity Government (NUG).
The clips reflect a wide range of spoken English contexts: interviews, announcements, sermons, and educational content. The speakers vary in tone, pace, and emotion — but all share the characteristic sound of Burmese-accented English.
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-accent-speech.shan_language_asr_voices
⭐ A Voice for the Shan People: The SHAN Herald Agency Audio Archive
This is an extensive 306-hour audio dataset of the Shan (Tai-Yai) language, meticulously curated from the public broadcasts of the Shan Herald Agency for News (SHAN). For over two decades, SHAN has been a vital, independent voice for the people of Shan State, Myanmar, chronicling their stories of culture, politics, and the enduring struggle for federal democracy.
This archive stands as one of the largest… See the full description on the dataset page: https://huggingface.co/datasets/myandev/shan_language_asr_voices.rfa_shan_language_voices
RFA Shan Language Voices
This dataset contains 20.58 hours of audio in the Shan (Tai-Yai) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Shan language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into 5,047 manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/myandev/rfa_shan_language_voices.myanmar-speech-dataset-for-asrPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset for ASR
This dataset is a comprehensive collection of Myanmar language speech data specifically curated for Automatic Speech Recognition (ASR) task. It combines following datasets:
Myanmar Speech Dataset (Google Fleurs)
Myanmar Speech Dataset (OpenSLR-80)
Ko-Yin-Maung/mig-burmese-audio-transcription
By merging these complementary resources, this dataset provides a more robust… See the full description on the dataset page: https://huggingface.co/datasets/myandev/myanmar-speech-dataset-for-asr.raw_1hr_myanmar_asr_audio
🇲🇲 Raw 1-Hour Burmese ASR Audio Dataset
A 1-hour dataset of Burmese (Myanmar language) spoken audio clips with transcripts, curated from official public-service media broadcasts by PVTV Myanmar — the media voice of Myanmar’s National Unity Government (NUG).
This dataset is intended for automatic speech recognition (ASR) and Burmese speech-processing research.
➡️ Author: freococo➡️ License: MIT➡️ Language: Burmese (my)
📦 Dataset Summary
Duration: ~1 hour
Chunks:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/raw_1hr_myanmar_asr_audio.70hours_myanmar_audio_jw_bible
📖 JW Myanmar Bible Audio-Text Dataset (New World Translation) — 56 Books / 70 Hours
This dataset is a fully aligned parallel corpus of Myanmar-language Bible audio and faithful Burmese transcriptions, taken from the New World Translation (NWT) published by Jehovah’s Witnesses (JW.org).
It combines a previously released 10-book collection with a newer 46-book expansion, now covering 56 books in total. Across 917 chapters, this dataset brings the Word to life — in crystal-clear… See the full description on the dataset page: https://huggingface.co/datasets/freococo/70hours_myanmar_audio_jw_bible.arabic_myanmar_quran_voices
📖 Arabic-Myanmar Quran Voice Dataset
🕌 Overview
This dataset contains high-quality MP3 audio recordings of the entire Holy Qur’an with:
Arabic recitation of each verse
Followed immediately by its Myanmar (Burmese) translation
It is the first complete Arabic-Myanmar Quran audio interpretation of its kind publicly released in Myanmar. The goal is to make the Qur’an more accessible to:
Elderly persons
Blind or visually impaired people
Myanmar speakers who wish… See the full description on the dataset page: https://huggingface.co/datasets/freococo/arabic_myanmar_quran_voices.jw_myanmar_bible_10books_audio
📖 JW Myanmar Bible Audio-Text Dataset (New World Translation)
This dataset is a parallel corpus of audio recordings and Myanmar (Burmese) transcriptions of the New World Translation (NWT) Bible, published by Jehovah’s Witnesses (JW.org). It covers selected books and chapters from both the Old and New Testaments, focused on spoken-style Burmese with clear narration.
✨ Dataset Highlights
🎧 High-quality chapter-based audio in Myanmar (Zawgyi-free Unicode)
📝 Aligned… See the full description on the dataset page: https://huggingface.co/datasets/freococo/jw_myanmar_bible_10books_audio.google_myanmar_asr_voices
Google Myanmar ASR Dataset (WebDataset Version)
This repository provides a clean, user-friendly, and robust version of the Google Myanmar ASR Dataset, which is derived from the OpenSLR-80 Burmese Speech Corpus.
This version has been carefully re-processed into the WebDataset format. Each sample consists of a .wav audio file and a clean .json metadata file, packaged into sharded .tar archives. This format is highly efficient for large-scale training of ASR models.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/google_myanmar_asr_voices.myanmar_bible_audio_46books_jw_version
📖 JW Myanmar Bible Audio-Text Dataset (New World Translation) — 46 Books Edition
This dataset is a meticulously aligned, high-quality parallel corpus of Myanmar-language Bible audio recordings and faithful Burmese transcriptions, drawn from the New World Translation (NWT) published by Jehovah’s Witnesses (JW.org).
Spanning 46 books from both the Old and New Testaments, this release represents the largest open-source Burmese Bible audio-text dataset of its kind — crafted with care… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_bible_audio_46books_jw_version.2hr_myanmar_asr_raw_audio
🇲🇲 Raw 2-Hour Burmese ASR Audio Dataset
A ~2-hour Burmese (Myanmar language) ASR dataset featuring 1,612 audio clips with aligned transcripts, curated from official public-service educational broadcasts by FOEIM Academy — a civic media arm of FOEIM.ORG, operating under the Myanmar National Unity Government (NUG).
This dataset is MIT-licensed as a public good — a shared asset for the Burmese-speaking world. It serves speech technology, education, and cultural preservation efforts… See the full description on the dataset page: https://huggingface.co/datasets/freococo/2hr_myanmar_asr_raw_audio.3hr_myanmar_asr_raw_audio
📚 3-Hour Burmese Speech Dataset from FOEIM Academy (ASR-ready)
This is a curated ~3-hour dataset of Burmese-language audio-transcript pairs derived from the official public-service educational media of FOEIM Academy, a civic platform affiliated with FOEIM.ORG.
It is structured for fine-grained automatic speech recognition (ASR) training and testing.All data is aligned from timestamped subtitle files (.srt) and segmented into high-quality .mp3 mono files with aligned transcripts.
➡️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/3hr_myanmar_asr_raw_audio.voa_myanmar_asr_audio_2⸻
Overview
This dataset was created by scraping and segmenting over 4,000 episodes of the VOA Burmese morning radio program. From that archive, 3,687 MP3 files were extracted and processed. This dataset contains sentence-level audio chunks suitable for ASR and speech-related model training.
The current release (voa_batch_001.tar and voa_batch_003.tar) contains a combined total of ~152,300 sentence-level audio chunks derived from the first 420 MP3 files in the archive, totaling… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_2.myanmar_bible_voices
🇲🇲 Myanmar Bible Speech 44k Corpus
A high-quality, sentence-level segmented Myanmar (Burmese) speech dataset containing 44,371 utterances (~60.1 hours) extracted from 917 Bible chapters.
All audio clips are aligned at the word and sentence level using Meta's MMS (Massively Multilingual Speech) Forced Aligner (mms-300m-1130-forced-aligner) and tokenized using mmdt-tokenizer.
📊 Dataset Summary
Total Utterances: 44,371
Total Duration: ~60.06 hours
Average… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_bible_voices.myanmarPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset (OpenSLR-80)
This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual OpenSLR dataset.
For the complete multilingual dataset and additional information, please visit the original dataset repository
of OpenSLR HuggingFace page.
Original Source
OpenSLR is a site devoted to hosting speech and language resources, such as training… See the full description on the dataset page: https://huggingface.co/datasets/kenmyz/myanmar.my-audio-dataset
Audio Transcription Dataset
This dataset contains audio file paths and their corresponding transcriptions for automatic speech recognition (ASR) tasks.
Dataset Description
This dataset is structured for audio transcription tasks with two main columns:
audio: Audio file paths (type: audio)
transcript: Text transcriptions (type: text)
Files
audio_dataset.csv: Main dataset file containing audio paths and transcriptions
Dataset Structure
audio… See the full description on the dataset page: https://huggingface.co/datasets/Aashish17405/my-audio-dataset.
