datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
my-audio-appmy-audio-dataset
Dataset Card for "my-audio-dataset"
More Information needed
nug_myanmar_asr
366 Hours NUG Myanmar ASR Dataset
The NUG Myanmar ASR Dataset is the first large-scale open Burmese speech dataset — now expanded to over 521,476 audio-text pairs, totaling ~366 hours of clean, segmented audio. All data was collected from public-service educational broadcasts by the National Unity Government (NUG) of Myanmar and the FOEIM Academy.
This dataset is released under a CC0 1.0 Universal license — fully open and public domain. No attribution required.
🕊️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/nug_myanmar_asr.voa_myanmar_asr_audio_1
📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology.
Overview
This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.myanmar-speech-dataset-for-asrPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset for ASR
This dataset is a comprehensive collection of Myanmar language speech data specifically curated for Automatic Speech Recognition (ASR) task. It combines following datasets:
Myanmar Speech Dataset (Google Fleurs)
Myanmar Speech Dataset (OpenSLR-80)
Ko-Yin-Maung/mig-burmese-audio-transcription
By merging these complementary resources, this dataset provides a more robust… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-for-asr.mya-tiktok-asr-120h
Burmese TikTok ASR (121h)
A weakly-supervised Burmese (Myanmar, my) speech corpus: 168,852 short audio clips / 121.1 hours, segmented from 3,902 public TikTok videos and paired with the Burmese subtitles TikTok generates for those videos.
Intended for pre-training and fine-tuning Burmese ASR models (e.g. Whisper) in a language with very little open speech data.
⚠️ Read this first. The transcripts are machine-generated, not human-verified — see Labels are ASR output. Treat that… See the full description on the dataset page: https://huggingface.co/datasets/t7188409/mya-tiktok-asr-120h.myanmar-shopvoice
Myanmar Shopvoice Dataset (300 Sample)
This is a randomly sampled subset of 300 Myanmar shop voice audio chunks for Whisper fine-tuning and evaluation.
Dataset Details
Total Audio Chunks: 300
Audio Format: 16 kHz WAV, mono
Language: Myanmar (Burmese)
Features
audio: Audio feature (16 kHz WAV audio player)
transcription: Myanmar sentence transcription text
source_audio: Source continuous recording file name
source_line: Index line of transcript… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/myanmar-shopvoice.my-app-assets115hours_pvtv_myanmar_asr
115 Hours PVTV Myanmar ASR
This dataset contains 156,262 audio-transcript pairs of spoken Burmese, totaling approximately 115.31 hours. The audio segments were extracted from publicly available YouTube videos published by PVTV and aligned using subtitle timestamps.
Dedication
This dataset would not exist without the persistent voices of PVTV editors, journalists, narrators, and production teams, who continue to speak to the people under difficult conditions. PVTV is the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/115hours_pvtv_myanmar_asr.myanmar-speech-dataset-openslr-80Please visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset (OpenSLR-80)
This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual OpenSLR dataset.
For the complete multilingual dataset and additional information, please visit the original dataset repository
of OpenSLR HuggingFace page.
Original Source
OpenSLR is a site devoted to hosting speech and language resources, such as training… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-openslr-80.myanmar-burmese-speech-datasetmyanmar-speech-dataset-google-fleursPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset (Google Fleurs)
This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual Google Fleurs dataset.
For the complete multilingual dataset and additional information, please visit the original dataset repository
of Google Fleurs HuggingFace page.
Original Source
Fleurs is the speech version of the FLoRes machine translation benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-google-fleurs.BioVITAA2TRetrieval
BioVITAA2TRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Measures whether an audio-text model can name the wild animal it is hearing. Each query is a field recording of a single animal, and the model ranks 100 candidate taxa -- the recorded taxon plus 99 distractors -- represented by their taxon names in a 325-entry text index. A taxon scores the highest similarity over its own index entries and the 100 taxa are ranked by that score, so the reported taxon_top_k_accuracy is… See the full description on the dataset page: https://huggingface.co/datasets/myang333/BioVITAA2TRetrieval.BioVITAA2IRetrieval
BioVITAA2IRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Measures whether a model can connect an animal's call to its appearance without text as an intermediary. Each query is a field recording of a single animal, and the model ranks 100 candidate taxa -- the recorded taxon plus 99 distractors -- over an index of 2,835 wildlife photographs, where a taxon is represented by every photograph of that taxon. A taxon scores its best-matching photograph and the 100 taxa are… See the full description on the dataset page: https://huggingface.co/datasets/myang333/BioVITAA2IRetrieval.myanmar_cele_voices
Myanmar Celebrity Voices
A high-quality speech dataset extracted from the official TikTok channel of Myanmar Celebrity TV.
Myanmar Celebrity Voices is a collection of 69,781 short audio segments (≈46 hours total) derived from public TikTok videos by The Official TikTok Channel of Myanmar Celebrity TV — one of the most popular digital media platforms in Myanmar.
The source channel regularly publishes:
Interviews with Myanmar’s top movie actors and actresses
Behind-the-scenes… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_cele_voices.myanmar-english-accent-speech
Myanmar English Accent Speech (PVTV & FOEIM)
This dataset contains English speech by Myanmar speakers, collected from public videos published by PVTV and FOEIM — two media channels operating under the National Unity Government (NUG).
The clips reflect a wide range of spoken English contexts: interviews, announcements, sermons, and educational content. The speakers vary in tone, pace, and emotion — but all share the characteristic sound of Burmese-accented English.
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-accent-speech.cs-dialogue-dpomyaudioPlease note this dataset is private
Using the data
You can stream the data data loader:
myaudio = load_dataset(
"evageon/myaudio",
use_auth_token=os.environ["HG_USER_TOKEN"], # replace this with your access token
streaming=True)
Then you can iterate over the dataset
# replace test with validation or train depending on split you need
print(next(iter(myaudio["test"])))
outputs:
{'path': 'CD93A8FF-C3ED-4AD4-95A6-8363CCB93B90_spk-0001_seg-0024467:0025150.wav', 'audio':… See the full description on the dataset page: https://huggingface.co/datasets/evageon/myaudio.BioVITAI2ARetrieval
BioVITAI2ARetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Measures whether a photograph of an animal can retrieve that animal's call. Each query is a wildlife photograph, and the model ranks 100 candidate taxa -- the photographed taxon plus 99 distractors -- over an index of 1,024 field recordings, where a taxon is represented by every recording of that taxon. A taxon scores its best-matching recording and the 100 taxa are ranked by that score, so the reported… See the full description on the dataset page: https://huggingface.co/datasets/myang333/BioVITAI2ARetrieval.my_awsome_ASR_dataBioVITAT2ARetrieval
BioVITAT2ARetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Measures whether a text-audio model can retrieve the call of a named animal. Each query is the taxon name of one held-out species or genus, and the model ranks 100 candidate taxa -- the queried taxon plus 99 distractors -- over an index of 1,024 field recordings, where a taxon is represented by every recording of that taxon. A taxon scores its best-matching recording and the 100 taxa are ranked by that score, so the… See the full description on the dataset page: https://huggingface.co/datasets/myang333/BioVITAT2ARetrieval.rfa_shan_language_voices
RFA Shan Language Voices
This dataset contains 20.58 hours of audio in the Shan (Tai-Yai) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Shan language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into 5,047 manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/myandev/rfa_shan_language_voices.myanmar-speech-dataset-for-asrPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset for ASR
This dataset is a comprehensive collection of Myanmar language speech data specifically curated for Automatic Speech Recognition (ASR) task. It combines following datasets:
Myanmar Speech Dataset (Google Fleurs)
Myanmar Speech Dataset (OpenSLR-80)
Ko-Yin-Maung/mig-burmese-audio-transcription
By merging these complementary resources, this dataset provides a more robust… See the full description on the dataset page: https://huggingface.co/datasets/myandev/myanmar-speech-dataset-for-asr.raw_1hr_myanmar_asr_audio
🇲🇲 Raw 1-Hour Burmese ASR Audio Dataset
A 1-hour dataset of Burmese (Myanmar language) spoken audio clips with transcripts, curated from official public-service media broadcasts by PVTV Myanmar — the media voice of Myanmar’s National Unity Government (NUG).
This dataset is intended for automatic speech recognition (ASR) and Burmese speech-processing research.
➡️ Author: freococo➡️ License: MIT➡️ Language: Burmese (my)
📦 Dataset Summary
Duration: ~1 hour
Chunks:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/raw_1hr_myanmar_asr_audio.my_audio_dataset70hours_myanmar_audio_jw_bible
📖 JW Myanmar Bible Audio-Text Dataset (New World Translation) — 56 Books / 70 Hours
This dataset is a fully aligned parallel corpus of Myanmar-language Bible audio and faithful Burmese transcriptions, taken from the New World Translation (NWT) published by Jehovah’s Witnesses (JW.org).
It combines a previously released 10-book collection with a newer 46-book expansion, now covering 56 books in total. Across 917 chapters, this dataset brings the Word to life — in crystal-clear… See the full description on the dataset page: https://huggingface.co/datasets/freococo/70hours_myanmar_audio_jw_bible.arabic_myanmar_quran_voices
📖 Arabic-Myanmar Quran Voice Dataset
🕌 Overview
This dataset contains high-quality MP3 audio recordings of the entire Holy Qur’an with:
Arabic recitation of each verse
Followed immediately by its Myanmar (Burmese) translation
It is the first complete Arabic-Myanmar Quran audio interpretation of its kind publicly released in Myanmar. The goal is to make the Qur’an more accessible to:
Elderly persons
Blind or visually impaired people
Myanmar speakers who wish… See the full description on the dataset page: https://huggingface.co/datasets/freococo/arabic_myanmar_quran_voices.jw_myanmar_bible_10books_audio
📖 JW Myanmar Bible Audio-Text Dataset (New World Translation)
This dataset is a parallel corpus of audio recordings and Myanmar (Burmese) transcriptions of the New World Translation (NWT) Bible, published by Jehovah’s Witnesses (JW.org). It covers selected books and chapters from both the Old and New Testaments, focused on spoken-style Burmese with clear narration.
✨ Dataset Highlights
🎧 High-quality chapter-based audio in Myanmar (Zawgyi-free Unicode)
📝 Aligned… See the full description on the dataset page: https://huggingface.co/datasets/freococo/jw_myanmar_bible_10books_audio.google_myanmar_asr_voices
Google Myanmar ASR Dataset (WebDataset Version)
This repository provides a clean, user-friendly, and robust version of the Google Myanmar ASR Dataset, which is derived from the OpenSLR-80 Burmese Speech Corpus.
This version has been carefully re-processed into the WebDataset format. Each sample consists of a .wav audio file and a clean .json metadata file, packaged into sharded .tar archives. This format is highly efficient for large-scale training of ASR models.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/google_myanmar_asr_voices.myanmar_bible_audio_46books_jw_version
📖 JW Myanmar Bible Audio-Text Dataset (New World Translation) — 46 Books Edition
This dataset is a meticulously aligned, high-quality parallel corpus of Myanmar-language Bible audio recordings and faithful Burmese transcriptions, drawn from the New World Translation (NWT) published by Jehovah’s Witnesses (JW.org).
Spanning 46 books from both the Old and New Testaments, this release represents the largest open-source Burmese Bible audio-text dataset of its kind — crafted with care… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_bible_audio_46books_jw_version.
