datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
burmese-speech-refined-openslr-80
Burmese Speech Refined OpenSLR-80
Summary
This dataset is a speech dataset developed based on the original OpenSLR Dataset (SLR80), with the text and audio data carefully reviewed and further refined for Burmese language applications.
In the original OpenSLR Dataset, the Burmese text was transcribed based on how the words were pronounced in the corresponding audio recordings. In this dataset, the original audio and text data were used as a reference, and the text… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/burmese-speech-refined-openslr-80.burmese-synthetic-speech-corpus
Burmese Synthetic Speech Corpus (DatarrX/burmese-synthetic-speech-corpus)
Overview
The Burmese Synthetic Speech Corpus is a high-fidelity, manually curated audio dataset specifically designed to advance Text-to-Speech (TTS) systems, speech recognition, and other audio-driven Machine Learning tasks for the Burmese (Myanmar) language.
Created by DatarrX, this dataset bridges the gap in low-resource speech technologies by providing highly natural, native-sounding… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/burmese-synthetic-speech-corpus.burmese-synthetic-speech-corpus
Burmese Synthetic Speech Corpus (DatarrX/burmese-synthetic-speech-corpus)
Overview
The Burmese Synthetic Speech Corpus is a high-fidelity, manually curated audio dataset specifically designed to advance Text-to-Speech (TTS) systems, speech recognition, and other audio-driven Machine Learning tasks for the Burmese (Myanmar) language.
Created by DatarrX, this dataset bridges the gap in low-resource speech technologies by providing highly natural, native-sounding… See the full description on the dataset page: https://huggingface.co/datasets/hackerlim7/burmese-synthetic-speech-corpus.openslr80-burmese
OpenSLR-80 – Brumese Transcribed Speech
Source: https://www.openslr.org/80/
This dataset contains transcribed high-quality audio of Burmese sentences recorded
by female volunteers. It is part of the OpenSLR collection of
free speech resources for low-resource languages.
The data was collected via the
Appen (formerly Figure Eight / CrowdFlower) crowdsourcing
platform and is intended for use in training automatic speech recognition (ASR)
and text-to-speech (TTS) systems.… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr80-burmese.9000hours_voa_burmese_audio
Overview
VOA Burmese radio news archive covering Morning (နံနက် ၅:၃၀ – ၆:၃၀) and Evening (ညပိုင်း ၉:၀၀ – ၁၀:၀၀) programmes for every calendar day from 2012-09-16 → 2025-06-09.
Metric
Value
Hours / rows
9 159
Files per day
2 (morning, evening)
Typical file size
15 – 50 MB
Licence
Public-domain (VOA staff recordings, U.S. 17 U.S.C. § 105)
This dataset upgrades Burmese from low-resource to mid-resource status for speech research, enabling self-supervised… See the full description on the dataset page: https://huggingface.co/datasets/freococo/9000hours_voa_burmese_audio.myX-Burmese-Numbers
📝 Burmese Numbers
An exhaustive, high-fidelity mapping dataset containing 10,000,001 rows that covers every single integer from 0 to 10,000,000 (One Kote / တစ်ကုဋေ).
This dataset provides a structural translation and text-normalization bridge between Western Arabic numerals and the Burmese numeral system, explicitly broken down into individual digit-by-digit readouts and contextual full-text linguistic expansions. It is designed primarily for Machine Learning engineers… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-Burmese-Numbers.
