datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voa-2-opus
Voice of America 2 for 🇺🇦 Ukrainian (OPUS)
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: x
Total duration: x
voa-opus
Voice of America for 🇺🇦 Ukrainian (OPUS)
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 326174
Total duration: 390h 59m 54s
Other
Labels generated by https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
voa_myanmar_asr_audio_1
📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology.
Overview
This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.voa
Voice of America for 🇺🇦 Ukrainian
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 326174
Total duration: 390h 59m 54s
Other
Labels generated by https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
9000hours_voa_burmese_audio
Overview
VOA Burmese radio news archive covering Morning (နံနက် ၅:၃၀ – ၆:၃၀) and Evening (ညပိုင်း ၉:၀၀ – ၁၀:၀၀) programmes for every calendar day from 2012-09-16 → 2025-06-09.
Metric
Value
Hours / rows
9 159
Files per day
2 (morning, evening)
Typical file size
15 – 50 MB
Licence
Public-domain (VOA staff recordings, U.S. 17 U.S.C. § 105)
This dataset upgrades Burmese from low-resource to mid-resource status for speech research, enabling self-supervised… See the full description on the dataset page: https://huggingface.co/datasets/freococo/9000hours_voa_burmese_audio.voa_myanmar_asr_audio_2⸻
Overview
This dataset was created by scraping and segmenting over 4,000 episodes of the VOA Burmese morning radio program. From that archive, 3,687 MP3 files were extracted and processed. This dataset contains sentence-level audio chunks suitable for ASR and speech-related model training.
The current release (voa_batch_001.tar and voa_batch_003.tar) contains a combined total of ~152,300 sentence-level audio chunks derived from the first 420 MP3 files in the archive, totaling… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_2.az-asr-voa-305h
Labelling
field
value
label_origin
script
speech_register
broadcast
channel
wideband-16k
provenance
inferred
Editorial broadcast text aligned to VOA audio. Terminal punctuation at 77.4% (against 98.1% for LocalDoc) is consistent with segments cut from continuous broadcast rather than at sentence ends.
Adding this to a call model made it worse. Chinar-F8 v4 mixed in 147.8 h of it and strict WER on human-transcribed calls moved 44.23% -> 56.69%. Register, not… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-voa-305h.
