datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voa-2-opus
Voice of America 2 for 🇺🇦 Ukrainian (OPUS)
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: x
Total duration: x
voa-opus
Voice of America for 🇺🇦 Ukrainian (OPUS)
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 326174
Total duration: 390h 59m 54s
Other
Labels generated by https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
voa_myanmar_voices
VOA Myanmar Voices
Burmese (Myanmar) speech corpus chunked into 20-second FLAC clips with transcripts. Derived from VOA Burmese radio broadcasts (public domain, U.S. 17 U.S.C. § 105).
Contents
File
Size
Description
voa-00000000.tar … voa-00000238.tar
498 GB
149 WebDataset shards
voa_transcripts.parquet
404 MB
1,424,257 (key, text) pairs
voa_transcripts.jsonl
1.5 GB
Same data, line-oriented
Audio: 16 kHz mono FLAC, exactly 20.00 s per chunk… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_voices.voa_myanmar_asr_audio_1
📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology.
Overview
This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.hausa_voa_topics
Dataset Card for Hausa VOA News Topic Classification dataset (hausa_voa_topics)
Dataset Summary
A news headline topic classification dataset, similar to AG-news, for Hausa. The news headlines were collected from VOA Hausa.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Hausa (ISO 639-1: ha)
Dataset Structure
Data Instances
An instance consists of a news title sentence and the corresponding topic label.… See the full description on the dataset page: https://huggingface.co/datasets/UdS-LSV/hausa_voa_topics.hausa_voa_nerThe Hausa VOA NER dataset is a labeled dataset for named entity recognition in Hausa. The texts were obtained from
Hausa Voice of America News articles https://www.voahausa.com/ . We concentrate on
four types of named entities: persons [PER], locations [LOC], organizations [ORG], and dates & time [DATE].
The Hausa VOA NER data files contain 2 columns separated by a tab ('\t'). Each word has been put on a separate line and
there is an empty line after each sentences i.e the CoNLL format. The first item on each line is a word, the second
is the named entity tag. The named entity tags have the format I-TYPE which means that the word is inside a phrase
of type TYPE. For every multi-word expression like 'New York', the first word gets a tag B-TYPE and the subsequent words
have tags I-TYPE, a word with tag O is not part of a phrase. The dataset is in the BIO tagging scheme.
For more details, see https://www.aclweb.org/anthology/2020.emnlp-main.204/voa
Voice of America for 🇺🇦 Ukrainian
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 326174
Total duration: 390h 59m 54s
Other
Labels generated by https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
VOA-ukr9000hours_voa_burmese_audio
Overview
VOA Burmese radio news archive covering Morning (နံနက် ၅:၃၀ – ၆:၃၀) and Evening (ညပိုင်း ၉:၀၀ – ၁၀:၀၀) programmes for every calendar day from 2012-09-16 → 2025-06-09.
Metric
Value
Hours / rows
9 159
Files per day
2 (morning, evening)
Typical file size
15 – 50 MB
Licence
Public-domain (VOA staff recordings, U.S. 17 U.S.C. § 105)
This dataset upgrades Burmese from low-resource to mid-resource status for speech research, enabling self-supervised… See the full description on the dataset page: https://huggingface.co/datasets/freococo/9000hours_voa_burmese_audio.voa_9000h_raw_mp3
VOA 9000h Raw MP3
Raw MP3 archive of VOA Burmese radio broadcasts (2012–2025). 8,234 full-length programs, ~4,000 hours, ~138 GB.
Split across 17 tar files (voa_raw_0000.tar … voa_raw_0016.tar), 500 MP3s per tar.
Source: freococo/9000hours_voa_burmese_audio — filtered to live URLs.
Derived datasets:
voa_myanmar_voices — 20s FLAC chunks + transcripts (498 GB)
myanmar_asr — ASR model trained on this audio
License
Public domain (VOA staff recordings, U.S. 17 U.S.C.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_9000h_raw_mp3.common_voice_resamplevoanewscastsvoa-test-compare-semambaAudios are enhanced by https://github.com/RoyChao19477/SEMamba
Generated by https://github.com/RustedBytes/audio-parquet-merger
donghovoa_myanmar_asr_audio_2⸻
Overview
This dataset was created by scraping and segmenting over 4,000 episodes of the VOA Burmese morning radio program. From that archive, 3,687 MP3 files were extracted and processed. This dataset contains sentence-level audio chunks suitable for ASR and speech-related model training.
The current release (voa_batch_001.tar and voa_batch_003.tar) contains a combined total of ~152,300 sentence-level audio chunks derived from the first 420 MP3 files in the archive, totaling… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_2.voa_news_amharic
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [Josiah Solomon]
Language(s) (NLP): [Amharic]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Uses
NLP, POS, NER
Direct Use
NLP, POS, NER
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/yosiasz/voa_news_amharic.burmese-VOA
Dataset Card for Burmese VOA News Dataset
This dataset is a comprehensive collection of Burmese news articles crawled from Voice of America (VOA) Burmese. It is specifically curated and processed for Natural Language Processing (NLP) tasks, focusing on high-quality news content, including the "Science and Technology" category.
Dataset Summary
The Burmese VOA Dataset contains 270,546 rows of news articles. The data has been meticulously scraped and structured into a… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/burmese-VOA.hausa_voa_topicsvoa-test-compare-sidonAudios are enhanced by https://github.com/sarulab-speech/Sidon
Generated by https://github.com/RustedBytes/audio-parquet-merger
az-asr-voa-305h
Labelling
field
value
label_origin
script
speech_register
broadcast
channel
wideband-16k
provenance
inferred
Editorial broadcast text aligned to VOA audio. Terminal punctuation at 77.4% (against 98.1% for LocalDoc) is consistent with segments cut from continuous broadcast rather than at sentence ends.
Adding this to a call model made it worse. Chinar-F8 v4 mixed in 147.8 h of it and strict WER on human-transcribed calls moved 44.23% -> 56.69%. Register, not… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-voa-305h.sft-dataset
Dataset Card for sft-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/voa-engines/sft-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/voa-engines/sft-dataset.voa_dataset
Dataset Card for "voa_dataset"
More Information needed
VOA_CSVFormatVOAJB-15VOA_newsvoa-citationsvoacantonesedVOAMAISData collected for the VOAMAIS project,
within the REX'22 exercise (link 1, link 2),
by the Air Force Academy Research Center.
For a calibration sample, see this dataset.
voakem
