CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01speech-uk /voa-2-opus Voice of America 2 for 🇺🇦 Ukrainian (OPUS) Community Discord: https://bit.ly/discord-uds Speech Recognition: https://t.me/speech_recognition_uk Speech Synthesis: https://t.me/speech_synthesis_uk Stats Total files processed: x Total duration: x audioautomatic-speech-recognition100K<n<1M0 likes339 downloads6mo agoHugging Face02speech-uk /voa-opus Voice of America for 🇺🇦 Ukrainian (OPUS) Community Discord: https://bit.ly/discord-uds Speech Recognition: https://t.me/speech_recognition_uk Speech Synthesis: https://t.me/speech_synthesis_uk Stats Total files processed: 326174 Total duration: 390h 59m 54s Other Labels generated by https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 audioautomatic-speech-recognition100K<n<1M0 likes331 downloads6mo agoHugging Face03freococo /voa_myanmar_voices VOA Myanmar Voices Burmese (Myanmar) speech corpus chunked into 20-second FLAC clips with transcripts. Derived from VOA Burmese radio broadcasts (public domain, U.S. 17 U.S.C. § 105). Contents File Size Description voa-00000000.tar … voa-00000238.tar 498 GB 149 WebDataset shards voa_transcripts.parquet 404 MB 1,424,257 (key, text) pairs voa_transcripts.jsonl 1.5 GB Same data, line-oriented Audio: 16 kHz mono FLAC, exactly 20.00 s per chunk… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_voices.automatic-speech-recognition1M<n<10M0 likes233 downloads3d agoHugging Face04freococo /voa_myanmar_asr_audio_1 📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology. Overview This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.audioautomatic-speech-recognition1M<n<10M1 likes221 downloads1y agoHugging Face05UdS-LSV /hausa_voa_topics Dataset Card for Hausa VOA News Topic Classification dataset (hausa_voa_topics) Dataset Summary A news headline topic classification dataset, similar to AG-news, for Hausa. The news headlines were collected from VOA Hausa. Supported Tasks and Leaderboards [More Information Needed] Languages Hausa (ISO 639-1: ha) Dataset Structure Data Instances An instance consists of a news title sentence and the corresponding topic label.… See the full description on the dataset page: https://huggingface.co/datasets/UdS-LSV/hausa_voa_topics.texttext-classification1K<n<10K0 likes177 downloads2y agoHugging Face06UdS-LSV /hausa_voa_nerThe Hausa VOA NER dataset is a labeled dataset for named entity recognition in Hausa. The texts were obtained from Hausa Voice of America News articles https://www.voahausa.com/ . We concentrate on four types of named entities: persons [PER], locations [LOC], organizations [ORG], and dates & time [DATE]. The Hausa VOA NER data files contain 2 columns separated by a tab ('\t'). Each word has been put on a separate line and there is an empty line after each sentences i.e the CoNLL format. The first item on each line is a word, the second is the named entity tag. The named entity tags have the format I-TYPE which means that the word is inside a phrase of type TYPE. For every multi-word expression like 'New York', the first word gets a tag B-TYPE and the subsequent words have tags I-TYPE, a word with tag O is not part of a phrase. The dataset is in the BIO tagging scheme. For more details, see https://www.aclweb.org/anthology/2020.emnlp-main.204/token-classification1K<n<10K3 likes169 downloads3y agoHugging Face07speech-uk /voa Voice of America for 🇺🇦 Ukrainian Community Discord: https://bit.ly/discord-uds Speech Recognition: https://t.me/speech_recognition_uk Speech Synthesis: https://t.me/speech_synthesis_uk Stats Total files processed: 326174 Total duration: 390h 59m 54s Other Labels generated by https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 audioautomatic-speech-recognition100K<n<1M1 likes113 downloads11mo agoHugging Face08robinhad /VOA-ukrgatedaudio10K<n<100K0 likes58 downloads2y agoHugging Face09freococo /9000hours_voa_burmese_audio Overview VOA Burmese radio news archive covering Morning (နံနက် ၅:၃၀ – ၆:၃၀) and Evening (ညပိုင်း ၉:၀၀ – ၁၀:၀၀) programmes for every calendar day from 2012-09-16 → 2025-06-09. Metric Value Hours / rows 9 159 Files per day 2 (morning, evening) Typical file size 15 – 50 MB Licence Public-domain (VOA staff recordings, U.S. 17 U.S.C. § 105) This dataset upgrades Burmese from low-resource to mid-resource status for speech research, enabling self-supervised… See the full description on the dataset page: https://huggingface.co/datasets/freococo/9000hours_voa_burmese_audio.textautomatic-speech-recognition1K<n<10K1 likes50 downloads1y agoHugging Face10freococo /voa_9000h_raw_mp3 VOA 9000h Raw MP3 Raw MP3 archive of VOA Burmese radio broadcasts (2012–2025). 8,234 full-length programs, ~4,000 hours, ~138 GB. Split across 17 tar files (voa_raw_0000.tar … voa_raw_0016.tar), 500 MP3s per tar. Source: freococo/9000hours_voa_burmese_audio — filtered to live URLs. Derived datasets: voa_myanmar_voices — 20s FLAC chunks + transcripts (498 GB) myanmar_asr — ASR model trained on this audio License Public domain (VOA staff recordings, U.S. 17 U.S.C.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_9000h_raw_mp3.audioautomatic-speech-recognition1K<n<10K0 likes39 downloads3d agoHugging Face11voa-engines /common_voice_resampleaudio100K<n<1M1 likes35 downloads2y agoHugging Face12EleutherAI /voanewscaststext10K<n<100K1 likes34 downloads1y agoHugging Face13speech-uk /voa-test-compare-semambaAudios are enhanced by https://github.com/RoyChao19477/SEMamba Generated by https://github.com/RustedBytes/audio-parquet-merger audio1K<n<10K0 likes28 downloads11mo agoHugging Face14voakit05 /donghoimagen<1K0 likes24 downloads10mo agoHugging Face15freococo /voa_myanmar_asr_audio_2⸻ Overview This dataset was created by scraping and segmenting over 4,000 episodes of the VOA Burmese morning radio program. From that archive, 3,687 MP3 files were extracted and processed. This dataset contains sentence-level audio chunks suitable for ASR and speech-related model training. The current release (voa_batch_001.tar and voa_batch_003.tar) contains a combined total of ~152,300 sentence-level audio chunks derived from the first 420 MP3 files in the archive, totaling… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_2.audioautomatic-speech-recognition100K<n<1M0 likes20 downloads1y agoHugging Face16yosiasz /voa_news_amharic Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [Josiah Solomon] Language(s) (NLP): [Amharic] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Uses NLP, POS, NER Direct Use NLP, POS, NER [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/yosiasz/voa_news_amharic.texttext-classificationn<1K0 likes14 downloads10mo agoHugging Face17DatarrX /burmese-VOA Dataset Card for Burmese VOA News Dataset This dataset is a comprehensive collection of Burmese news articles crawled from Voice of America (VOA) Burmese. It is specifically curated and processed for Natural Language Processing (NLP) tasks, focusing on high-quality news content, including the "Science and Technology" category. Dataset Summary The Burmese VOA Dataset contains 270,546 rows of news articles. The data has been meticulously scraped and structured into a… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/burmese-VOA.texttext-generation100K<n<1M5 likes14 downloads5mo agoHugging Face18Aletheia-ng /hausa_voa_topicstext1K<n<10K0 likes12 downloads2y agoHugging Face19speech-uk /voa-test-compare-sidonAudios are enhanced by https://github.com/sarulab-speech/Sidon Generated by https://github.com/RustedBytes/audio-parquet-merger audio1K<n<10K0 likes12 downloads11mo agoHugging Face20cillegio /az-asr-voa-305hgated Labelling field value label_origin script speech_register broadcast channel wideband-16k provenance inferred Editorial broadcast text aligned to VOA audio. Terminal punctuation at 77.4% (against 98.1% for LocalDoc) is consistent with segments cut from continuous broadcast rather than at sentence ends. Adding this to a call model made it worse. Chinar-F8 v4 mixed in 147.8 h of it and strict WER on human-transcribed calls moved 44.23% -> 56.69%. Register, not… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-voa-305h.tabularautomatic-speech-recognition100K<n<1M0 likes10 downloads5d agoHugging Face21voa-engines /sft-dataset Dataset Card for sft-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/voa-engines/sft-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/voa-engines/sft-dataset.textn<1K0 likes9 downloads2y agoHugging Face22alamin05 /voa_dataset Dataset Card for "voa_dataset" More Information needed textn<1K0 likes8 downloads1y agoHugging Face23Bruce-Azar-Wayne /VOA_CSVFormattextn<1K0 likes7 downloads2y agoHugging Face24CKQUED01 /VOAJB-150 likes7 downloads1y agoHugging Face25Bruce-Azar-Wayne /VOA_newstextn<1K0 likes4 downloads2y agoHugging Face26voa-engines /voa-citationstextn<1K0 likes4 downloads1y agoHugging Face27lucien /voacantonesed0 likes1 downloads5y agoHugging Face28ciafa /VOAMAISgatedData collected for the VOAMAIS project, within the REX'22 exercise (link 1, link 2), by the Air Force Academy Research Center. For a calibration sample, see this dataset. 0 likes1 downloads7mo agoHugging Face29soulseekerok /voakem0 likes1 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.