datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mls_hq_urgent_track1emo_webds_2emo_parleremo_webdsOpenDialog
OpenDialog
OpenDialog is a 6.8k hours spoken dialogue dataset, introduced in the paper ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching.
Paper: https://arxiv.org/abs/2507.09318
GitHub: https://github.com/k2-fsa/ZipVoice
Project Page: https://zipvoice-dialog.github.io
OpenDialog is the first large-scale (6.8k hours) open-source spoken dialogue dataset derived from in-the-wild speech data. It consists of:
English data: 5074 hours
Chinese data: 1759… See the full description on the dataset page: https://huggingface.co/datasets/k2-fsa/OpenDialog.laion-coco-13m-tarmls_hqemo_speech_filtered_v12 second filtered emotional speech in webdataset format
https://huggingface.co/datasets/EQ4You/Emotional_Speech
vocal_bursts_taxonomy_100_clean_wdslaion-audio-preview-splitEmilia-YODAS-KO-filteredkhit_thit_news_voices
Khit Thit News Voices
In the fight for truth, these are the voices that refuse to be silenced.
Khit Thit News Voices is a focused collection of 15,841 audio segments (≈14.7 hours total) from Khit Thit News, one of Myanmar's most vital and trusted independent media outlets. Founded by renowned journalist Mr. Thar Lun Zaung Htet, Khit Thit News stands as a pillar of reliable information and a primary voice for democratic forces within the country.
This dataset primarily features the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/khit_thit_news_voices.kenya-philippines-twospeaker-english-dialogue
Kenya/Philippines English Dialogue
Two-speaker dialogues in English, recorded on split tracks.
Changelog
Jan 2026: v1 release - vad-segmented WebRTC tracks
Specs
Speakers: >150; ~15 PH, remaining KE
Total duration: ~65 hours
Files sample rate: 48kHz
Actual sample rate: TBD
Language: English (PH, KE accents)
Topics: day-to-day conversation
Collection method
The dataset is built to capture the variety in the Kenyan accent.
The Philippino interviewers… See the full description on the dataset page: https://huggingface.co/datasets/Reord-AI/kenya-philippines-twospeaker-english-dialogue.earssagaw_karen_asrThis is the first public Sagaw Karen language ASR dataset in AI history.
Sagaw Karen ASR
This dataset contains audio recordings and aligned metadata in the Sagaw Karen language (ISO 639-3: ksw), a major Sgaw Karenic language spoken throughout southern and eastern Myanmar. The language is sometimes also referred to as Sgaw Karen or Sakaw Karen in English transliterations.
All audio segments in this dataset were sourced from publicly available news broadcasts published by PVTV… See the full description on the dataset page: https://huggingface.co/datasets/freococo/sagaw_karen_asr.karenni_language_asr_audio
RFA Karenni (Kayah) Language Voices
This dataset contains 17 hours of audio in the Karenni (Kayah) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Karenni language family, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
This dataset was created by freococo.
The audio has been automatically segmented… See the full description on the dataset page: https://huggingface.co/datasets/freococo/karenni_language_asr_audio.kachin_asr_audio
Dataset Summary
This is the first public Kachin language ASR dataset in history.
Kachin ASR Audio is a collection of speech data in the Kachin (Jinghpaw) language, sourced entirely from publicly available PVTV (People’s Voice Television) broadcasts. The dataset includes narration, interviews, and spoken reports intended to support the development of automatic speech recognition (ASR) systems for rare-resource indigenous languages in Myanmar.
Each audio file is paired with metadata… See the full description on the dataset page: https://huggingface.co/datasets/freococo/kachin_asr_audio.Speech-IFEvalwestern_poe_karen_asrThis is the first public Western Poe Karen language ASR dataset in AI history.
Western Poe Karen ASR
This dataset contains audio recordings and aligned transcriptions in the Western Poe Karen language (also known in linguistic literature as Western Pwo or Delta Pwo, ISO 639-3: pwo), a Karenic language spoken primarily in the Ayeyarwady Delta region of Myanmar. Although linguists commonly refer to this language as Western Pwo Karen, the community and this project prefer the spelling… See the full description on the dataset page: https://huggingface.co/datasets/freococo/western_poe_karen_asr.kimi-audio-dpoeastern_poe_karen_asrThis is the first public Eastern Poe Karen language ASR dataset in AI history.
Eastern Poe Karen ASR
This dataset contains audio recordings and aligned metadata in the Eastern Poe Karen language (a regional variety of Eastern Pwo, ISO 639-3: pwo), a Karenic language spoken primarily in Mon State and Kayin State in southeastern Myanmar. While linguistically described as Eastern Pwo Karen, the community and this project prefer the term Poe as a community-endorsed spelling.
All audio… See the full description on the dataset page: https://huggingface.co/datasets/freococo/eastern_poe_karen_asr.COIG-Kun-Aug-Audio本数据集基于https://huggingface.co/datasets/m-a-p/COIG-Kun作为种子问题,使用Qwen2.5-72B-Instruct-GPTQ-Int4继续生成更多轮次的问题。然后使用Qwen2.5-72B-Instruct-GPTQ-Int4生成问题的答案(每轮答案生成都会将之前的问题和答案当作上下文,确保当前的答案和历史相关)。见sharegpt.json文件。问题使用cosyvoice生成对应音频。audio_part1-3.tar.gz分别是三部分压缩包,分别解压后合并成一个audio文件夹,也可只下载一部分使用。
wds_vocal_burst_100Emilia-YODAS-KOdynamic-superb-train-noise-reverbyoutube_ka_rajesh_raw_tempyoutube_kn_rajesh_raw_tempEmilia-KOyoutube_ka_bookbrahma_raw_tempBEAF-Audio
