datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-Voice-Dataset-Train
X-Voice Training Dataset
Overview
The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling.
Also the train set of X-Voice Model.
Core Statistics
Total Speech Duration: 420K hours
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.Japanese-Eroge-Voice
Japanese-Eroge-Voice
Description
This dataset contains pairs of audio data and corresponding transcriptions extracted from Japanese eroge (adult games) that I have personally purchased. The transcriptions are generated using the litagin/anime-whisper model.
Preprocessing Steps
The raw audio data has undergone the following preprocessing steps:
Loudness Normalization:
Audio loudness is normalized using ffmpeg's 2-pass loudnorm filter to target parameters of… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice.emolia
emolia-balanced-5M-subset · flac 48 kHz · WebDataset (paired)
This is the emolia-balanced-5M-subset corpus repackaged for high-quality
audio–text contrastive training. Audio is re-encoded as mono FLAC at 48 kHz
(PCM 16-bit) and stored as a WebDataset of paired <key>.flac + <key>.json
samples.
The JSON sidecar carries the full annotation stack:
Original metadata (id, text, duration, speaker, language, dnsmos).
A free-text emotion_caption derived from the emotion-annotation scalars.… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia.shan_language_asr_voices
⭐ A Voice for the Shan People: The SHAN Herald Agency Audio Archive
This is an extensive 306-hour audio dataset of the Shan (Tai-Yai) language, meticulously curated from the public broadcasts of the Shan Herald Agency for News (SHAN). For over two decades, SHAN has been a vital, independent voice for the people of Shan State, Myanmar, chronicling their stories of culture, politics, and the enduring struggle for federal democracy.
This archive stands as one of the largest publicly… See the full description on the dataset page: https://huggingface.co/datasets/freococo/shan_language_asr_voices.X-Voice-TestsetX-Voice Multilingual Test Set
High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages
This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model.
Dataset Summary
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.media_queen_entertaiment_voices
Media Queen Entertainment Voices
Where the stars speak, and their stories come to life.
Media Queen Entertainment Voices is a massive, large-scale collection of 190,013 short audio segments (totaling approximately 125 hours of speech) derived from public videos by Media Queen Entertainment — a prominent digital media channel in Myanmar focused on celebrity news, lifestyle content, and in-depth interviews.
The source channel regularly features:
Interviews with artists, actors… See the full description on the dataset page: https://huggingface.co/datasets/freococo/media_queen_entertaiment_voices.khit_thit_news_voices
Khit Thit News Voices
In the fight for truth, these are the voices that refuse to be silenced.
Khit Thit News Voices is a focused collection of 15,841 audio segments (≈14.7 hours total) from Khit Thit News, one of Myanmar's most vital and trusted independent media outlets. Founded by renowned journalist Mr. Thar Lun Zaung Htet, Khit Thit News stands as a pillar of reliable information and a primary voice for democratic forces within the country.
This dataset primarily features the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/khit_thit_news_voices.myanmar_cele_voices
Myanmar Celebrity Voices
A high-quality speech dataset extracted from the official TikTok channel of Myanmar Celebrity TV.
Myanmar Celebrity Voices is a collection of 69,781 short audio segments (≈46 hours total) derived from public TikTok videos by The Official TikTok Channel of Myanmar Celebrity TV — one of the most popular digital media platforms in Myanmar.
The source channel regularly publishes:
Interviews with Myanmar’s top movie actors and actresses
Behind-the-scenes… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_cele_voices.mrtv_news_voices
🗣️ Overview
MRTV Voices is a large-scale Burmese speech dataset built from publicly available news broadcasts and programs aired on Myanma Radio and Television (MRTV) — the official state-run media channel of Myanmar.
🎙️ It contains over 130,000 short audio clips (≈117 hours) with aligned transcripts derived from auto-generated subtitles.
This dataset captures:
Formal Burmese used in government bulletins and official reports
Clear pronunciation, enunciation, and pacing —… See the full description on the dataset page: https://huggingface.co/datasets/freococo/mrtv_news_voices.Japanese-Eroge-Voice
Japanese-Eroge-Voice
Description
This dataset contains pairs of audio data and corresponding transcriptions extracted from Japanese eroge (adult games) that I have personally purchased. The transcriptions are generated using the litagin/anime-whisper model.
Preprocessing Steps
The raw audio data has undergone the following preprocessing steps:
Loudness Normalization:
Audio loudness is normalized using ffmpeg's 2-pass loudnorm filter to target parameters of… See the full description on the dataset page: https://huggingface.co/datasets/test12313/Japanese-Eroge-Voice.sunday_journal_voices
Sunday Journal Voices
Voices that shape our week, stories that define our time.
Sunday Journal Voices is a large-scale collection of 26,288 short audio segments (≈18 hours total) derived from public videos by Sunday Journal — a leading digital media platform in Myanmar known for its in-depth reporting and interviews.
The source channel regularly publishes:
News analysis and commentary on current events
In-depth interviews with public figures, experts, and community leaders… See the full description on the dataset page: https://huggingface.co/datasets/freococo/sunday_journal_voices.mozilla-common-voice-spontaneous-speech-asr-shared-task
Mozilla Common Voice Spontaneous Speech ASR Shared Task
This repository combines the Mozilla Data Collective Common Voice spontaneous speech ASR shared-task
train/dev and test archives in one place.
Locales present across the combined train/dev and test packages: ady, aln, bas, bew, bxk,
cgg, el-CY, hch, kbd, kcn, koo, led, lke, lth, meh, mmc, pne, qxp, ruc,
rwm, sco, tob, top, ttj, ukv, ush.
Split package
Mozilla Data Collective dataset ID
Hub archive
Original MDC archive… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/mozilla-common-voice-spontaneous-speech-asr-shared-task.rfa_rakhine_language_voices
RFA Rakhine Language Voices
This dataset contains 14.53 hours of audio in the Rakhine (Arakanese) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Rakhine language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rfa_rakhine_language_voices.mrtv4_voices
MRTV4 Official Voices
The official voice of Myanmar's national broadcaster, ready for the age of AI.
MRTV4 Official Voices is a focused collection of 2,368 high-quality audio segments (totaling 1h 51m 54s of speech) derived from the public broadcasts of MRTV4 Official — one of Myanmar's leading television channels.
The source channel regularly features:
National news broadcasts and public announcements.
Educational programming and cultural segments.
Formal presentations and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/mrtv4_voices.rfa_shan_language_voices
RFA Shan Language Voices
This dataset contains 20.58 hours of audio in the Shan (Tai-Yai) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Shan language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into 5,047 manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rfa_shan_language_voices.google_myanmar_asr_voices
Google Myanmar ASR Dataset (WebDataset Version)
This repository provides a clean, user-friendly, and robust version of the Google Myanmar ASR Dataset, which is derived from the OpenSLR-80 Burmese Speech Corpus.
This version has been carefully re-processed into the WebDataset format. Each sample consists of a .wav audio file and a clean .json metadata file, packaged into sharded .tar archives. This format is highly efficient for large-scale training of ASR models.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/google_myanmar_asr_voices.zello-public-channels-voice-sample
Zello Public Channels Voice Dataset Sample
Dataset summary
Total audio: 48.23 hours
Total messages: 13207
Breakdown by language
Language
Hours
Messages
Speakers
Channels
ms
10.65
2056
112
7
en
9.54
2990
362
26
id
6.98
1597
157
11
es
6.86
2144
224
21
tl
4.28
1404
78
6
pt
3.92
863
66
11
sw
2.43
182
21
1
ru
1.30
282
42
12
th
0.95
55090
3
zh
0.87
828
312
10
it
0.15
25
19
7
is
0.09
76
25
10
ko
0.05
75
43
13
vi
0.04
42
34
8
fr… See the full description on the dataset page: https://huggingface.co/datasets/zello/zello-public-channels-voice-sample.
