datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-Voice-Dataset-Train
X-Voice Training Dataset
Overview
The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling.
Also the train set of X-Voice Model.
Core Statistics
Total Speech Duration: 420K hours
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.voice-data
Voice-Data: a curated multi-corpus voice dataset for voice–text contrastive training
voice-data is a single, globally-shuffled WebDataset that bundles several voice/speech corpora into one ready-to-train mixture for voice–text contrastive (CLAP-style) models such as VoiceCLAP. Each clip pairs 48 kHz mono FLAC audio with a natural-language text caption describing the voice — its emotion, prosody, timbre, speaking style, recording context, and speaker traits.
The distinguishing… See the full description on the dataset page: https://huggingface.co/datasets/gijs/voice-data.unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1Japanese-Eroge-Voice
Japanese-Eroge-Voice
Description
This dataset contains pairs of audio data and corresponding transcriptions extracted from Japanese eroge (adult games) that I have personally purchased. The transcriptions are generated using the litagin/anime-whisper model.
Preprocessing Steps
The raw audio data has undergone the following preprocessing steps:
Loudness Normalization:
Audio loudness is normalized using ffmpeg's 2-pass loudnorm filter to target parameters of… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice.voiceclap-data
VoiceCLAP Data
The audio + dense-caption mixture used to train
laion/voiceclap-small and
laion/voiceclap-large.
Each tar shard is a WebDataset of
paired <key>.flac (48 kHz mono audio) + <key>.json (caption + metadata)
samples. Captions and structured attribute annotations are produced
automatically by a pipeline of audio-aware LLMs — Qwen-Audio, Gemini Flash 2.5,
and a thinking-mode reasoning model that scores emotion under the EmoNet
taxonomy plus per-clip vocal-burst, timbre… See the full description on the dataset page: https://huggingface.co/datasets/laion/voiceclap-data.common-voice-subset-for-clapemolia
emolia-balanced-5M-subset · flac 48 kHz · WebDataset (paired)
This is the emolia-balanced-5M-subset corpus repackaged for high-quality
audio–text contrastive training. Audio is re-encoded as mono FLAC at 48 kHz
(PCM 16-bit) and stored as a WebDataset of paired <key>.flac + <key>.json
samples.
The JSON sidecar carries the full annotation stack:
Original metadata (id, text, duration, speaker, language, dnsmos).
A free-text emotion_caption derived from the emotion-annotation scalars.… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia.tachiwin_voice_raw
Tachiwin Indigenous Languages of Mexico Voice Pretrain Dataset (RAW)
This dataset collects a large amount of speech hours in indigenous languages of Mexico recorded from 23 public radio stations that broadcast to the indigenous communities.
How it was obtained
After the raw recording of all transmissions, the music and inaudible sound were removed and then the samples of pure speech were classified by language. However, as there are no existing speech language classifiers… See the full description on the dataset page: https://huggingface.co/datasets/tachiwin/tachiwin_voice_raw.common_voice_21_ru
Dataset Description
Набор данных validated.tsv отфильтрованный по down_votes = 0
📊 Статистика датасета
Информация по сплитам
🔹 Тренировочный набор (train)
Метрика
Значение
Количество семплов
93,531
Общая продолжительность
132.25 часов (476,089.70 секунд)
Средняя продолжительность семпла
5.09 секунд
🔹 Валидационный набор (validate)
Метрика
Значение
Количество семплов
38,836
Общая продолжительность
55.21… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/common_voice_21_ru.common_voice_englishcommon_voice_21_0_yuecantonese only
shan_language_asr_voices
⭐ A Voice for the Shan People: The SHAN Herald Agency Audio Archive
This is an extensive 306-hour audio dataset of the Shan (Tai-Yai) language, meticulously curated from the public broadcasts of the Shan Herald Agency for News (SHAN). For over two decades, SHAN has been a vital, independent voice for the people of Shan State, Myanmar, chronicling their stories of culture, politics, and the enduring struggle for federal democracy.
This archive stands as one of the largest publicly… See the full description on the dataset page: https://huggingface.co/datasets/freococo/shan_language_asr_voices.majestrinoEmotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: 1.54 TB
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.voice-annotation-data-v2
Voice Annotation Data v2
A curated dataset of 18,632 audio samples (9,391 positives + 9,241 negatives) across 58 voice dimensions. Each bucket contains up to 25 positive examples (audio that clearly fits the bucket) and 25 negative examples (audio confirmed to NOT fit the bucket by Gemini 2.0 Flash).
Changes from v1
Positive + Negative pairs: Every bucket now has up to 25 confirmed negative examples alongside 25 positives
EXPL redefined: Content Appropriateness reduced… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/voice-annotation-data-v2.X-Voice-TestsetX-Voice Multilingual Test Set
High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages
This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model.
Dataset Summary
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.media_queen_entertaiment_voices
Media Queen Entertainment Voices
Where the stars speak, and their stories come to life.
Media Queen Entertainment Voices is a massive, large-scale collection of 190,013 short audio segments (totaling approximately 125 hours of speech) derived from public videos by Media Queen Entertainment — a prominent digital media channel in Myanmar focused on celebrity news, lifestyle content, and in-depth interviews.
The source channel regularly features:
Interviews with artists, actors… See the full description on the dataset page: https://huggingface.co/datasets/freococo/media_queen_entertaiment_voices.common_voice_16_1_fr_smallkhit_thit_news_voices
Khit Thit News Voices
In the fight for truth, these are the voices that refuse to be silenced.
Khit Thit News Voices is a focused collection of 15,841 audio segments (≈14.7 hours total) from Khit Thit News, one of Myanmar's most vital and trusted independent media outlets. Founded by renowned journalist Mr. Thar Lun Zaung Htet, Khit Thit News stands as a pillar of reliable information and a primary voice for democratic forces within the country.
This dataset primarily features the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/khit_thit_news_voices.myanmar_cele_voices
Myanmar Celebrity Voices
A high-quality speech dataset extracted from the official TikTok channel of Myanmar Celebrity TV.
Myanmar Celebrity Voices is a collection of 69,781 short audio segments (≈46 hours total) derived from public TikTok videos by The Official TikTok Channel of Myanmar Celebrity TV — one of the most popular digital media platforms in Myanmar.
The source channel regularly publishes:
Interviews with Myanmar’s top movie actors and actresses
Behind-the-scenes… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_cele_voices.clustered-reference-voices
Clustered Reference Voices (EMOLIA 3K)
3,000 enhanced reference voice MP3s — one high-quality representative sample per speaker cluster, selected and scored by a multi-expert neural quality model.
Overview
Property
Value
Total clips
3,000
Total duration
11.3 hours
Mean duration
13.5 s (range: 3.5 – 29.9 s)
Format
MP3, 192 kbps, 48 kHz
Language
English
Naming
{cluster_id}.mp3 (0 – 2999)
Source
The source data is… See the full description on the dataset page: https://huggingface.co/datasets/laion/clustered-reference-voices.mrtv_news_voices
🗣️ Overview
MRTV Voices is a large-scale Burmese speech dataset built from publicly available news broadcasts and programs aired on Myanma Radio and Television (MRTV) — the official state-run media channel of Myanmar.
🎙️ It contains over 130,000 short audio clips (≈117 hours) with aligned transcripts derived from auto-generated subtitles.
This dataset captures:
Formal Burmese used in government bulletins and official reports
Clear pronunciation, enunciation, and pacing —… See the full description on the dataset page: https://huggingface.co/datasets/freococo/mrtv_news_voices.reference-voices-enhanced
Reference Voices Enhanced
2,004 AI voice samples enhanced with ClearerVoice-Studio MossFormer2_SE_48K speech enhancement, annotated with Empathic Insight Voice Plus (59 quality + emotion scores).
Dataset Summary
Source: laion/ai-voices-deduplicated (2,004 speaker-deduplicated, quality-filtered AI voice samples)
Speech Enhancement: ClearerVoice MossFormer2_SE_48K — background noise removal and speech clarity improvement
Output Format: Enhanced WAV files at 48kHz… See the full description on the dataset page: https://huggingface.co/datasets/laion/reference-voices-enhanced.Japanese-Eroge-Voice
Japanese-Eroge-Voice
Description
This dataset contains pairs of audio data and corresponding transcriptions extracted from Japanese eroge (adult games) that I have personally purchased. The transcriptions are generated using the litagin/anime-whisper model.
Preprocessing Steps
The raw audio data has undergone the following preprocessing steps:
Loudness Normalization:
Audio loudness is normalized using ffmpeg's 2-pass loudnorm filter to target parameters of… See the full description on the dataset page: https://huggingface.co/datasets/test12313/Japanese-Eroge-Voice.sunday_journal_voices
Sunday Journal Voices
Voices that shape our week, stories that define our time.
Sunday Journal Voices is a large-scale collection of 26,288 short audio segments (≈18 hours total) derived from public videos by Sunday Journal — a leading digital media platform in Myanmar known for its in-depth reporting and interviews.
The source channel regularly publishes:
News analysis and commentary on current events
In-depth interviews with public figures, experts, and community leaders… See the full description on the dataset page: https://huggingface.co/datasets/freococo/sunday_journal_voices.mozilla-common-voice-spontaneous-speech-asr-shared-task
Mozilla Common Voice Spontaneous Speech ASR Shared Task
This repository combines the Mozilla Data Collective Common Voice spontaneous speech ASR shared-task
train/dev and test archives in one place.
Locales present across the combined train/dev and test packages: ady, aln, bas, bew, bxk,
cgg, el-CY, hch, kbd, kcn, koo, led, lke, lth, meh, mmc, pne, qxp, ruc,
rwm, sco, tob, top, ttj, ukv, ush.
Split package
Mozilla Data Collective dataset ID
Hub archive
Original MDC archive… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/mozilla-common-voice-spontaneous-speech-asr-shared-task.rfa_rakhine_language_voices
RFA Rakhine Language Voices
This dataset contains 14.53 hours of audio in the Rakhine (Arakanese) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Rakhine language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rfa_rakhine_language_voices.mrtv4_voices
MRTV4 Official Voices
The official voice of Myanmar's national broadcaster, ready for the age of AI.
MRTV4 Official Voices is a focused collection of 2,368 high-quality audio segments (totaling 1h 51m 54s of speech) derived from the public broadcasts of MRTV4 Official — one of Myanmar's leading television channels.
The source channel regularly features:
National news broadcasts and public announcements.
Educational programming and cultural segments.
Formal presentations and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/mrtv4_voices.japanese-singing-voice
Japanese Singing Voice Dataset / 日本語歌声データセット
English | 日本語
English
A large-scale Japanese singing voice dataset for training voice conversion models.
Dataset Description
This dataset contains Japanese singing voice audio files collected for training singing voice conversion (SVC) models such as Seed-VC, RVC, So-VITS-SVC, and similar architectures.
Dataset Statistics
Metric
Value
Total Duration
~1,000 hours
Number of Files
15,311… See the full description on the dataset page: https://huggingface.co/datasets/tts-dataset/japanese-singing-voice.common_voice_16_1
