datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
acl-voice-cloning-fr-expandedtry
ACL Voice Cloning FR — Expanded Pairs
Every (reference, target) segment pair within each speaker.
Audio is embedded directly — playable on the dataset viewer.
Column
Type
Role
ref_en_voice
🎤 Audio
English audio of reference segment (voice to clone)
ref_fr_voice
🎤 Audio
French cloned audio of reference segment
ref_en_text
string
English text of reference
ref_fr_text
string
French text of reference
trg_en_voice
🎤 Audio
English audio of target segment… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/acl-voice-cloning-fr-expandedtry.acl-voice-cloning-fr-expanded
ACL Voice Cloning FR — Expanded Pairs
Every (reference, target) segment pair within each speaker.
Audio is embedded directly — playable on the dataset viewer.
Column
Type
Role
ref_en_voice
🎤 Audio
English audio of reference segment (voice to clone)
ref_fr_voice
🎤 Audio
French cloned audio of reference segment
ref_en_text
string
English text of reference
ref_fr_text
string
French text of reference
trg_en_voice
🎤 Audio
English audio of target segment… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/acl-voice-cloning-fr-expanded.voice_cloning_2Qwen3-TTS-Cloning-VoicesEarningsCallVoice-Cloning-Listening
Voice Cloning Listening Study
Open the listening app or
download the offline package.
This package contains 100 earnings-call listening sets. Each set provides a
reference voice, an analyst question and nine anonymous candidate answers:
one recorded answer and eight voice clones. All 1,000 WAV files are mono,
16 kHz, PCM-16. No answer key, previous rankings or participant responses are
included. Please do not look up candidate identities while labeling.
Participate… See the full description on the dataset page: https://huggingface.co/datasets/gmarti/EarningsCallVoice-Cloning-Listening.voice_cloning_style_transfer
Voice "Cloning" is Style Transfer — Audio Dataset
Companion dataset for the preprint
"Voice 'Cloning' is Style Transfer" (Zhou, Bianchi, Bartelds, Pot, Kwon, Zou; 2026).
Code, notebooks, and reproduction figures live at
github.com/kzhou-cloud/voice-cloning-public.
🎧 Listen to a small set of paired examples on the
project page.
What's in here
Split
# files
Description
original
699
QC-validated human recordings of the Grandfather Passage from 86 non-native… See the full description on the dataset page: https://huggingface.co/datasets/kzhou/voice_cloning_style_transfer.voice-emo-cloning-dataset
Emotion-Cloning TTS Training Dataset
Location
/home/deployer/laion/echo-tts-training-main/emotion_eval/dataset_output/
Overview
This dataset contains ~22,518 training triplets for fine-tuning a zero-shot voice+emotion cloning TTS model. Each sample provides everything needed to train a model that can clone both a speaker's voice identity AND their emotional delivery from separate reference audio clips.
The data is stored as WebDataset .tar shards, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/voice-emo-cloning-dataset.acl-voice-cloning-fr-cleaned-v2
ACL Voice Cloning FR — Cleaned & Expanded (V2)
Source filtered with GPU-accelerated quality checks (SNR≥10.0dB, silence≤65%),
then expanded into all (reference, target) pairs per speaker.
Filter
Threshold
Duration
1.0-20.0s
SNR
≥ 10.0 dB
Silence
≤ 65%
Text length
5-500 chars
Split
Source kept
Expanded pairs
Shards
train
748
70,006
43
test
119
7,022
7
acl-voice-cloning-fr-expanded1acl6060-voice-cloningacl-voice-cloning-fr-expanded2
ACL Voice Cloning FR — Expanded Pairs
Every (reference, target) segment pair within each speaker.
Audio is embedded directly — playable on the dataset viewer.
Column
Type
Role
ref_en_voice
🎤 Audio
English audio of reference segment (voice to clone)
ref_fr_voice
🎤 Audio
French cloned audio of reference segment
ref_en_text
string
English text of reference
ref_fr_text
string
French text of reference
trg_en_voice
🎤 Audio
English audio of target segment… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/acl-voice-cloning-fr-expanded2.acl6060-voice-cloning-fracl-voice-cloning-fr-datavoice_cloningvishal_voice_cloning_15secsacl-voice-cloning-fr-cleaneden_and_de_reference_voice_files_for_emotion_cloningacl6060-voice-cloning-fr-new-datavoice_cloning_dataset_fr_ar_zhvishal_voice_cloning_30secs_newQwen3-TTS-Cloning-VoicesAudio_cloning_Al_pacinoacl6060-voice-cloning-fr-newacl6060-voice-cloning-multilingualmax_voice_cloning_30secs_newtemp_data_cloning_outputmax_voice_cloning_30secsmax_voice_cloning_30secs_testacl6060-voice-cloning-fr-dataorpheus_voice_cloning
