datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
acl-voice-cloning-fr-expandedtry
ACL Voice Cloning FR — Expanded Pairs
Every (reference, target) segment pair within each speaker.
Audio is embedded directly — playable on the dataset viewer.
Column
Type
Role
ref_en_voice
🎤 Audio
English audio of reference segment (voice to clone)
ref_fr_voice
🎤 Audio
French cloned audio of reference segment
ref_en_text
string
English text of reference
ref_fr_text
string
French text of reference
trg_en_voice
🎤 Audio
English audio of target segment… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/acl-voice-cloning-fr-expandedtry.acl-voice-cloning-fr-expanded
ACL Voice Cloning FR — Expanded Pairs
Every (reference, target) segment pair within each speaker.
Audio is embedded directly — playable on the dataset viewer.
Column
Type
Role
ref_en_voice
🎤 Audio
English audio of reference segment (voice to clone)
ref_fr_voice
🎤 Audio
French cloned audio of reference segment
ref_en_text
string
English text of reference
ref_fr_text
string
French text of reference
trg_en_voice
🎤 Audio
English audio of target segment… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/acl-voice-cloning-fr-expanded.robot-tts-profiles
Robot TTS Profiles (zero-train)
Damaged prompt clips + Chatterbox samples generated with cfg_weight=0.1 so the
model largely inherits the prompt's degradation instead of repairing it.
NOTE: cfg_weight=0.0 crashes chatterbox-tts 0.1.7 (t3.py hardcodes a CFG batch
of two); use 0.05-0.1 as the low-CFG setting.
prompts/ - one damaged reference clip per profile (scifi_robot, retro_8bit, intercom, telephone)
samples/ - generated speech per profile
base_clean.wav - the clean control… See the full description on the dataset page: https://huggingface.co/datasets/ACloudCenter/robot-tts-profiles.acl-6060-openvoiceacl-6060
ACL 60/60
Dataset details
ACL 60/60 evaluation sets for multilingual translation of ACL 2022 technical presentations into 10 target languages.
Citation
@inproceedings{salesky-etal-2023-evaluating,
title = "Evaluating Multilingual Speech Translation under Realistic Conditions with Resegmentation and Terminology",
author = "Salesky, Elizabeth and
Darwish, Kareem and
Al-Badrashiny, Mohamed and
Diab, Mona and
Niehues, Jan"… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/acl-6060.Sukuma-Voices-ACL
Sukuma Voices Dataset 🎙️
The first publicly available speech corpus for Sukuma (Kisukuma), a Bantu language spoken by approximately 10 million people in northern Tanzania. This dataset supports speech-to-text, text-to-speech, and speech evaluation tasks.
Dataset Description
Sukuma Voices addresses the critical gap in speech technology resources for one of Africa's most severely under-resourced languages. The dataset includes both human recordings and… See the full description on the dataset page: https://huggingface.co/datasets/sartifyllc/Sukuma-Voices-ACL.acl-voice-cloning-fr-cleaned-v2
ACL Voice Cloning FR — Cleaned & Expanded (V2)
Source filtered with GPU-accelerated quality checks (SNR≥10.0dB, silence≤65%),
then expanded into all (reference, target) pairs per speaker.
Filter
Threshold
Duration
1.0-20.0s
SNR
≥ 10.0 dB
Silence
≤ 65%
Text length
5-500 chars
Split
Source kept
Expanded pairs
Shards
train
748
70,006
43
test
119
7,022
7
acl-abstracts-3k-audioacl-abstracts-3k-audio-deacl-voice-cloning-fr-expanded1acl6060-voice-cloningacl-voice-cloning-fr-expanded2
ACL Voice Cloning FR — Expanded Pairs
Every (reference, target) segment pair within each speaker.
Audio is embedded directly — playable on the dataset viewer.
Column
Type
Role
ref_en_voice
🎤 Audio
English audio of reference segment (voice to clone)
ref_fr_voice
🎤 Audio
French cloned audio of reference segment
ref_en_text
string
English text of reference
ref_fr_text
string
French text of reference
trg_en_voice
🎤 Audio
English audio of target segment… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/acl-voice-cloning-fr-expanded2.acl6060-voice-cloning-fracl-voice-cloning-fr-dataacl-voice-cloning-fr-cleanedacl6060-voice-cloning-fr-new-datarasst-demo-acl6060-zh-segments
RASST Demo ACL6060 Chinese Streaming Segments
This dataset contains the prepared ACL6060 en-to-zh streaming evaluation
segments used by the RASST demo and sglang-omni thinker decode profiling
workflow.
It is a convenience bundle for reproducing N=32 streaming load experiments. The
source ACL6060 data is public; this repository packages derived 16 kHz mono WAV
clips plus SimulEval-compatible source and target lists.
Contents
seg/: 468 cropped WAV segments.… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rasst-demo-acl6060-zh-segments.acl6060-voice-cloning-multilingualacl6060-voice-cloning-fr-newacl6060-voice-cloning-fr-dataacl-voice-cloning-fr-sparse-v1
Sparse French Voice Cloning Dataset
Reduced redundancy (1 ref per target) for fast LoRA experimentation.
