env-tts
Datasets
All datasets matching “env-tts”Env-TTS-Clean
Env-TTS-Clean
Environment-aware text-to-speech training corpus (clean release). Each row
pairs four short 24 kHz mono FLAC clips with aligned transcripts:
an environment sample (different speaker, same acoustic scene),
a speaker reference (same speaker as the target utterance),
a speaker-enhanced copy of the reference (MossFormer2 enhancement — or, for
the DDS source, the real clean-studio recording of the speaker reference),
the target speech to synthesise,
so a model can… See the full description on the dataset page: https://huggingface.co/datasets/humanify/Env-TTS-Clean.Env-TTS-SD-Corpus-24K-Enhancedenv_tts_data_samplesEnv-TTS-SD-Corpus
env-tts-sd-corpus
Environment-aware text-to-speech training corpus. Each row pairs three short
16 kHz mono FLAC clips with a transcript:
an environment sample (different speaker, same acoustic scene),
a speaker reference (same speaker, optionally with augmented acoustics),
the target speech,
so a TTS model can learn to synthesise a target utterance with both a specified
voice and a specified environment.
Schema
column
type
description… See the full description on the dataset page: https://huggingface.co/datasets/humanify/Env-TTS-SD-Corpus.envtts_evaltts_env
