humanify/Env-TTS-Clean
Env-TTS-Clean Environment-aware text-to-speech training corpus (clean release). Each row pairs four short 24 kHz mono FLAC clips with aligned transcripts: an environment sample (different speaker, same acoustic scene), a speaker reference (same speaker as the target utterance), a speaker-enhanced copy of the reference (MossFormer2 enhancement — or, for the DDS source, the real clean-studio recording of the speaker reference), the target speech to synthesise, so a model can… See the full description on the dataset page: https://huggingface.co/datasets/humanify/Env-TTS-Clean.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face