humanify/Env-TTS-SD-Corpus
env-tts-sd-corpus Environment-aware text-to-speech training corpus. Each row pairs three short 16 kHz mono FLAC clips with a transcript: an environment sample (different speaker, same acoustic scene), a speaker reference (same speaker, optionally with augmented acoustics), the target speech, so a TTS model can learn to synthesise a target utterance with both a specified voice and a specified environment. Schema column type description… See the full description on the dataset page: https://huggingface.co/datasets/humanify/Env-TTS-SD-Corpus.
This repository belongs to humanify on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
