BertilBraun/voice-light-synthetic-audio
Voice-Light Synthetic Audio English-only synthetic conversational speech for training and evaluating streaming turn-taking models. The corpus focuses on end-of-turn prediction, continuation holds, short backchannels, interruptions, and response timing. The dataset contains user-side FLAC speech units plus typed conversation plans, rendering provenance, quality ledgers, and deterministic reconstruction metadata. Assistant speech is represented as a time-varying… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-synthetic-audio.
01.9k
