datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voice-light-synthetic-audio
Voice-Light Synthetic Audio
English-only synthetic conversational speech for training and evaluating streaming
turn-taking models. The corpus focuses on end-of-turn prediction, continuation holds,
short backchannels, interruptions, and response timing.
The dataset contains user-side FLAC speech units plus typed conversation plans,
rendering provenance, quality ledgers, and deterministic reconstruction metadata.
Assistant speech is represented as a time-varying… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-synthetic-audio.commonvoice_16_1_bert_vits2
Cantonese Common Voice 16.1 for Bert-VITS2 fine tuning format
This dataset contains 14.5 hours of validated speech data in Cantonese (yue and zh-hk) from the Common Voice project, but with some cleansing and fixing of common Chinese characters, and used facebook/seamless-m4t-v2-large to cross check the data. The dataset is in the format required for fine-tuning the Bert-VITS2.
For more detail of cleansing, fixing and filtering, please refer to the notebook.
Data… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/commonvoice_16_1_bert_vits2.w2v-bert-2.0-nepali-transliteratorw2v-bert-2.0-nepaliKishidaFumio_voicedata_for_Bert-VITS2AbeShinzo_voicedata_for_Bert-VITS2SugaYosihide_voicedata_for_Bert-VITS2HerbeetcelsoHerbeetHenriquebertespanholherbetbetovozmgiprtimvozherbertvipHerbertNarativavozprtimhebertespanholprtimespanlholberteHerbeetVanderleybert-vits2bertespanholHerbertbettoHerbeebetoHerbeetoGcelsobetbert_vits2
