Bert-VITS2
commonvoice_16_1_bert_vits2
Cantonese Common Voice 16.1 for Bert-VITS2 fine tuning format
This dataset contains 14.5 hours of validated speech data in Cantonese (yue and zh-hk) from the Common Voice project, but with some cleansing and fixing of common Chinese characters, and used facebook/seamless-m4t-v2-large to cross check the data. The dataset is in the format required for fine-tuning the Bert-VITS2.
For more detail of cleansing, fixing and filtering, please refer to the notebook.
Data… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/commonvoice_16_1_bert_vits2.Style-Bert-VITS2-DatasetsKishidaFumio_voicedata_for_Bert-VITS2Style-Bert-VITS2-Datasets2Style-Bert-VITS2-bert_deberta-v2-large-japanese-char-wwmAbeShinzo_voicedata_for_Bert-VITS2
