datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
so-vits-svc-4.0-ru-The_Witcher_3_Wild_HuntЭто тренировочные данные моделей голосов персонажей из "Ведьмак 3: Дикая охота" для so-vits-svc-4.1.1
commonvoice_16_1_bert_vits2
Cantonese Common Voice 16.1 for Bert-VITS2 fine tuning format
This dataset contains 14.5 hours of validated speech data in Cantonese (yue and zh-hk) from the Common Voice project, but with some cleansing and fixing of common Chinese characters, and used facebook/seamless-m4t-v2-large to cross check the data. The dataset is in the format required for fine-tuning the Bert-VITS2.
For more detail of cleansing, fixing and filtering, please refer to the notebook.
Data… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/commonvoice_16_1_bert_vits2.VITS_DATASET_1_60KVITS_DATASET_60k_100Kvits-datasetv14-fake-vitsv14-fake-vits-1KishidaFumio_voicedata_for_Bert-VITS2AbeShinzo_voicedata_for_Bert-VITS2so-vits-svc-4.0-Scrap_MechanicSugaYosihide_voicedata_for_Bert-VITS2wasm-dataset-sa-vitspro-v1
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/wasmdashai/wasm-dataset-sa-vitspro-v1.so-vits-svc-4.0-ru-Warcraft_III_ReforgedЭто тренировочные данные моделей голосов персонажей из "Warcraft III: Reforged" для so-vits-svc-4.1.1
vitsshata_rustaveli_vitsyaz_u_tygravai_shkury_output_original
Віцязь у тыгравай скуры — арыгінальнае аўдыё
Аўтар / Author: Шата РуставеліМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
shata_rustaveli_vitsyaz_u_tygravai_shkury_output
Доўгасць аўдыё
3h31m
Радкоў у датасеце
1,027
Структура
Кожны радок змяшчае:
audio —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/shata_rustaveli_vitsyaz_u_tygravai_shkury_output_original.viktar-tkachenka-nezvychainyia-prygody-vitsi-bobryka
Незвычайныя прыгоды Віці Бобрыка
Metadata
Author: Віктар Ткачэнка
Title: Незвычайныя прыгоды Віці Бобрыка
Narrator:
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/viktar-tkachenka-nezvychainyia-prygody-vitsi-bobryka.shata_rustaveli_vitsyaz_u_tygravai_shkury_all
Віцязь у тыгравай скуры
Аўтар / Author: Шата РуставеліМова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
1,091
Працягласць
3 гадз 24 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона… See the full description on the dataset page: https://huggingface.co/datasets/fosters/shata_rustaveli_vitsyaz_u_tygravai_shkury_all.ukr-male-speaker-0-v0-vitsVITS_NGOCHUYENinfore1-mms-vitsvits2so-vits-svc-4.1-7_Days_To_DieRus = Это данные для тренировки моделей голосов персонажей из "7 Days To Die", обученные для so-vits-svc-4.1.21
Eng = This is data for training character voice models from "7 Days To Die", trained for so-vits-svc-4.1.21
so-vits-svc-4.1-Dyson_Sphere_Programdarija-tts-vits-ready
Darija TTS Dataset (VITS-Ready)
This is a processed version of the Fahd1199/darija-tts dataset,
formatted specifically for finetuning VITS/MMS TTS models.
Dataset Structure
The dataset contains:
1740 training samples
436 validation samples
Each sample has:
audio: Path to a 16kHz WAV file
text: Darija (Moroccan Arabic) transcription
speaker_id: Speaker identifier (single speaker)
How to Use
This dataset is ready to be used with the finetune-hf-vits… See the full description on the dataset page: https://huggingface.co/datasets/Fahd1199/darija-tts-vits-ready.Shinku_Yuuki_so-vits-svc4.1
Nikaidou Shinku & Yuuki (so-vits-svc 4.1)
音声来源:いろとりどりのセカイ(福圆美里)、アストラエアの白き永遠(新田惠海)。
数据集时长: 4 小时
免责声明:本作品仅作为娱乐目的发布,可能造成的后果与使用的音声转换项目的作者、贡献者无关。
Zeta-Voice-ID-Vits-Finetuningdarija-tts-vits-readynepali_dataset_modified_for_vitsvits-owsmvits_dataset
