datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LemonfootVoiceDatasets
Kit Lemonfoot's Voice Datasets
This repository aims to house every dataset used in my AI voice models. Please credit me if you use these datasets.
This includes the following:
Every dataset used in my RVC models
Every dataset used in my Style BertVITS2 models
Every dataset used in my GPT-SoVITS models
Multiple datasets submitted to 15.ai (RIP)
Reference
Hand-transcribed datasets were built for 15.ai and as such are in the LJSpeech transcription format.… See the full description on the dataset page: https://huggingface.co/datasets/Kit-Lemonfoot/LemonfootVoiceDatasets.LEMAS-Dataset-eval
Overview
This dataset is part of LEMAS-Project(lemas-project.github.io/LEMAS-Project).
It contains a large-scale training set (150k+ hours) and a curated evaluation set
(500 utterances per language) covering 10 languages, all with word-level alignment.
Fields
key: unique utterance identifier; the first two characters indicate the language ID
audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-eval.ECG_PCGportuguese_audioslemas-italian-train-speechnoisy-datasetLEMAS-Dataset-eval-esportuguese_audios_maleFine-Tuning-03conradovozDataset-11-01-2026LemonSoda200voz_lembredemim
