CoolFace
Modelpublic

aaaaaaaaaff/free-svc

sourceHugging Facecc-by-nc-sa-4.0updated 6mo agoView on Hugging Face
0likes10downloads
Model Card

FreeSVC: Zero-shot Multilingual Singing Voice Conversion

FreeSVC is a promising multilingual zero-shot singing voice conversion model. It enables the conversion of singing voices across languages without the need for extensive language-specific training. GitHub repository. Paper arXiv pre-print.

Supported Languages

LanguageIDStatusSpeech DataSinging Data
Chinese0✅ Full255h70h
Dutch1✅ FullPart of CML-
English2✅ Full921h47h
French3✅ FullPart of CML-
German4✅ FullPart of CML-
Italian5✅ FullPart of CML-
Japanese6✅ Full30h-
Other*7⚠️ Partial-10h
Polish8✅ FullPart of CML-
Portuguese9✅ FullPart of CML-
Spanish10✅ FullPart of CML-

*Note: The "Other" category is used for vocal techniques without content.

Model Overview

FreeSVC leverages an enhanced VITS architecture integrated with Speaker-invariant Clustering (SPIN) and the ECAPA2 speaker encoder. This combination effectively separates speaker characteristics from linguistic content, ensuring high-quality and natural-sounding voice conversions across multiple languages.

Training Datasets

FreeSVC was trained on a diverse set of speech and singing datasets covering multiple languages:

**Dataset****Hours****Language****Type**
AISHELL-1170hChineseSpeech
AISHELL-385hChineseSpeech
CML-TTS3.1k7 LanguagesSpeech
HiFiTTS292hEnglishSpeech
JVS30hJapaneseSpeech
LibriTTS-R585hEnglishSpeech
NUS (NHSS)7hEnglishSpeech, Singing
OpenSinger50hChineseSinging
Opencpop5hChineseSinging
PopBuTFy10h, 40hChinese, EnglishSinging
POPCS5hChineseSinging
VCTK44hEnglishSpeech
VocalSet10hOtherSinging

License

FreeSVC is released under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license. This means:

  • The model can only be used for research and non-commercial purposes. Any commercial use is strictly prohibited.
  • Any derivative works must be shared under the same license.
  • Proper attribution must be given when using the model.

Users must also comply with the licenses of the original datasets used for training. Some datasets may have additional restrictions beyond CC BY-NC-SA 4.0. Ensure you review and adhere to their terms before using the model.

For full details, refer to the CC BY-NC-SA 4.0 License.

Citation

@INPROCEEDINGS{10890068,
  author={Ferreira, Alef Iury and Gris, Lucas Rafael and Da Rosa, Augusto and Oliveira, Frederico and Casanova, Edresson and Sousa, Rafael and Junior, Arnaldo and Soares, Anderson and Filho, Arlindo Galvão},
  booktitle={ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, 
  title={FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion}, 
  year={2025},
  volume={},
  number={},
  pages={1-5},
  keywords={Training;Source coding;Zero shot learning;Refining;Signal processing;Data models;Acoustics;Multilingual;Data mining;Speech synthesis;Singing Voice Conversion;Synthesis of Singing Voices;Cross-lingual and multilingual aspects in speech synthesis},
  doi={10.1109/ICASSP49660.2025.10890068}}