yasoukyoku/utai-runtimes
Utai runtimes and models Download mirror for Utai (UtaiSynthesizer): Python training runtimes (wheels/, runtime packs), and the model files the app fetches on demand (models/). The application's own source licence (AGPL-3.0) covers the application code only and conveys no rights over the files below; each keeps the terms stated here. Models trained by the Utai project models/auxiliary/score2cv_768.onnx, score2cv_256.onnx (ScoreToCV) and autotune_a1.onnx (automatic… See the full description on the dataset page: https://huggingface.co/datasets/yasoukyoku/utai-runtimes.
Utai runtimes and models
Download mirror for Utai (UtaiSynthesizer): Python training runtimes (wheels/, runtime packs), and the model files the app fetches on demand (models/). The application's own source licence (AGPL-3.0) covers the application code only and conveys no rights over the files below; each keeps the terms stated here.
Models trained by the Utai project
models/auxiliary/score2cv_768.onnx, score2cv_256.onnx (ScoreToCV) and autotune_a1.onnx (automatic pitch tuning), plus the copies under models/aux/.
These are the project's own weights, but they were trained on third-party singing corpora whose terms carry over. Training set: 44,947 clips.
Credits (verbatim, as required by the corpora's terms): 『©SSS』 · 『歌声DB制作:アマノケイ 音声提供者: 霧野蒼太』 · 『DB制作:おふとんP』 · 『御丹宮くるみ歌声データべース』 · GTSinger (Zhang et al., 2024) · M4Singer (Zhang et al., 2022) · PJS (Koguchi & Takamichi, 2020) · 日本声優統計学会.
Licence position / use: consistent with Creative Commons' guidance that in many cases an AI model is not an adaptation of its training works (https://creativecommons.org/using-cc-licensed-works-for-ai-training/), we do not treat these weights as Adapted Material of the CC-licensed corpora, so the ShareAlike conditions of GTSinger / M4Singer (BY-NC-SA) and PJS (BY-SA) do not attach to them. The models are nevertheless provided for non-commercial use only, with the attribution above: several Japanese corpora's usage agreements require it, and 94% of the training data is NonCommercial-licensed.
Audio generated with these models is the 「出力音声」 (output voice) of the Natsume Yuuri DB and falls under 「夏目悠李の出力音声に関する利用規約」 (ATSUYA; https://ksdcm1ng.wixsite.com/njksofficial/%E8%A6%8F%E7%B4%84-rules), shipped unmodified beside the model files as NATSUME_OUTPUT_VOICE_TERMS.txt: commercial use needs separate permission, and generated audio must not be used to build acoustic/pitch models (including as training data in Utai's own training feature).
Third-party model weights mirrored here
Mirrored for availability, under their own licences (attribution preserved, same terms):
- NSF-HiFiGAN (OpenVPI) vocoder — CC BY-NC-SA 4.0 (
NOTICE.txt/NOTICE.zh-CN.txtbeside the files). - GAME vocal-to-MIDI — CC BY-NC-SA.
- ContentVec, RMVPE, and the RVC / so-vits-svc pretrained bases under
models/training/— their upstream licences.
Runtime packs and wheels
Third-party Python packages redistributed as-is under their own licences (PyTorch, ROCm/TheRock wheels, etc.).
