Gigsu/vocoloco-onnx
6
VocoLoco — OmniVoice ONNX Models
ONNX exports of k2-fsa/OmniVoice for browser-based text-to-speech inference via ONNX Runtime Web.
Models
Usage
These models are designed to run in the browser via VocoLoco, a fully client-side TTS application. No server required.
Architecture
- Backbone: Qwen3-0.6B (28 transformer layers)
- Audio codec: HiggsAudioV2 (8 codebooks, 24kHz output)
- Generation: Iterative masked diffusion (configurable 8-32 steps)
- Voice cloning: Zero-shot via reference audio encoding
- Voice design: Text-based control (gender, pitch, accent)
License
Apache 2.0 — same as the original OmniVoice model.
Attribution
Based on OmniVoice by Xiaomi Corp (k2-fsa).
