CoolFace
Modelpublic

soniqo/Supertonic-3-LiteRT

sourceHugging Faceopenrailupdated 23d agoView on Hugging Face
3likes447downloads
Model Card

SupertonicTTS-3 — LiteRT (.tflite, Android / Qualcomm NPU)

First-party LiteRT export of Supertonic-3's four non-autoregressive flow-matching graphs. Built by our own pipeline (speech-models/stmodels): weights lifted from the `Supertone/supertonic-3` ONNX initializers → PyTorch nn.Module`litert_torch.convert` (torch.export → StableHLO → TFLite). This avoids the onnx2tf NCHW/ConvNeXt layout failures that block direct ONNX→TFLite for this model.

Graphs & parity (FP32, vs ONNX Runtime)

Moduletfliteparity max\Δ\
duration_predictor.tflite3.4 MB4.1e-05 ✓
vector_estimator.tflite (ODE denoiser)244 MB5.6e-03 ✓
vocoder.tflite97 MB2.6e-04 ✓
text_encoder.tflite34 MB1.1e-01 (localized; mean ~2.5e-4) ⚠️
vector_estimator_L128.tflite (L=128 bucket)244 MB1.7e-02 (relative 2.7e-3) ✓
vocoder_L128.tflite (L=128 bucket)97 MB2.7e-04 ✓

Shapes are fixed (T=128 text tokens; latent length L per graph) — litert_torch cannot lower the relpos attention + ConvNeXt stack with a symbolic L, so the latent axis ships in buckets: the base pair is L=64 (≈4.5 s of audio), and *_L128.tflite is the same two L-dependent graphs exported at L=128 (≈9 s). Pick the smallest bucket whose window holds a piece's predicted duration — short sentences run on the cheap L=64 pair, a long sentence is generated in one pass on L=128 instead of being split (speech-core's LiteRTSupertonicTts does this automatically and loads the L=128 pair on first use). The host runs the flow-matching ODE loop (vector_estimator ×total_steps). Assets to drive them: tts.json, unicode_indexer.json (G2P-free tokenizer), voice_styles/*.json.

Graph inputs are bound by tensor name (serving_default_args_N, N = signature position): the converter permutes input slots differently per toolchain version, so do not rely on slot order.

Running on Android / Qualcomm NPU

  • CPU/GPU: LiteRT (ai_edge_litert / TFLite) interpreter with XNNPACK/GPU delegate.
  • Qualcomm HTP/NPU: the LiteRT QNN delegate at runtime, or compile to a QNN context binary via Qualcomm AI Hub (qai_hub) from these graphs (static shapes are HTP-friendly). int8/int4 PTQ via ai-edge-quantizer for full HTP residency is a follow-up.

Attribution & license

Other Supertonic-3 formats

Ecosystem

  • **soniqo.audio** — website / use-case explorer (transcription, voice cloning, live ASR, voice agents).
  • **speech-core** — C++ orchestration library; Supertonic plugs in as a TTSInterface LiteRT model.
  • **speech-swift** — Apple Silicon MLX + CoreML runtime.
  • **speech-android** — Android SDK consuming on-device LiteRT bundles.

Other LiteRT models in this collection