CoolFace
Modelpublic

vlapky/TeraTTSv2

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes7downloads
Model Card

TeraTTSv2 ONNX

TeraTTSv2 is a self-contained ONNX Runtime text-to-speech release with selectable diffusion samplers, ten voice styles, Russian stress marking, and streamed audio output.

This release uses the clean English/Russian 25-second teacher and its matching eight-step CFG-3 distilled student.

Important — Russian stress is automatic. Text inside <ru>…</ru> receives stress markers automatically by default. Explicit + markers always win. Important — cross-language prompts. When an English reference voice is speaking Russian, experiment with duration_scale below 1 (for example 0.8). It is usually a better starting point than the default 1. Recommended voices: ru_f1 and ru_m5 are the preferred Russian voice prompts.

Installation

bash
pip install -r requirements.txt

sounddevice is only required for direct speaker playback. On Linux, install the system PortAudio library if it is not already present.

Load with Transformers

python
from transformers import AutoModel

tts = AutoModel.from_pretrained(
    "TeraSpace/TeraTTSv2",
    trust_remote_code=True,
    provider="CPUExecutionProvider",
    threads=6
)

waveform = tts.generate_speech(
    "<ru>Привет от TeraTTS.</ru>",
    voice="ru_f1",
    duration_scale=1,
)
tts.save_wav("teratts.wav", waveform)

waveform is a mono float32 NumPy array at 44,100 Hz. save_wav writes standard signed-16-bit PCM WAV without an extra audio package.

To inspect the exact text passed to the encoder after number expansion, stress marking, and Unicode normalization, call tts.normalize_text(text).

Controls

ControlValuesEffect
voiceru_f1 ★, ru_m5 ★, ru_f2, ru_m1, eng_f3, eng_f4_whisper, eng_f5, eng_m2_whisper, eng_m3, eng_m4Selects a bundled precomputed voice style named after its reference audio. ★ marks the recommended Russian prompts.
duration_scalePositive float, default 1Higher values produce slower, longer speech.
diffusion_modeldistilled (default), teacherDistilled is faster; teacher supports adjustable CFG.
ruaccent_modefull (default), dictionaryFull uses RUAccent neural ONNX graphs plus dictionaries; dictionary mode loads dictionaries only.

The default diffusion_model="distilled" is the fast eight-step sampler. To use the teacher sampler, choose it while loading:

python
teacher_tts = AutoModel.from_pretrained(
    "TeraSpace/TeraTTSv2",
    trust_remote_code=True,
    provider="CPUExecutionProvider",
    threads=6,
    diffusion_model="teacher",
)

guidance can be adjusted when generating with the teacher sampler. The distilled sampler has CFG 3 baked into its graph.

Language tags, numbers, and Russian stress

Language tags are required: wrap text in <en>…</en> or <ru>…</ru>. The runtime rejects untagged or unbalanced input with a tag-specific error. Before number expansion and stress marking, it inserts spaces after punctuation and between a number and a following word. Characters outside the model vocabulary are skipped with a runtime warning. Numbers inside language tags are expanded to words in the matching language before synthesis:

python
waveform = tts.generate_speech(
    "<ru>У меня 21 яблоко.</ru> <en>I have 42 apples.</en>",
    voice="ru_f1",
    duration_scale=1,
)

Russian text is automatically stress-marked by the bundled RUAccent-derived runtime. Manual + markers remain authoritative. For a lower-memory, deterministic dictionary-only path, choose the mode while loading:

python
dictionary_tts = AutoModel.from_pretrained(
    "TeraSpace/TeraTTSv2",
    trust_remote_code=True,
    provider="CPUExecutionProvider",
    threads=6,
    ruaccent_mode="dictionary",
)

Dictionary mode does not load RUAccent neural ONNX graphs. It marks known words and applies deterministic ё replacements, while unknown words and ambiguous homographs are left unchanged. Set russian_stress=False to disable automatic Russian stress processing entirely.

When using an English voice such as eng_f3 for Russian text, start by trying duration_scale=0.8 and adjust by ear:

python
waveform = tts.generate_speech(
    "<ru>Это русский текст английским голосом.</ru>",
    voice="eng_f3",
    duration_scale=0.8,
)

Stream audio

python
for chunk in tts.generate_speech_stream(
    "<en>Streaming speech is ready.</en>",
    voice="eng_f3",
    duration_scale=1,
):
    # Send float32 mono chunks (44,100 Hz) to a player or network client.
    consume(chunk)

The remote code loads only the selected sampler graph plus shared ONNX graphs. For security, pin a specific Hub commit when using trust_remote_code=True.

Attribution

The local Russian stress annotator and its assets are adapted from RUAccent, Copyright 2026 Denis Petrov, under the MIT License. See RUACCENT_NOTICE.txt.