nguyen-brat/nguyen-ngoc-ngan-vieneu-tts-fine-tune
VieNeu-TTS Vietnamese Fine-tuned
A Vietnamese text-to-speech model fine-tuned from pnnbao-ump/VieNeu-TTS on ~80k segments of Vietnamese audiobook speech.
Architecture: Qwen2-0.5B language model backbone + NeuCodec neural vocoder (24 kHz). The LoRA adapter (r=32, alpha=64) is merged into the base weights for zero-overhead inference.
๐๏ธ Live demo: nguyen-brat/vieneu-tts-demo ๐ LoRA adapter only: nguyen-brat/vieneu-tts-vi
End-to-End Guide
Step 1: Install dependencies
# System packages (Ubuntu/Debian)
sudo apt-get install -y build-essential cmake libsndfile1 ffmpeg
# Install Rust (needed for sea-g2p phonemizer)
curl https://sh.rustup.rs -sSf | sh -s -- -y
source ~/.cargo/env
# Python packages
pip install vieneu soundfile numpyPython version: Requires Python 3.10โ3.12 (not 3.13) due to native extension wheels.
Step 2: Run inference
Choose one of the methods below depending on your hardware.
Method A: PyTorch inference (GPU recommended)
Best for GPU machines. Uses the full BF16 model (1.1 GB).
from vieneu import Vieneu
import soundfile as sf
# Load the fine-tuned model (auto-downloads from HF Hub)
tts = Vieneu(backbone_repo="nguyen-brat/VieNeu-TTS-Vietnamese-Finetuned")
# Generate speech
audio = tts.infer("Xin chao, day la giong noi tieng Viet duoc tong hop bang AI.")
sf.write("output.wav", audio, samplerate=24000)
print("Saved output.wav")With GPU explicitly:
import torch
from vieneu import Vieneu
device = "cuda" if torch.cuda.is_available() else "cpu"
tts = Vieneu(
backbone_repo="nguyen-brat/VieNeu-TTS-Vietnamese-Finetuned",
backbone_device=device,
codec_device=device,
)
audio = tts.infer("Xin chao the gioi!")Method B: GGUF inference (CPU or GPU)
Best for machines without a GPU, or for faster inference with quantized weights (441 MB). 3x faster than PyTorch on GPU, and runs well on CPU.
# Extra dependency for GGUF
pip install llama-cpp-python transformers huggingface-hubGPU GGUF: To use GPU acceleration, install with CUDA support: ``bash CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python --force-reinstall ``from huggingface_hub import hf_hub_download
from llama_cpp import Llama
from transformers import AutoTokenizer
from vieneu import Vieneu
import soundfile as sf
import numpy as np
# --- Config ---
REPO = "nguyen-brat/VieNeu-TTS-Vietnamese-Finetuned"
SPEECH_OFFSET = 151671 # <|speech_0|> token ID
SPEECH_END = 151670 # <|SPEECH_GENERATION_END|> token ID
SPEECH_MAX = 65535
# --- Step 1: Load model components ---
gguf_path = hf_hub_download(repo_id=REPO, filename="vieneu_q4_k_m.gguf")
# n_gpu_layers=-1 for full GPU offload, 0 for CPU-only
llm = Llama(model_path=gguf_path, n_gpu_layers=0, n_ctx=32768, verbose=False)
tokenizer = AutoTokenizer.from_pretrained(REPO)
vieneu = Vieneu(backbone_repo=REPO) # for codec + phonemizer + voices
# --- Step 2: Prepare input ---
text = "Xin chao, day la giong noi tieng Viet."
# Get a preset voice (or use ref_audio for voice cloning)
voice = vieneu.get_preset_voice() # uses default voice
ref_codes = voice["codes"]
ref_text = voice.get("text", "")
# Normalize and phonemize
from vieneu_utils.phonemize_text import phonemize_with_dict
text_normalized = vieneu.normalizer.normalize(text)
ref_phonemes = vieneu.get_ref_phonemes(ref_text)
chunk_phonemes = phonemize_with_dict(text_normalized, skip_normalize=True)
# --- Step 3: Build prompt ---
codes_str = "".join(f"<|speech_{int(c)}|>" for c in ref_codes.flatten())
prompt = (
f"user: Convert the text to speech:"
f"<|TEXT_PROMPT_START|>{ref_phonemes} {chunk_phonemes}<|TEXT_PROMPT_END|>\n"
f"assistant:<|SPEECH_GENERATION_START|>{codes_str}"
)
prompt_ids = tokenizer.encode(prompt, add_special_tokens=False)
# --- Step 4: Generate speech tokens ---
speech_ids = []
for token_id in llm.generate(prompt_ids, top_k=50, temp=1.0, reset=True):
# Ignore end token until at least 50 speech tokens collected
if token_id == SPEECH_END and len(speech_ids) >= 50:
break
if SPEECH_OFFSET <= token_id <= SPEECH_OFFSET + SPEECH_MAX:
speech_ids.append(token_id)
if len(speech_ids) >= 16000:
break
# --- Step 5: Decode to audio ---
decode_str = "".join(f"<|speech_{tid - SPEECH_OFFSET}|>" for tid in speech_ids)
audio = vieneu._decode(decode_str)
sf.write("output.wav", audio, samplerate=24000)
print(f"Generated {len(audio)/24000:.1f}s of audio ({len(speech_ids)} tokens)")Voice Cloning (zero-shot)
Clone any speaker's voice from a 3-10 second reference clip:
from vieneu import Vieneu
import soundfile as sf
tts = Vieneu(backbone_repo="nguyen-brat/VieNeu-TTS-Vietnamese-Finetuned")
audio = tts.infer(
text="Ngay xua, co mot chu be ten la An song trong mot ngoi lang nho ven song.",
ref_audio="reference.wav", # 3-10 second WAV of target speaker
ref_text="Transcript of reference.", # optional but improves quality
)
sf.write("cloned.wav", audio, samplerate=24000)Preset Voices
6 built-in Vietnamese speaker voices โ no reference audio needed:
from vieneu import Vieneu
tts = Vieneu(backbone_repo="nguyen-brat/VieNeu-TTS-Vietnamese-Finetuned")
# List available voices
for display_name, key in tts.list_preset_voices():
print(f" {key}: {display_name}")
# Use a specific preset
voice = tts.get_preset_voice("binh") # binh | tuyen | nguyen | huong | ngoc | doan
audio = tts.infer(text="Xin chao!", voice=voice)Long-form Synthesis (stories / articles)
For texts longer than ~250 characters, split into chunks with automatic pause insertion:
from vieneu import Vieneu
import soundfile as sf
tts = Vieneu(backbone_repo="nguyen-brat/VieNeu-TTS-Vietnamese-Finetuned")
story = """Ngay xua, co mot chu be ten la An song trong mot ngoi lang nho ven song.
Moi sang, An thuong day som de giup me ganh nuoc va nau com.
Cau thich nhat la duoc chay ra bo song, ngoi ngam nhung con ca boi loi duoi lan nuoc trong vat."""
# Synthesize each chunk (max 100 chars) with pauses between sentences/paragraphs
chunks = []
for i in range(0, len(story), 200):
chunk = story[i:i+200].strip()
if chunk:
audio = tts.infer(chunk)
chunks.append(audio)
import numpy as np
silence = np.zeros(int(24000 * 0.3)) # 300ms pause
full_audio = np.concatenate([np.concatenate([c, silence]) for c in chunks])
sf.write("story.wav", full_audio, samplerate=24000)Performance
RTF (real-time factor): lower is faster. 0.4x means 10s of audio generated in 4s.
Training Details
Dataset
- Source: Vietnamese audiobooks from YouTube
- Pipeline: Demucs source separation, MossFormer2 denoising, WhisperX diarization, speaker verification, PhoWhisper transcription, DNSMOS quality filter
- Size: ~79,743 training segments + ~4,200 validation segments
- Segment length: 3-15 seconds
- Quality threshold: DNSMOS >= 3.0, Whisper log-prob >= -1.0
Fine-tuning Config
Training Curve
Best validation loss: 5.826 at checkpoint 24500.
Hardware
- 2x NVIDIA GPU
- DeepSpeed ZeRO-2 with optimizer CPU offload
Files in This Repo
License
Apache 2.0 โ see base model pnnbao-ump/VieNeu-TTS for upstream licensing.
