mohammed-bahumaish/vibevoice-realtime-0.5b-with-encoder
4438
VibeVoice-Realtime-0.5B — with encoder (voice cloning)
microsoft/VibeVoice-Realtime-0.5B with the missing acoustic encoder added — enabling voice cloning from your own audio.
Usage
pip install "transformers==4.51.3" torch soundfile
pip install git+https://github.com/microsoft/VibeVoice
# get the scripts (the model itself downloads automatically on first run)
huggingface-cli download mohammed-bahumaish/vibevoice-realtime-0.5b-with-encoder \
make_voice_prompt.py run_tts.py --local-dir .
# 1) build a voice prompt from ~15-30s of reference audio
python make_voice_prompt.py \
--voice_wav my_voice.wav \
--transcript "exact transcript of the reference audio" \
--output my_voice.pt
# 2) speak anything in that voice
python run_tts.py \
--voice_pt my_voice.pt \
--text "Hello! This works with the stock Microsoft inference code." \
--output out.wavThe .pt files are drop-in compatible with Microsoft's own demos, like the prebaked demo/voices/streaming_model/*.pt voices.
Tips
transformersmust be 4.51.x — 5.x silently breaks the model.- Use a true 24 kHz+ recording, ≥ 15 s, clean single speaker.
- Pass
--transcriptexplicitly for best results (auto-transcription is English-only).
License
MIT. Base model by Microsoft; its responsible-use guidelines apply — clone only voices you have the right to use.
