VibeVoice-Realtime-0.5B — with encoder (voice cloning)

microsoft/VibeVoice-Realtime-0.5B with the missing acoustic encoder added — enabling voice cloning from your own audio.

Usage

pip install "transformers==4.51.3" torch soundfile
pip install git+https://github.com/microsoft/VibeVoice

# get the scripts (the model itself downloads automatically on first run)
huggingface-cli download mohammed-bahumaish/vibevoice-realtime-0.5b-with-encoder \
    make_voice_prompt.py run_tts.py --local-dir .

# 1) build a voice prompt from ~15-30s of reference audio
python make_voice_prompt.py \
    --voice_wav my_voice.wav \
    --transcript "exact transcript of the reference audio" \
    --output my_voice.pt

# 2) speak anything in that voice
python run_tts.py \
    --voice_pt my_voice.pt \
    --text "Hello! This works with the stock Microsoft inference code." \
    --output out.wav

The .pt files are drop-in compatible with Microsoft's own demos, like the prebaked demo/voices/streaming_model/*.pt voices.

Tips

  • transformers must be 4.51.x — 5.x silently breaks the model.
  • Use a true 24 kHz+ recording, ≥ 15 s, clean single speaker.
  • Pass --transcript explicitly for best results (auto-transcription is English-only).

License

MIT. Base model by Microsoft; its responsible-use guidelines apply — clone only voices you have the right to use.

Downloads last month
415
Safetensors
Model size
1B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mohammed-bahumaish/vibevoice-realtime-0.5b-with-encoder

Finetuned
(17)
this model

Space using mohammed-bahumaish/vibevoice-realtime-0.5b-with-encoder 1

MiniMax H3 Video Generator 20 free credits · Text & image to video Try Free →