CoolFace
Modelpublic

hanxxing/Bernini-R-S2V

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes8downloads
Model Card

Bernini-R-S2V

ComfyUI Bernini-R-S2V custom node v2 (update)

Bernini S2V Conditioning v2 - Bernini in-context video/image conditioning with masked lip-sync for one or two speakers. <video controls width="100%" height="480"> <source src="https://huggingface.co/rzgar/Bernini-R-S2V/resolve/main/video/ComfyUI__00003-audio.mp4" type="video/mp4"> Your browser does not support the video tag. </video>

Unzip ComfyUI-WanBerniniS2V_v2.zip into ComfyUI/custom_nodes/, then restart ComfyUI.

or save the Python files in: ComfyUI/custom_nodes/ComfyUI-WanBerniniS2V_v2/

Disable or remove the older ComfyUI-WanBerniniS2V folder if you only want v2.

One speaker

  • audio_1 + mask_1
  • Leave audio_2 / mask_2 unwired

Two speakers (dialog)

  • audio_1 + mask_1 - first speaker
  • audio_2 + mask_2 - second speaker
  • speaker_2_start_frame = -1 - second audio starts when the first clip ends

Masks

Paint on the output frame where each speaker's face appears. Masks control lip-sync placement only, they are not tied to reference_image_N slots. <img src="https://huggingface.co/rzgar/Bernini-R-S2V/resolve/main/ComfyUI-WanBerniniS2Vv2/demoassets/ScreenshotComfyUI-WanBerniniS2Vv2.png" width="1280" height="720" /> _______ <video controls width="100%" height="480"> <source src="https://huggingface.co/rzgar/Bernini-R-S2V/resolve/main/Bernini-R-S2V-FP8/video/ComfyUI_00005-audio.mp4" type="video/mp4"> Your browser does not support the video tag. </video>

Speech-driven video on Bernini-R , T2V, I2V, and V2V with lip-sync.

This model adds single-speaker audio support to Bernini-R, so you can drive video with speech in text-to-video, image-to-video, and video-to-video setups. It is not state-of-the-art audio-to-video, but it removes the need for post-processing or extra models just to add speech to Wan videos. For basic talking-head work, or longer videos built from short sequences, it is a handy all-in-one option on top of Bernini's motion and editing strengths.

ComfyUI detects these as `WAN22_S2V` / 'WanModel_S2V' (audio keys trigger S2V model type).

Usage

  1. 1.Download the checkpoints.
  2. 2.Place in:
   ComfyUI/models/diffusion_models/
  1. 1.Add wav2vec2 to:
   ComfyUI/models/audio_encoders/
  1. 1.Install [ComfyUI-WanBerniniS2V](https://huggingface.co/rzgar/Bernini-R-S2V/tree/main/ComfyUI-WanBerniniS2V) custom node:
   (Create a folder in 'ComfyUI/custom_nodes/'  named 'ComfyUI-WanBerniniS2V', then save the Python files in that folder.)
  1. 1.Restart ComfyUI.
  2. 2.Search for the Bernini S2V Conditioning node or use the example workfllow

Audio tips

SettingRecommendation
ChannelsMono, wav2vec2 downmixes stereo internally
Sample rate44.1 khz or 48 khz (resampled to 16 kHz)
ContentClear speech, less background music = better sync
Lengthin my testing, max 9 to 15 seconds