cicada-ai/Chanjing-Avatar-14B
117
Chanjing-Avatar 14B
Chanjing-Avatar 14B is an audio-driven 720p avatar video generation model based on Wan2.1-T2V-14B. It adds audio conditioning and LoRA adapters to the Wan video diffusion model.
Source code and complete inference instructions: chanjing-ai/Chanjing-Avatar
Chanjing-Avatar Model Family
- Chanjing-Avatar 14B: 720p image-to-video generation from a reference image and driving audio.
- Chanjing-Avatar V2V 5B: video-to-video generation that preserves source motion and regenerates the speaking face.
- Chanjing-Avatar V2V 1.3B: a lighter video-to-video model for audio-driven face animation.
The checkpoint contains audio modules, input projection, and LoRA adapters in BF16. The Wan2.1 base model and Wav2Vec audio encoder are required separately.
Chanjing-Avatar-14B/
|-- config.json
`-- diffusion_pytorch_model.safetensorshf download cicada-ai/Chanjing-Avatar-14B \
--local-dir models/Chanjing-Avatar-14BUsers are responsible for obtaining consent for source images and voices and for clearly disclosing synthetic media.
