CoolFace
Modelpublic

DonatoDiffusion/Wan2.1-Fun-1.3B-InP

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
1likes42downloads
README_en.md249 linesDownload Raw Back to root
1---2license: apache-2.03language:4- en5- zh6pipeline_tag: text-to-video7library_name: diffusers8tags:9- video10- video-generation11---12 13# Wan-Fun14 15๐Ÿ˜Š Welcome!16 17[![Hugging Face Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Spaces-yellow)](https://huggingface.co/spaces/alibaba-pai/Wan2.1-Fun-1.3B-InP)18 19[![Github](https://img.shields.io/badge/๐ŸŽฌ%20Code-Github-blue)](https://github.com/aigc-apps/VideoX-Fun)20 21[English](./README_en.md) | [็ฎ€ไฝ“ไธญๆ–‡](./README.md)22 23# Table of Contents24- [Table of Contents](#table-of-contents)25- [Model zoo](#model-zoo)26- [Video Result](#video-result)27- [Quick Start](#quick-start)28- [How to use](#how-to-use)29- [Reference](#reference)30- [License](#license)31 32# Model zoo33V1.0:34| Name | Storage Space | Hugging Face | Model Scope | Description |35|--|--|--|--|--|36| Wan2.1-Fun-1.3B-InP | 19.0 GB | [๐Ÿค—Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-InP) | [๐Ÿ˜„Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-InP) | Wan2.1-Fun-1.3B text-to-video weights, trained at multiple resolutions, supporting start and end frame prediction. |37| Wan2.1-Fun-14B-InP | 47.0 GB | [๐Ÿค—Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-InP) | [๐Ÿ˜„Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-InP) | Wan2.1-Fun-14B text-to-video weights, trained at multiple resolutions, supporting start and end frame prediction. |38| Wan2.1-Fun-1.3B-Control | 19.0 GB | [๐Ÿค—Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-Control) | [๐Ÿ˜„Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-Control) | Wan2.1-Fun-1.3B video control weights, supporting various control conditions such as Canny, Depth, Pose, MLSD, etc., and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction at 81 frames, trained at 16 frames per second, with multilingual prediction support. |39| Wan2.1-Fun-14B-Control | 47.0 GB | [๐Ÿค—Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-Control) | [๐Ÿ˜„Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-Control) | Wan2.1-Fun-14B video control weights, supporting various control conditions such as Canny, Depth, Pose, MLSD, etc., and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction at 81 frames, trained at 16 frames per second, with multilingual prediction support. |40 41# Video Result42 43### Wan2.1-Fun-14B-InP && Wan2.1-Fun-1.3B-InP44 45<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">46  <tr>47      <td>48          <video src="https://github.com/user-attachments/assets/bd72a276-e60e-4b5d-86c1-d0f67e7425b9" width="100%" controls autoplay loop></video>49      </td>50       <td>51          <video src="https://github.com/user-attachments/assets/cb7aef09-52c2-4973-80b4-b2fb63425044" width="100%" controls autoplay loop></video>52     </td>53      <td>54          <video src="https://github.com/user-attachments/assets/4e10d491-f1cf-4b08-a7c5-1e01e5418140" width="100%" controls autoplay loop></video>55      </td>56      <td>57          <video src="https://github.com/user-attachments/assets/f7e363a9-be09-4b72-bccf-cce9c9ebeb9b" width="100%" controls autoplay loop></video>58     </td>59  </tr>60</table>61 62<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">63  <tr>64      <td>65          <video src="https://github.com/user-attachments/assets/28f3e720-8acc-4f22-a5d0-ec1c571e9466" width="100%" controls autoplay loop></video>66      </td>67      <td>68          <video src="https://github.com/user-attachments/assets/fb6e4cb9-270d-47cd-8501-caf8f3e91b5c" width="100%" controls autoplay loop></video>69      </td>70       <td>71          <video src="https://github.com/user-attachments/assets/989a4644-e33b-4f0c-b68e-2ff6ba37ac7e" width="100%" controls autoplay loop></video>72     </td>73      <td>74          <video src="https://github.com/user-attachments/assets/9c604fa7-8657-49d1-8066-b5bb198b28b6" width="100%" controls autoplay loop></video>75     </td>76  </tr>77</table>78 79### Wan2.1-Fun-14B-Control && Wan2.1-Fun-1.3B-Control80 81<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">82  <tr>83      <td>84          <video src="https://github.com/user-attachments/assets/f35602c4-9f0a-4105-9762-1e3a88abbac6" width="100%" controls autoplay loop></video>85      </td>86      <td>87          <video src="https://github.com/user-attachments/assets/8b0f0e87-f1be-4915-bb35-2d53c852333e" width="100%" controls autoplay loop></video>88      </td>89       <td>90          <video src="https://github.com/user-attachments/assets/972012c1-772b-427a-bce6-ba8b39edcfad" width="100%" controls autoplay loop></video>91     </td>92  <tr>93</table>94 95<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">96  <tr>97      <td>98          <video src="https://github.com/user-attachments/assets/53002ce2-dd18-4d4f-8135-b6f68364cabd" width="100%" controls autoplay loop></video>99      </td>100      <td>101          <video src="https://github.com/user-attachments/assets/a1a07cf8-d86d-4cd2-831f-18a6c1ceee1d" width="100%" controls autoplay loop></video>102      </td>103       <td>104          <video src="https://github.com/user-attachments/assets/3224804f-342d-4947-918d-d9fec8e3d273" width="100%" controls autoplay loop></video>105     </td>106  <tr>107      <td>108          <video src="https://github.com/user-attachments/assets/c6c5d557-9772-483e-ae47-863d8a26db4a" width="100%" controls autoplay loop></video>109      </td>110      <td>111          <video src="https://github.com/user-attachments/assets/af617971-597c-4be4-beb5-f9e8aaca2d14" width="100%" controls autoplay loop></video>112      </td>113       <td>114          <video src="https://github.com/user-attachments/assets/8411151e-f491-4264-8368-7fc3c5a6992b" width="100%" controls autoplay loop></video>115     </td>116  </tr>117</table>118 119# Quick Start120### 1. Cloud usage: AliyunDSW/Docker121#### a. From AliyunDSW122DSW has free GPU time, which can be applied once by a user and is valid for 3 months after applying.123 124Aliyun provide free GPU time in [Freetier](https://free.aliyun.com/?product=9602825&crowd=enterprise&spm=5176.28055625.J_5831864660.1.e939154aRgha4e&scm=20140722.M_9974135.P_110.MO_1806-ID_9974135-MID_9974135-CID_30683-ST_8512-V_1), get it and use in Aliyun PAI-DSW to start CogVideoX-Fun within 5min!125 126[![DSW Notebook](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/dsw.png)](https://gallery.pai-ml.com/#/preview/deepLearning/cv/cogvideox_fun)127 128#### b. From ComfyUI129Our ComfyUI is as follows, please refer to [ComfyUI README](comfyui/README.md) for details.130![workflow graph](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/cogvideoxfunv1_workflow_i2v.jpg)131 132#### c. From docker133If you are using docker, please make sure that the graphics card driver and CUDA environment have been installed correctly in your machine.134 135Then execute the following commands in this way:136 137```138# pull image139docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun140 141# enter image142docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun143 144# clone code145git clone https://github.com/aigc-apps/CogVideoX-Fun.git146 147# enter CogVideoX-Fun's dir148cd CogVideoX-Fun149 150# download weights151mkdir models/Diffusion_Transformer152mkdir models/Personalized_Model153 154# Please use the hugginface link or modelscope link to download the model.155# CogVideoX-Fun156# https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP157# https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP158 159# Wan160# https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-InP161# https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-InP162```163 164### 2. Local install: Environment Check/Downloading/Installation165#### a. Environment Check166We have verified this repo execution on the following environment:167 168The detailed of Windows:169- OS: Windows 10170- python: python3.10 & python3.11171- pytorch: torch2.2.0172- CUDA: 11.8 & 12.1173- CUDNN: 8+174- GPU๏ผš Nvidia-3060 12G & Nvidia-3090 24G175 176The detailed of Linux:177- OS: Ubuntu 20.04, CentOS178- python: python3.10 & python3.11179- pytorch: torch2.2.0180- CUDA: 11.8 & 12.1181- CUDNN: 8+182- GPU๏ผšNvidia-V100 16G & Nvidia-A10 24G & Nvidia-A100 40G & Nvidia-A100 80G183 184We need about 60GB available on disk (for saving weights), please check!185 186#### b. Weights187We'd better place the [weights](#model-zoo) along the specified path:188 189```190๐Ÿ“ฆ models/191โ”œโ”€โ”€ ๐Ÿ“‚ Diffusion_Transformer/192โ”‚   โ”œโ”€โ”€ ๐Ÿ“‚ CogVideoX-Fun-V1.1-2b-InP/193โ”‚   โ”œโ”€โ”€ ๐Ÿ“‚ CogVideoX-Fun-V1.1-5b-InP/194โ”‚   โ”œโ”€โ”€ ๐Ÿ“‚ Wan2.1-Fun-14B-InP195โ”‚   โ””โ”€โ”€ ๐Ÿ“‚ Wan2.1-Fun-1.3B-InP/196โ”œโ”€โ”€ ๐Ÿ“‚ Personalized_Model/197โ”‚   โ””โ”€โ”€ your trained trainformer model / your trained lora model (for UI load)198```199 200# How to Use201 202<h3 id="video-gen">1. Generation</h3>203 204#### a. GPU Memory Optimization205Since Wan2.1 has a very large number of parameters, we need to consider memory optimization strategies to adapt to consumer-grade GPUs. We provide `GPU_memory_mode` for each prediction file, allowing you to choose between `model_cpu_offload`, `model_cpu_offload_and_qfloat8`, and `sequential_cpu_offload`. This solution is also applicable to CogVideoX-Fun generation.206 207- `model_cpu_offload`: The entire model is moved to the CPU after use, saving some GPU memory.208- `model_cpu_offload_and_qfloat8`: The entire model is moved to the CPU after use, and the transformer model is quantized to float8, saving more GPU memory.209- `sequential_cpu_offload`: Each layer of the model is moved to the CPU after use. It is slower but saves a significant amount of GPU memory.210 211`qfloat8` may slightly reduce model performance but saves more GPU memory. If you have sufficient GPU memory, it is recommended to use `model_cpu_offload`.212 213#### b. Using ComfyUI214For details, refer to [ComfyUI README](comfyui/README.md).215 216#### c. Running Python Files217- **Step 1**: Download the corresponding [weights](#model-zoo) and place them in the `models` folder.218- **Step 2**: Use different files for prediction based on the weights and prediction goals. This library currently supports CogVideoX-Fun, Wan2.1, and Wan2.1-Fun. Different models are distinguished by folder names under the `examples` folder, and their supported features vary. Use them accordingly. Below is an example using CogVideoX-Fun:219  - **Text-to-Video**:220    - Modify `prompt`, `neg_prompt`, `guidance_scale`, and `seed` in the file `examples/cogvideox_fun/predict_t2v.py`.221    - Run the file `examples/cogvideox_fun/predict_t2v.py` and wait for the results. The generated videos will be saved in the folder `samples/cogvideox-fun-videos`.222  - **Image-to-Video**:223    - Modify `validation_image_start`, `validation_image_end`, `prompt`, `neg_prompt`, `guidance_scale`, and `seed` in the file `examples/cogvideox_fun/predict_i2v.py`.224    - `validation_image_start` is the starting image of the video, and `validation_image_end` is the ending image of the video.225    - Run the file `examples/cogvideox_fun/predict_i2v.py` and wait for the results. The generated videos will be saved in the folder `samples/cogvideox-fun-videos_i2v`.226  - **Video-to-Video**:227    - Modify `validation_video`, `validation_image_end`, `prompt`, `neg_prompt`, `guidance_scale`, and `seed` in the file `examples/cogvideox_fun/predict_v2v.py`.228    - `validation_video` is the reference video for video-to-video generation. You can use the following demo video: [Demo Video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/play_guitar.mp4).229    - Run the file `examples/cogvideox_fun/predict_v2v.py` and wait for the results. The generated videos will be saved in the folder `samples/cogvideox-fun-videos_v2v`.230  - **Controlled Video Generation (Canny, Pose, Depth, etc.)**:231    - Modify `control_video`, `validation_image_end`, `prompt`, `neg_prompt`, `guidance_scale`, and `seed` in the file `examples/cogvideox_fun/predict_v2v_control.py`.232    - `control_video` is the control video extracted using operators such as Canny, Pose, or Depth. You can use the following demo video: [Demo Video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1.1/pose.mp4).233    - Run the file `examples/cogvideox_fun/predict_v2v_control.py` and wait for the results. The generated videos will be saved in the folder `samples/cogvideox-fun-videos_v2v_control`.234- **Step 3**: If you want to integrate other backbones or Loras trained by yourself, modify `lora_path` and relevant paths in `examples/{model_name}/predict_t2v.py` or `examples/{model_name}/predict_i2v.py` as needed.235 236#### d. Using the Web UI237The web UI supports text-to-video, image-to-video, video-to-video, and controlled video generation (Canny, Pose, Depth, etc.). This library currently supports CogVideoX-Fun, Wan2.1, and Wan2.1-Fun. Different models are distinguished by folder names under the `examples` folder, and their supported features vary. Use them accordingly. Below is an example using CogVideoX-Fun:238 239- **Step 1**: Download the corresponding [weights](#model-zoo) and place them in the `models` folder.240- **Step 2**: Run the file `examples/cogvideox_fun/app.py` to access the Gradio interface.241- **Step 3**: Select the generation model on the page, fill in `prompt`, `neg_prompt`, `guidance_scale`, and `seed`, click "Generate," and wait for the results. The generated videos will be saved in the `sample` folder.242 243# Reference244- CogVideo: https://github.com/THUDM/CogVideo/245- EasyAnimate: https://github.com/aigc-apps/EasyAnimate246- Wan2.1: https://github.com/Wan-Video/Wan2.1/247 248# License249This project is licensed under the [Apache License (Version 2.0)](https://github.com/modelscope/modelscope/blob/master/LICENSE).