alibaba-pai/Z-Image-Fun-Lora-Distill
Z-Image-Fun-Lora-Distill

Model Card
a. 2603 Models
b. 2602 Models && Models Before 2602
Model Features
- This is a Distill LoRA for Z-Image that distills both steps and CFG. It does not use any Z-Image-Turbo related weights and is trained from scratch. It is compatible with other Z-Image LoRAs and Controls.
- This model will slightly reduce the output quality and change the output composition of the model. For specific comparisons, please refer to the Results section.
- The purpose of this model is to provide fast generation compatibility for Z-Image derivative models, not to replace Z-Image-Turbo.
Results
The difference between the 2603 version model and the 2602 version model
The 2602 model tends to produce blurry images with sigmas below 0.500, as the distillation model was not trained on certain steps. The 2603 model introduces a random timesteps strategy, making it better adapted to sigmas below 0.500.
As shown below, when using kloptimal, many sigmas fall below 0.500. The 2603 model handles these cases correctly, while the 2602 model does not. Note that although kloptimal is used in the figure, we still recommend using the simple scheduler for inference.
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr> <td>Z-Image-Fun-Lora-Distill-8-Steps-2602</td> <td>Z-Image-Fun-Lora-Distill-8-Steps-2603</td> </tr> <tr> <td><img src="results/2602.png" width="100%" /></td> <td><img src="results/2603.png" width="100%" /></td> </tr> </table>
The difference between the 2602 version model and the previous model
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr> <td>Z-Image-Fun-Lora-Distill-8-Steps-2602</td> <td>Z-Image-Fun-Lora-Distill-4-Steps-2602</td> <td>Z-Image-Fun-Lora-Distill-8-Steps</td> </tr> <tr> <td><img src="results/260218steps.png" width="100%" /><img src="results/260228steps.png" width="100%" /><img src="results/260238steps.png" width="100%" /><img src="results/260248steps.png" width="100%" /><img src="results/260258steps.png" width="100%" /></td> <td><img src="results/260214steps.png" width="100%" /><img src="results/260224steps.png" width="100%" /><img src="results/260234steps.png" width="100%" /><img src="results/260244steps.png" width="100%" /><img src="results/260254steps.png" width="100%" /></td> <td><img src="results/old18steps.png" width="100%" /><img src="results/old28steps.png" width="100%" /><img src="results/old38steps.png" width="100%" /><img src="results/old48steps.png" width="100%" /><img src="results/old58steps.png" width="100%" /></td> </tr> </table>
Work itself
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr> <td>Output 25 steps</td> <td>Output 8-Steps-2602</td> <td>Output 4-Steps-2602</td> </tr> <tr> <td><img src="results/output4.png" width="100%" /></td> <td><img src="results/output426028steps.png" width="100%" /></td> <td><img src="results/output426024steps.png" width="100%" /></td> </tr> </table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr> <td>Output 25 steps</td> <td>Output 8-Steps-2602</td> <td>Output 4-Steps-2602</td> </tr> <tr> <td><img src="results/output1.png" width="100%" /></td> <td><img src="results/output126028steps.png" width="100%" /></td> <td><img src="results/output126024steps.png" width="100%" /></td> </tr> </table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr> <td>Output 25 steps</td> <td>Output 8-Steps-2602</td> <td>Output 4-Steps-2602</td> </tr> <tr> <td><img src="results/output2.png" width="100%" /></td> <td><img src="results/output226028steps.png" width="100%" /></td> <td><img src="results/output226024steps.png" width="100%" /></td> </tr> </table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr> <td>Output 25 steps</td> <td>Output 8-Steps-2602</td> <td>Output 4-Steps-2602</td> </tr> <tr> <td><img src="results/output3.png" width="100%" /></td> <td><img src="results/output326028steps.png" width="100%" /></td> <td><img src="results/output326024steps.png" width="100%" /></td> </tr> </table>
Work with Controlnet
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr> <td>Pose + Inpaint</td> <td>Output 25 steps</td> <td>Output 8-Steps-2602</td> <td>Output 4-Steps-2602</td> </tr> <tr> <td><img src="asset/inpaint.jpg" width="100%" /><img src="asset/mask.jpg" width="100%" /></td> <td><img src="results/inpaint.png" width="100%" /></td> <td><img src="results/inpaint26028steps.png" width="100%" /></td> <td><img src="results/inpaint26024steps.png" width="100%" /></td> </tr> </table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr> <td>Pose + Inpaint</td> <td>Output 25 steps</td> <td>Output 8-Steps-2602</td> <td>Output 4-Steps-2602</td> </tr> <tr> <td><img src="asset/inpaint.jpg" width="100%" /><img src="asset/mask.jpg" width="100%" /><img src="asset/pose.jpg" width="100%" /></td> <td><img src="results/poseinpaint.png" width="100%" /></td> <td><img src="results/poseinpaint26028steps.png" width="100%" /></td> <td><img src="results/poseinpaint2602_4steps.png" width="100%" /></td> </tr> </table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr> <td>Pose</td> <td>Output 25 steps</td> <td>Output 8-Steps-2602</td> <td>Output 4-Steps-2602</td> </tr> <tr> <td><img src="asset/pose2.jpg" width="100%" /></td> <td><img src="results/pose2.png" width="100%" /></td> <td><img src="results/pose226028steps.png" width="100%" /></td> <td><img src="results/pose226024steps.png" width="100%" /></td> </tr> </table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr> <td>Canny</td> <td>Output</td> <td>Output 8-Steps-2602</td> <td>Output 4-Steps-2602</td> </tr> <tr> <td><img src="asset/canny.jpg" width="100%" /></td> <td><img src="results/canny.png" width="100%" /></td> <td><img src="results/canny26028steps.png" width="100%" /></td> <td><img src="results/canny26024steps.png" width="100%" /></td> </tr> </table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr> <td>Depth</td> <td>Output</td> <td>Output 8-Steps-2602</td> <td>Output 4-Steps-2602</td> </tr> <tr> <td><img src="asset/gray.jpg" width="100%" /></td> <td><img src="results/gray.png" width="100%" /></td> <td><img src="results/gray26028steps.png" width="100%" /></td> <td><img src="results/gray26024steps.png" width="100%" /></td> </tr> </table>
Inference
Go to the VideoX-Fun repository for more details.
Please clone the VideoX-Fun repository and create the required directories:
# Clone the code
git clone https://github.com/aigc-apps/VideoX-Fun.git
# Enter VideoX-Fun's directory
cd VideoX-Fun
# Create model directories
mkdir -p models/Diffusion_Transformer
mkdir -p models/Personalized_ModelThen download the weights into models/DiffusionTransformer and models/PersonalizedModel.
๐ฆ models/
โโโ ๐ Diffusion_Transformer/
โ โโโ ๐ Z-Image/
โโโ ๐ Personalized_Model/
โ โโโ ๐ฆ Z-Image-Fun-Lora-Distill-4-Steps-2602.safetensors
โ โโโ ๐ฆ Z-Image-Fun-Lora-Distill-8-Steps-2602.safetensors
โ โโโ ๐ฆ Z-Image-Fun-Controlnet-Union-2.1.safetensors
โ โโโ ๐ฆ Z-Image-Fun-Controlnet-Union-2.1-lite.safetensorsTo run the model, first set the lorapath in `examples/zimage/predictt2i.py` to: `PersonalizedModel/Z-Image-Fun-Lora-Distill-8-Steps.safetensors`
Then, run the file: examples/z_image/predict_t2i.py
The following scripts are also supported:
- examples/zimagefun/predictt2icontrol_2.1.py
- examples/zimagefun/predicti2iinpaint_2.1.py
Recommended Settings:
- cfg = 1.0
- steps = 8
- lora_weight = 0.8 (suggested range: 0.7 ~ 0.9)
