CoolFace
Modelpublic

zhengchong/CatVTON

sourceHugging Facecc-by-nc-sa-4.0updated 2y agoView on Hugging Face
86likes5.5kdownloads
Model Card

๐Ÿˆ CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models

<div style="display: flex; justify-content: center; align-items: center;"> <a href="http://arxiv.org/abs/2407.15886" style="margin: 0 2px;"> <img src='https://img.shields.io/badge/arXiv-2407.15886-red?style=flat&logo=arXiv&logoColor=red' alt='arxiv'> </a> <a href='https://huggingface.co/zhengchong/CatVTON' style="margin: 0 2px;"> <img src='https://img.shields.io/badge/Hugging Face-ckpts-orange?style=flat&logo=HuggingFace&logoColor=orange' alt='huggingface'> </a> <a href="https://github.com/Zheng-Chong/CatVTON" style="margin: 0 2px;"> <img src='https://img.shields.io/badge/GitHub-Repo-blue?style=flat&logo=GitHub' alt='GitHub'> </a> <a href="http://120.76.142.206:8888" style="margin: 0 2px;"> <img src='https://img.shields.io/badge/Demo-Gradio-gold?style=flat&logo=Gradio&logoColor=red' alt='Demo'> </a> <a href="https://huggingface.co/spaces/zhengchong/CatVTON" style="margin: 0 2px;"> <img src='https://img.shields.io/badge/Space-ZeroGPU-orange?style=flat&logo=Gradio&logoColor=red' alt='Demo'> </a> <a href='https://zheng-chong.github.io/CatVTON/' style="margin: 0 2px;"> <img src='https://img.shields.io/badge/Webpage-Project-silver?style=flat&logo=&logoColor=orange' alt='webpage'> </a> <a href="https://github.com/Zheng-Chong/CatVTON/LICENCE" style="margin: 0 2px;"> <img src='https://img.shields.io/badge/License-CC BY--NC--SA--4.0-lightgreen?style=flat&logo=Lisence' alt='License'> </a> </div>

CatVTON is a simple and efficient virtual try-on diffusion model with *1) Lightweight Network (899.06M parameters totally), 2) Parameter-Efficient Training (49.57M parameters trainable) and 3) Simplified Inference (< 8G VRAM for 1024X768 resolution)*.

Updates

  • โ€”`2024/10/17`:**Mask-free version**๐Ÿค— of CatVTON is release and please try it in our **Online Demo**.
  • โ€”`2024/10/13`: We have built a repo **Awesome-Try-On-Models** that focuses on image, video, and 3D-based try-on models published after 2023, aiming to provide insights into the latest technological trends. If you're interested, feel free to contribute or give it a ๐ŸŒŸ star!
  • โ€”`2024/08/13`: We localize DensePose & SCHP to avoid certain environment issues.
  • โ€”`2024/08/10`: Our ๐Ÿค— **HuggingFace Space** is available now! Thanks for the grant from **ZeroGPU**๏ผ
  • โ€”`2024/08/09`: **Evaluation code** is provided to calculate metrics ๐Ÿ“š.
  • โ€”`2024/07/27`: We provide code and workflow for deploying CatVTON on **ComfyUI** ๐Ÿ’ฅ.
  • โ€”`2024/07/24`: Our **Paper on ArXiv** is available ๐Ÿฅณ!
  • โ€”`2024/07/22`: Our **App Code** is released, deploy and enjoy CatVTON on your mechine ๐ŸŽ‰!
  • โ€”`2024/07/21`: Our **Inference Code** and **Weights** ๐Ÿค— are released.
  • โ€”`2024/07/11`: Our **Online Demo** is released ๐Ÿ˜.

Installation

Create a conda environment & Install requirments

shell
conda create -n catvton python==3.9.0
conda activate catvton
cd CatVTON-main  # or your path to CatVTON project dir
pip install -r requirements.txt

Deployment

ComfyUI Workflow

We have modified the main code to enable easy deployment of CatVTON on ComfyUI. Due to the incompatibility of the code structure, we have released this part in the Releases, which includes the code placed under custom_nodes of ComfyUI and our workflow JSON files.

To deploy CatVTON to your ComfyUI, follow these steps:

  1. 1.Install all the requirements for both CatVTON and ComfyUI, refer to Installation Guide for CatVTON and Installation Guide for ComfyUI.
  2. 2.Download `ComfyUI-CatVTON.zip` and unzip it in the custom_nodes folder under your ComfyUI project (clone from ComfyUI).
  3. 3.Run the ComfyUI.
  4. 4.Download `catvton_workflow.json` and drag it into you ComfyUI webpage and enjoy ๐Ÿ˜†!
Problems under Windows OS, please refer to issue#8.

When you run the CatVTON workflow for the first time, the weight files will be automatically downloaded, usually taking dozens of minutes.

<div align="center"> <img src="resource/img/comfyui-1.png" width="100%" height="100%"/> </div>

<!-- <div align="center"> <img src="resource/img/comfyui.png" width="100%" height="100%"/> </div> -->

Gradio App

To deploy the Gradio App for CatVTON on your machine, run the following command, and checkpoints will be automatically downloaded from HuggingFace.

PowerShell
CUDA_VISIBLE_DEVICES=0 python app.py \
--output_dir="resource/demo/output" \
--mixed_precision="bf16" \
--allow_tf32 

When using bf16 precision, generating results with a resolution of 1024x768 only requires about 8G VRAM.

Inference

1. Data Preparation

Before inference, you need to download the VITON-HD or DressCode dataset. Once the datasets are downloaded, the folder structures should look like these:

โ”œโ”€โ”€ VITON-HD
|   โ”œโ”€โ”€ test_pairs_unpaired.txt
โ”‚   โ”œโ”€โ”€ test
|   |   โ”œโ”€โ”€ image
โ”‚   โ”‚   โ”‚   โ”œโ”€โ”€ [000006_00.jpg | 000008_00.jpg | ...]
โ”‚   โ”‚   โ”œโ”€โ”€ cloth
โ”‚   โ”‚   โ”‚   โ”œโ”€โ”€ [000006_00.jpg | 000008_00.jpg | ...]
โ”‚   โ”‚   โ”œโ”€โ”€ agnostic-mask
โ”‚   โ”‚   โ”‚   โ”œโ”€โ”€ [000006_00_mask.png | 000008_00.png | ...]
...
โ”œโ”€โ”€ DressCode
|   โ”œโ”€โ”€ test_pairs_paired.txt
|   โ”œโ”€โ”€ test_pairs_unpaired.txt
โ”‚   โ”œโ”€โ”€ [dresses | lower_body | upper_body]
|   |   โ”œโ”€โ”€ test_pairs_paired.txt
|   |   โ”œโ”€โ”€ test_pairs_unpaired.txt
โ”‚   โ”‚   โ”œโ”€โ”€ images
โ”‚   โ”‚   โ”‚   โ”œโ”€โ”€ [013563_0.jpg | 013563_1.jpg | 013564_0.jpg | 013564_1.jpg | ...]
โ”‚   โ”‚   โ”œโ”€โ”€ agnostic_masks
โ”‚   โ”‚   โ”‚   โ”œโ”€โ”€ [013563_0.png| 013564_0.png | ...]
...

For the DressCode dataset, we provide script to preprocessed agnostic masks, run the following command:

PowerShell
CUDA_VISIBLE_DEVICES=0 python preprocess_agnostic_mask.py \
--data_root_path <your_path_to_DressCode> 

2. Inference on VTIONHD/DressCode

To run the inference on the DressCode or VITON-HD dataset, run the following command, checkpoints will be automatically downloaded from HuggingFace.

PowerShell
CUDA_VISIBLE_DEVICES=0 python inference.py \
--dataset [dresscode | vitonhd] \
--data_root_path <path> \
--output_dir <path> 
--dataloader_num_workers 8 \
--batch_size 8 \
--seed 555 \
--mixed_precision [no | fp16 | bf16] \
--allow_tf32 \
--repaint \
--eval_pair  

3. Calculate Metrics

After obtaining the inference results, calculate the metrics using the following command:

PowerShell
CUDA_VISIBLE_DEVICES=0 python eval.py \
--gt_folder <your_path_to_gt_image_folder> \
--pred_folder <your_path_to_predicted_image_folder> \
--paired \
--batch_size=16 \
--num_workers=16 
  • โ€”--gt_folder and --pred_folder should be folders that contain only images.
  • โ€”To evaluate the results in a paired setting, use --paired; for an unpaired setting, simply omit it.
  • โ€”--batch_size and --num_workers should be adjusted based on your machine.

Acknowledgement

Our code is modified based on Diffusers. We adopt Stable Diffusion v1.5 inpainting as the base model. We use SCHP and DensePose to automatically generate masks in our Gradio App and ComfyUI workflow. Thanks to all the contributors!

License

All the materials, including code, checkpoints, and demo, are made available under the Creative Commons BY-NC-SA 4.0 license. You are free to copy, redistribute, remix, transform, and build upon the project for non-commercial purposes, as long as you give appropriate credit and distribute your contributions under the same license.

Citation

bibtex
@misc{chong2024catvtonconcatenationneedvirtual,
 title={CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models}, 
 author={Zheng Chong and Xiao Dong and Haoxiang Li and Shiyue Zhang and Wenqing Zhang and Xujie Zhang and Hanqing Zhao and Xiaodan Liang},
 year={2024},
 eprint={2407.15886},
 archivePrefix={arXiv},
 primaryClass={cs.CV},
 url={https://arxiv.org/abs/2407.15886}, 
}