CoolFace
Modelpublic

VisualCloze/VisualClozePipeline-LoRA-384

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
4likes18downloads
Model Card

VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning (LoRA weight for with <strong><span style="color:red">Diffusers</span></strong>)

<div align="center">

[Paper] &emsp; [Project Page] &emsp; [Github]

</div>

<div align="center">

[๐Ÿค— <strong><span style="color:hotpink">Diffusers</span></strong> Implementation]

</div>

<div align="center">

[๐Ÿค— Online Demo] &emsp; [๐Ÿค— Full Model Card] &emsp; [๐Ÿค— Dataset Card]

</div>

Examples

If you find VisualCloze is helpful, please consider to star โญ the <strong><span style="color:hotpink">Github Repo</span></strong>. Thanks!

๐Ÿ“ฐ News

๐ŸŒ  Key Features

An in-context learning based universal image generation framework.

  1. 1.Support various in-domain tasks.
  2. 2.Generalize to <strong><span style="color:hotpink"> unseen tasks</span></strong> through in-context learning.
  3. 3.Unify multiple tasks into one step and generate both target image and intermediate results.
  4. 4.Support reverse-engineering a set of conditions from a target image.

๐Ÿ”ฅ Examples are shown in the project page.

๐Ÿ”ง Installation

<strong><span style="color:hotpink">You can install the official </span></strong> diffusers.

bash
pip install git+https://github.com/huggingface/diffusers.git

๐Ÿ’ป Diffusers Usage

![Huggingface VisualCloze](https://huggingface.co/spaces/VisualCloze/VisualCloze)

The weights in this model contains the LoRA weights supporting diffusers and Full Model Card provides the full weights.

While this model uses the resolution of 384, a model trained with the resolution of 512 is released at Full Model Card 512 and LoRA Model Card 512. The resolution means that each image will be resized to it before being concatenated to avoid the out-of-memory error. To generate high-resolution images, we use the SDEdit technology for upsampling the generated results.

Example with Depth-to-Image:

<img src="./visualclozediffusersexample_depthtoimage.jpg" width="60%" height="50%" alt="Example with Depth-to-Image"/>

python
import torch
from diffusers import VisualClozePipeline
from diffusers.utils import load_image


# Load in-context images (make sure the paths are correct and accessible)
image_paths = [
    # in-context examples
    [
        load_image('https://github.com/lzyhha/VisualCloze/raw/main/examples/examples/93bc1c43af2d6c91ac2fc966bf7725a2/93bc1c43af2d6c91ac2fc966bf7725a2_depth-anything-v2_Large.jpg'),
        load_image('https://github.com/lzyhha/VisualCloze/raw/main/examples/examples/93bc1c43af2d6c91ac2fc966bf7725a2/93bc1c43af2d6c91ac2fc966bf7725a2.jpg'),
    ],
    # query with the target image
    [
        load_image('https://github.com/lzyhha/VisualCloze/raw/main/examples/examples/79f2ee632f1be3ad64210a641c4e201b/79f2ee632f1be3ad64210a641c4e201b_depth-anything-v2_Large.jpg'),
        None,  # No image needed for the query in this case
    ],
]

# Task and content prompt
task_prompt = "Each row outlines a logical process, starting from [IMAGE1] gray-based depth map with detailed object contours, to achieve [IMAGE2] an image with flawless clarity."
content_prompt = """A serene portrait of a young woman with long dark hair, wearing a beige dress with intricate 
gold embroidery, standing in a softly lit room. She holds a large bouquet of pale pink roses in a black box, 
positioned in the center of the frame. The background features a tall green plant to the left and a framed artwork 
on the wall to the right. A window on the left allows natural light to gently illuminate the scene. 
The woman gazes down at the bouquet with a calm expression. Soft natural lighting, warm color palette, 
high contrast, photorealistic, intimate, elegant, visually balanced, serene atmosphere."""

# Load the VisualClozePipeline
pipe = VisualClozePipeline.from_pretrained("black-forest-labs/FLUX.1-Fill-dev", resolution=384, torch_dtype=torch.bfloat16)
pipe.load_lora_weights('VisualCloze/VisualClozePipeline-LoRA-384', weight_name='visualcloze-lora-384.safetensors')
pipe.to("cuda")

# Run the pipeline
image_result = pipe(
    task_prompt=task_prompt,
    content_prompt=content_prompt,
    image=image_paths,
    upsampling_width=1024,
    upsampling_height=1024,
    upsampling_strength=0.4,
    guidance_scale=30,
    num_inference_steps=30,
    max_sequence_length=512,
    generator=torch.Generator("cpu").manual_seed(0)
).images[0][0]

# Save the resulting image
image_result.save("visualcloze.png")
Example with Virtual Try-On:

<img src="./visualclozediffusersexample_tryon.jpg" width="60%" height="50%" alt="Example with Virtual Try-On"/>

python
import torch
from diffusers import VisualClozePipeline
from diffusers.utils import load_image


# Load in-context images (make sure the paths are correct and accessible)
# The images are from the VITON-HD dataset at https://github.com/shadow2496/VITON-HD
image_paths = [
    # in-context examples
    [
        load_image('https://github.com/lzyhha/VisualCloze/raw/main/examples/examples/tryon/00700_00.jpg'),
        load_image('https://github.com/lzyhha/VisualCloze/raw/main/examples/examples/tryon/03673_00.jpg'),
        load_image('https://github.com/lzyhha/VisualCloze/raw/main/examples/examples/tryon/00700_00_tryon_catvton_0.jpg'),
    ],
    # query with the target image
    [
        load_image('https://github.com/lzyhha/VisualCloze/raw/main/examples/examples/tryon/00555_00.jpg'),
        load_image('https://github.com/lzyhha/VisualCloze/raw/main/examples/examples/tryon/12265_00.jpg'),
        None
    ],
]

# Task and content prompt
task_prompt = "Each row shows a virtual try-on process that aims to put [IMAGE2] the clothing onto [IMAGE1] the person, producing [IMAGE3] the person wearing the new clothing."
content_prompt = None

# Load the VisualClozePipeline
pipe = VisualClozePipeline.from_pretrained("black-forest-labs/FLUX.1-Fill-dev", resolution=384, torch_dtype=torch.bfloat16)
pipe.load_lora_weights('VisualCloze/VisualClozePipeline-LoRA-384', weight_name='visualcloze-lora-384.safetensors')
pipe.to("cuda")

# Run the pipeline
image_result = pipe(
    task_prompt=task_prompt,
    content_prompt=content_prompt,
    image=image_paths,
    upsampling_height=1632,
    upsampling_width=1232,
    upsampling_strength=0.3,
    guidance_scale=30,
    num_inference_steps=30,
    max_sequence_length=512,
    generator=torch.Generator("cpu").manual_seed(0)
).images[0][0]

# Save the resulting image
image_result.save("visualcloze.png")

Citation

If you find VisualCloze useful for your research and applications, please cite using this BibTeX:

bibtex
@InProceedings{Li_2025_ICCV,
    author    = {Li, Zhong-Yu and Du, Ruoyi and Yan, Juncheng and Zhuo, Le and Li, Zhen and Gao, Peng and Ma, Zhanyu and Cheng, Ming-Ming},
    title     = {VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning},
    booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
    month     = {October},
    year      = {2025},
    pages     = {18969-18979}
}

Acknowledgement

The LoRA weights are converted with the help of NagaSaiAbhinay at the discussion.