CoolFace
Modelpublic

prithivMLmods/PhotoCleanser-i2i

sourceHugging Faceotherupdated 1y agoView on Hugging Face
9likes750downloads
Model Card

1.png

PhotoCleanser-i2i [Image-to-Image] (experimental)

<Gallery />

PhotoCleanser-i2i is an adapter for black-forest-lab's FLUX.1-Kontext-dev. It is an experimental LoRA designed for removing specified object(s) while preserving the remaining content in the image. The model was trained on 36 image pairs (18 start images, 18 end images). Synthetic result nodes were generated using NanoBanana from Google and labeled with DeepCaption-VLA-7B. The adapter is triggered with the following prompt:

[!note]

[photo content], remove the specified object(s) from the image while preserving the background and remaining elements, maintaining realism and original details.

Some use cases

[photo content], remove all humans from the image while preserving the background and remaining elements, maintaining realism and original details.
[photo content], remove the ball from the image while preserving the background and remaining elements, maintaining realism and original details.
[photo content], remove the cat from the image while preserving the background and remaining elements, maintaining realism and original details.

Sample Inference Comparing the Base Model with the Adapter

Note: In over 25 inferences across various scale settings, the base model struggles to properly reconstruct the basketball net after the ball is removed.
FLUX.1-Kontext-devPhotoCleanser-i2i
FLUX.1 KontextPhotoCleanser-i2i
No desired effect was achieved with the base model after testing many settings, but PhotoCleanser-i2i performed its best to remove the human character from the image.
FLUX.1-Kontext-devPhotoCleanser-i2i
FLUX.1 KontextPhotoCleanser-i2i
No desired effect was achieved with the base model after testing many settings, but PhotoCleanser-i2i performed its best to remove the human/animal character from the image.
FLUX.1-Kontext-devPhotoCleanser-i2i
FLUX.1 KontextPhotoCleanser-i2i

Parameter Settings

SettingValue
Module TypeAdapter
Base ModelFLUX.1 Kontext Dev - fp8
Trigger Words[photo content], remove the specified object(s) from the image while preserving the background and remaining elements, maintaining realism and original details.
Image Processing Repeats50
Epochs32
Save Every N Epochs1

Labeling: DeepCaption-VLA-7B(natural language & English)

Total Images Used for Training : 36 Image Pairs (18 Start, 18 End)

Synthetic Result Node generated by NanoBanana from Google (Image Result Sets Dataset)

Training Parameters

SettingValue
Seed-
Clip Skip-
Text Encoder LR0.00001
UNet LR0.00005
LR Schedulerconstant
OptimizerAdamW8bit
Network Dimension64
Network Alpha32
Gradient Accumulation Steps-

Label Parameters

SettingValue
Shuffle Caption-
Keep N Tokens-

Advanced Parameters

SettingValue
Noise Offset0.03
Multires Noise Discount0.1
Multires Noise Iterations10
Conv Dimension-
Conv Alpha-
Batch Size-
Steps4370 (Low(700))
Samplereuler

Trigger words

You should use [photo content] to trigger the image generation.

You should use remove the specified object(s) from the image while preserving the background and remaining elements to trigger the image generation.

You should use maintaining realism and original details. to trigger the image generation.

Download model

Download them in the Files & versions tab.