sywang/TPIPS-EarlyFusion-Qwen3VL-8B
TPIPS — Early Fusion (R1/R2 register tokens)
Text-conditioned perceptual image similarity, built on Qwen/Qwen3-VL-Embedding-8B. This repo holds the `early_fusion` checkpoint (one of three TPIPS models, each in its own repo — see the table at the bottom). Code and full docs: https://github.com/adobe-research/TPIPS.
Early fusion. (img0, img1, text) are encoded jointly with two register tokens R1 / R2 under a block-sparse attention mask. The pairwise score is cos(R1, R2) (symmetrized over both orderings) (higher = more similar). Odd-one-out probabilities are a softmax over the three "other-pair" scores divided by the temperature; 2AFC compares the two reference-candidate scores. The pairwise score is the model's raw output (temperature is applied at the probability step).
Usage
TPIPS supports Python 3.10 and later. Install matching PyTorch and torchvision builds from the official PyTorch installer, then install TPIPS:
pip install tpipsimport tpips
from PIL import Image
model = tpips.load_model("early_fusion", device="cuda")
a = Image.open("a.jpg").convert("RGB")
b = Image.open("b.jpg").convert("RGB")
similarity = model.similarity(a, b, factor="lighting") # higher is more similar
distance = model.distance(a, b, factor="lighting") # lower is more similarThe first call downloads the selected TPIPS checkpoint and its Qwen backbone. A CUDA GPU is recommended; FlashAttention is optional.
The TPIPS models
License
TPIPS is provided under the Adobe Research License for noncommercial research use. See the license for the complete terms.
