tencent/WeMM-Embedding-2B
WeMM-Embedding-2B
  
WeMM-Embedding-2B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 2,048-dimensional L2-normalized embedding. Audio input is not supported.
Installation
pip install torch transformers==5.2.0 "qwen-vl-utils[decord]==0.0.14" \
"sentence-transformers>=5.7.0" "accelerate>=1.1.0"Transformers
import torch
from qwen_vl_utils import process_vision_info
from transformers import AutoModel, AutoProcessor
model_id = "tencent/WeMM-Embedding-2B"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModel.from_pretrained(
model_id, trust_remote_code=True, dtype=torch.bfloat16
).cuda().eval()
messages = [{"role": "user", "content": [
{"type": "image", "image": "/path/to/image.jpg"},
{"type": "video", "video": "/path/to/video.mp4"},
{"type": "text", "text": "This can be any text input."},
]}]
text = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=False
)
images, videos, video_kwargs = process_vision_info(
messages,
image_patch_size=16,
return_video_kwargs=True,
return_video_metadata=True,
)
if videos is not None:
videos, video_metadata = zip(*videos)
videos, video_metadata = list(videos), list(video_metadata)
else:
video_metadata = None
inputs = processor(
text=text,
images=images,
videos=videos,
video_metadata=video_metadata,
return_tensors="pt",
**video_kwargs,
).to("cuda")
with torch.inference_mode():
embedding = model.embedding(**inputs)Use any subset of the content items to encode text, image, or video independently.
Sentence Transformers
from sentence_transformers import SentenceTransformer
model_id = "tencent/WeMM-Embedding-2B"
model = SentenceTransformer(model_id, trust_remote_code=True)
queries = [
"Which Llama 4 model variants are available?",
"How is mapo tofu prepared?",
]
documents = [
"Mapo tofu is a Sichuan dish of soft tofu simmered in a spicy, numbing sauce of chili bean paste and Sichuan peppercorn.",
{
"image": "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/llama4_hgf.png",
},
{
"video": "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/mapo_tofu.mp4",
},
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# (2, 2048) (3, 2048)
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[0.0685, 0.5379, 0.0423],
# [0.7829, 0.1852, 0.4729]])Each input is a string, a URL or path, a PIL.Image, or a dict combining image, video, and text keys. Chat messages such as {"role": "user", "content": [{"type": "image", "image": ...}, {"type": "text", "text": ...}]} are also accepted, which is the way to interleave several images or videos in one input.
Matryoshka Embeddings
embedding_256 = torch.nn.functional.normalize(embedding[..., :256], dim=-1)With Sentence Transformers, pass truncate_dim and let it renormalize:
embeddings_256 = model.encode_document(documents, truncate_dim=256, normalize_embeddings=True)Use a dimension listed in model.config.matryoshka_dimensions. On MMEB-v2, 256-dimensional embeddings retain 98.7% of the full-dimensional image and video performance.
Serving
vLLM 0.27.0:
MODEL_PATH=/path/to/WeMM-Embedding-2B
vllm serve "$MODEL_PATH" \
--runner pooling \
--chat-template "$MODEL_PATH/embedding_chat_template.jinja"SGLang 0.5.9:
MODEL_PATH=/path/to/WeMM-Embedding-2B
python patch_sglang_video.py
python -m sglang.launch_server \
--model-path "$MODEL_PATH" \
--is-embedding \
--enable-precise-embedding-interpolationEvaluation
MMEB-v2
Results on 78 datasets from Table 1 of the technical report. Image and video tasks use Hit@1, while visual-document tasks use NDCG@5. Higher is better.
† Closed-source leaderboard submission without publicly released model weights or a public inference endpoint.
MMEB-v3
Results on all 190 tasks from Table 2 of the technical report. V3-All includes the 78 MMEB-v2 tasks, 53 text tasks, 47 agent tasks, 11 audio tasks, and MCMR. Unsupported tasks are assigned a score of zero.
Text results use NDCG@5; agent, MCMR, and audio results use Hit@1.
Citation
If you find this repository useful, please consider giving a star ⭐ and citation
@article{wemm-embedding,
title={WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report},
author={Junjie Zhou and Ke Mei and Lei Li and Tianyi Wang and Fengyun Rao and Jing Lyu},
year={2026},
eprint={2608.24053},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.24053},
}License
WeMM-Embedding-2B, including the code, model parameters, and weights made publicly available by Tencent, is licensed under the Apache License 2.0. Third-party components remain subject to their respective original licenses.
