CoolFace
Modelpublic

rp-yu/Qwen2-VL-2b-VPT-Seg

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
2likes6downloads
Model Card

Introducing Visual Perception Token into Multimodal Large Language Model

This repository contains models based on the paper Introducing Visual Perception Token into Multimodal Large Language Model. These models utilize Visual Perception Tokens to enhance the visual perception capabilities of multimodal large language models (MLLMs).

Code: https://github.com/yu-rp/VisualPerceptionToken