CoolFace
Modelpublic

Alibaba-DAMO-Academy/PixelRefer-2B

sourceHugging Faceapache-2.0updated 11mo agoView on Hugging Face
4likes27downloads
Model Card

<p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/64a3fe3dde901eb01df12398/CiwESNwyy7VTooOifRgiQ.png" width="70%" style="margin-bottom: 0.2;"/> <p>

<h3 align="center"><a href="http://arxiv.org/abs/2510.23603" style="color:#4D2B24"> PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity</a></h3>

<div align=center>

Static Badge ![arXiv preprint](https://arxiv.org/abs/2510.23603) ![Dataset](https://huggingface.co/datasets/DAMO-NLP-SG/VideoRefer-700K) ![Model](https://huggingface.co/collections/Alibaba-DAMO-Academy/pixelrefer) ![Benchmark](https://huggingface.co/datasets/DAMO-NLP-SG/VideoRefer-Bench)

![Homepage](https://circleradon.github.io/PixelRefer/) ![Huggingface](https://huggingface.co/spaces/lixin4ever/PixelRefer) </div>

📰 News

🌏 Model Zoo

📑 Citation

If you find VideoRefer Suite useful for your research and applications, please cite using this BibTeX:

bibtex
@article{yuan2025pixelrefer,
  title     = {PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity},
  author    = {Yuqian Yuan and Wenqiao Zhang and Xin Li and Shihao Wang and Kehan Li and Wentong Li and Jun Xiao and Lei Zhang and Beng Chin Ooi},
  year      = {2025},
  journal   = {arXiv},
}

@inproceedings{yuan2025videorefer,
  title     = {Videorefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM},
  author    = {Yuqian Yuan and Hang Zhang and Wentong Li and Zesen Cheng and Boqiang Zhang and Long Li and Xin Li and Deli Zhao and Wenqiao Zhang and Yueting Zhuang and others},
  booktitle = {Proceedings of the Computer Vision and Pattern Recognition Conference},
  pages     = {18970--18980},
  year      = {2025},
}