CoolFace
Modelpublic

CodeGoat24/UnifiedReward-2.0-qwen3vl-32b

sourceHugging Facemitupdated 10mo agoView on Hugging Face
2likes1.2kdownloads
Model Card

Model Summary

UnifiedReward-2.0-qwen3vl-32b is the first unified reward model based on Qwen/Qwen3-VL-32B-Instruct for multimodal understanding and generation assessment, enabling both pairwise ranking and pointwise scoring, which can be employed for vision model preference alignment.

For further details, please refer to the following resources:

  • โ€”๐Ÿ“ฐ Paper: https://arxiv.org/pdf/2503.05236
  • โ€”๐Ÿช Project Page: https://codegoat24.github.io/UnifiedReward/
  • โ€”๐Ÿค— Model Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-models-67c3008148c3a380d15ac63a
  • โ€”๐Ÿค— Dataset Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-training-data-67c300d4fd5eff00fa7f1ede
  • โ€”๐Ÿ‘‹ Point of Contact: Yibin Wang

๐Ÿ Compared with Current Reward Models

Reward ModelMethodImage GenerationImage UnderstandingVideo GenerationVideo Understanding
PickScorePointโˆš
HPSPointโˆš
ImageRewardPointโˆš
LLaVA-CriticPair/Pointโˆš
IXC-2.5-RewardPair/Pointโˆšโˆš
VideoScorePointโˆš
LiFTPointโˆš
VisionRewardPointโˆšโˆš
VideoRewardPointโˆš
UnifiedReward (Ours)Pair/Pointโˆšโˆšโˆšโˆš

Citation

@article{unifiedreward,
  title={Unified reward model for multimodal understanding and generation},
  author={Wang, Yibin and Zang, Yuhang and Li, Hao and Jin, Cheng and Wang, Jiaqi},
  journal={arXiv preprint arXiv:2503.05236},
  year={2025}
}