CoolFace
Modelpublic

csfufu/Revisual-R1-final

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
8likes27downloads
Model Card

🌟 ReVisual-R1 (7B) β€” Open-Source Multimodal Reasoner

One cold-start, two RL stages, endless reasoning power.

πŸ”‘ Highlights

  • β€”SOTA on 9 tough benchmarks covering visual–math + text reasoning.
  • β€”Three-Stage SRO Training
  1. 1.Text Cold-Start β€” seed deep reflection
  2. 2.Multimodal RL β€” align vision & logic
  3. 3.Text RL β€” polish fluency & brevity
  4. 4.PAD (Prioritized Advantage Distillation) keeps gradients alive.
  5. 5.Efficient-Length Reward = concise, self-reflective CoT.

πŸ“š Resources


πŸ“Œ Citation

bibtex
@article{chen2025advancing,
  title={Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning},
  author={Chen, Shuang and Guo, Yue and Su, Zhaochen and Li, Yafu and Wu, Yulun and Chen, Jiacheng and Chen, Jiayu and Wang, Weijie and Qu, Xiaoye and Cheng, Yu},
  journal={arXiv preprint arXiv:2506.04207},
  year={2025}
}

Take ReVisual-R1 for a spin and let us know what you build! 🎯