CewEhao/OPD-Aha-4B
<h2 align="center">๐๏ธ OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation</h2>
<p align="center"> <a href="https://arxiv.org/pdf/2609.16459"><img alt="Paper" src="https://img.shields.io/badge/arXiv-2609.16459-B31B1B?logo=arxiv"></a> <a href="https://github.com/Echochef/OPD-Aha"><img alt="Code" src="https://img.shields.io/badge/Code-GitHub-black?logo=github"></a> <a href="https://huggingface.co/CewEhao/OPD-Aha-4B"><img alt="HF Model" src="https://img.shields.io/badge/%F0%9F%A4%97%20HuggingFace-OPD--Aha--4B-yellow"></a> </p>
<p align="center"> ๐ค Hugging Face model: <a href="https://huggingface.co/CewEhao/OPD-Aha-4B">CewEhao/OPD-Aha-4B</a> ยท ๐ป Code: <a href="https://github.com/Echochef/OPD-Aha">Echochef/OPD-Aha</a> </p>
๐ Introduction
This is the official model card for OPD-Aha-4B, built on `Qwen/Qwen3.5-4B`.
OPD-Aha is an on-policy self-distillation framework for improving fine-grained visual perception and multimodal mathematical reasoning. It trains the model with a frozen visual teacher and a counterfactual visual input so that learning focuses on evidence that changes the teacher distribution.
โก Serving
The project provides a vLLM serving entrypoint:
git clone https://github.com/Echochef/OPD-Aha.git
cd OPD-Aha
MODEL_PATH=CewEhao/OPD-Aha-4B \
SERVED_MODEL_NAME=opd-aha-4b \
bash scripts/serve_model.sh๐๏ธ Training and evaluation
Training, checkpoint merging, inference, and evaluation code is available in `Echochef/OPD-Aha`. The repository includes fine-grained perception evaluation for V*Bench, HR-Bench, MME-RealWorld, and ZoomBench, together with mathematical reasoning evaluation for MathVista, MathVerse, WeMath, MathVision, and DynaMath.
๐ Acknowledgements
OPD-Aha builds on `Qwen`, `verl`, `vLLM`, and `Vision-OPD`.
๐ License
This model is released under the Apache-2.0 License. The base model and datasets remain subject to their respective licenses.
