CoolFace
Modelpublic

Zhang199/TinyLLaVA-Video-R1

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
3likes35downloads
Model Card

<center><span style="font-size:2em;">TinyLLaVA-Video-R1</span></center>

![arXiv](https://arxiv.org/abs/2504.09641)![Github](https://github.com/ZhangXJ199/TinyLLaVA-Video-R1)

Here, we introduce a small-scale video reasoning model TinyLLaVA-Video-R1, based on the traceably trained model TinyLLaVA-Video. After reinforcement learning on general Video-QA datasets, the model not only significantly improves its reasoning and thinking abilities, but also exhibits the emergent characteristic of “aha moments”.

Result

Model (HF Path)Video-MME(wo sub)MVBenchMLVUMMVU(mc)
Zhang199/TinyLLaVA-Video-R146.649.552.446.9