CoolFace
Modelpublic

Zhang199/TinyLLaVA-Video-Qwen2.5-3B-Group-16-512

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
1likes34downloads
Model Card

<center><span style="font-size:2em;">TinyLLaVA-Video</span></center>

![arXiv](https://arxiv.org/abs/2501.15513)![Github](https://github.com/ZhangXJ199/TinyLLaVA-Video)

Here, we introduce TinyLLaVA-Video-Qwen2.5-3B-Group-16-512. For LLM and vision tower, we choose Qwen2.5-3B and siglip-so400m-patch14-384, respectively. The model adopts the Video-Level Group Resampler, samples 16 frames from each video, and represents the video sequence using 512 tokens.

Result

Model (HF Path)#Frame/QueryVideo-MMEMVBenchLongVideoBenchMLVU
Zhang199/TinyLLaVA-Video-Qwen2.5-3B-Group-1fps-5121fps/51247.747.042.052.6
Zhang199/TinyLLaVA-Video-Qwen2.5-3B-Group-16-51216/51247.045.542.452.5
Zhang199/TinyLLaVA-Video-Qwen2.5-3B-Naive-16-51216/51244.742.537.648.1
Zhang199/TinyLLaVA-Video-Phi2-Naive-16-51216/51242.742.042.246.5