Zhang199/TinyLLaVA-Video-Qwen2.5-3B-Group-16-512
134
<center><span style="font-size:2em;">TinyLLaVA-Video</span></center>

Here, we introduce TinyLLaVA-Video-Qwen2.5-3B-Group-16-512. For LLM and vision tower, we choose Qwen2.5-3B and siglip-so400m-patch14-384, respectively. The model adopts the Video-Level Group Resampler, samples 16 frames from each video, and represents the video sequence using 512 tokens.
