CoolFace
Modelpublic

PommesPeter/prism-qwen25-extra-siglip-224px-0_5b

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes12downloads
Model Card

Prism with Qwen 2.5 0.5B backbone (Prismatic-Compatible Version)

This model is trained on the Llava-1.5-Instruct dataset. The official MiniVLA model uses a single vision encoder (Siglip), which is the key difference.

Usage Instructions

See the MiniVLA GitHub README for instructions on how to use this checkpoint for downstream training and finetuning.

Reference

BibTeX:

bibtex
@article{belkhale24minivla,
    title={MiniVLA: A Better VLA with a Smaller Footprint},
    author={Suneel Belkhale and Dorsa Sadigh},
    url={https://github.com/Stanford-ILIAD/openvla-mini}
    year={2024}
}