CoolFace
Modelpublic

variante/llava-1.5-7b-llara-D-inBC-VIMA-80k

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
1likes8downloads
Model Card

<br> <be>

LLaRA Model Card

This model is released with paper [LLaRA: Supercharging Robot Learning Data for Vision-Language Policy](https://arxiv.org/abs/2406.20095)

Xiang Li<sup>1</sup>, Cristina Mata<sup>1</sup>, Jongwoo Park<sup>1</sup>, Kumara Kahatapitiya<sup>1</sup>, Yoo Sung Jang<sup>1</sup>, Jinghuan Shang<sup>1</sup>, Kanchana Ranasinghe<sup>1</sup>, Ryan Burgert<sup>1</sup>, Mu Cai<sup>2</sup>, Yong Jae Lee<sup>2</sup>, and Michael S. Ryoo<sup>1</sup>

<sup>1</sup>Stony Brook University <sup>2</sup>University of Wisconsin-Madison

Model details

Model type: LLaRA is an open-source visuomotor policy trained by fine-tuning LLaVA-7b-v1.5 on instruction-following data D-inBC, converted from VIMA-Data. For the conversion code, please refer to convert_vima.ipynb

Model date: llava-1.5-7b-llara-D-inBC-VIMA-80k was trained in June 2024.

Paper or resources for more information: https://github.com/LostXine/LLaRA

Where to send questions or comments about the model: https://github.com/LostXine/LLaRA/issues

Intended use

Primary intended uses: The primary use of LLaRA is research on large multimodal models for robotics.

Primary intended users: The primary intended users of the model are researchers and hobbyists in robotics, computer vision, natural language processing, machine learning, and artificial intelligence.