CoolFace
Modelpublic

lmms-lab/llava-critic-7b

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
15likes342downloads
Model Card

LLaVA-Critic-7B

Model Summary

llava-critic-7b is the first open-source large multimodal model (LMM) designed as a generalist evaluator for assessing model performance across diverse multimodal scenarios. Built on the foundation of llava-onevision-7b-ov, it has been finetuned on LLaVA-Critic-113k dataset to develop its "critic" capacities.

LLaVA-Critic excels in two primary scenarios:

  • 1️⃣ LMM-as-a-Judge: It delivers judgments closely aligned with human, and provides concrete, image-grounded reasons. An open-source alternative to GPT for evaluations.
  • 2️⃣ Preference Learning: Reliable reward signals power up visual chat, leading to LLaVA-OV-Chat 7B/72B.

For further details, please refer to the following resources:

  • 📰 Paper: https://arxiv.org/abs/2410.02712
  • 🪐 Project Page: https://llava-vl.github.io/blog/2024-10-03-llava-critic/
  • 📦 Datasets: https://huggingface.co/datasets/lmms-lab/llava-critic-113k
  • 🤗 Model Collections: https://huggingface.co/collections/lmms-lab/llava-critic-66fe3ef8c6e586d8435b4af8
  • 👋 Point of Contact: Tianyi Xiong

Use

Intended Use

The model demonstrates general capacities in providing quantitative judgments and qualitative justifications for evaluating LMM-generated responses. It mainly focuses on two evaluation settings:

  • Pointwise scoring, where it assigns a score to an individual candidate response.
  • Pairwise ranking, where it compares two candidate responses to determine their relative quality.

Quick Start

~~~python

pip install git+https://github.com/LLaVA-VL/LLaVA-NeXT.git

from llava.model.builder import loadpretrainedmodel from llava.mmutils import getmodelnamefrompath, processimages, tokenizerimagetoken from llava.constants import IMAGETOKENINDEX, DEFAULTIMAGETOKEN, DEFAULTIMSTARTTOKEN, DEFAULTIMENDTOKEN, IGNOREINDEX from llava.conversation import convtemplates, SeparatorStyle

from PIL import Image import requests import copy import torch

import sys import warnings import os

warnings.filterwarnings("ignore") pretrained = "lmms-lab/llava-critic-7b" modelname = "llavaqwen" device = "cuda" devicemap = "auto" tokenizer, model, imageprocessor, maxlength = loadpretrainedmodel(pretrained, None, modelname, devicemap=devicemap) # Add any other thing you want to pass in llavamodelargs

model.eval()

url = "https://github.com/LLaVA-VL/blog/blob/main/2024-10-03-llava-critic/static/images/criticimgseven.png?raw=True" image = Image.open(requests.get(url, stream=True).raw) imagetensor = processimages([image], imageprocessor, model.config) imagetensor = [image.to(dtype=torch.float16, device=device) for image in image_tensor]

convtemplate = "qwen1_5" # Make sure you use correct chat template for different models

pairwise ranking

critic_prompt = "Given an image and a corresponding question, please serve as an unbiased and fair judge to evaluate the quality of the answers provided by a Large Multimodal Model (LMM). Determine which answer is better and explain your reasoning with specific details. Your task is provided as follows:\nQuestion: [What this image presents?]\nThe first response: [The image is a black and white sketch of a line that appears to be in the shape of a cross. The line is a simple and straightforward representation of the cross shape, with two straight lines intersecting at a point.]\nThe second response: [This is a handwritten number seven.]\nASSISTANT:\n"

pointwise scoring

critic_prompt = "Given an image and a corresponding question, please serve as an unbiased and fair judge to evaluate the quality of answer answers provided by a Large Multimodal Model (LMM). Score the response out of 100 and explain your reasoning with specific details. Your task is provided as follows:\nQuestion: [What this image presents?]\nThe LMM response: [This is a handwritten number seven.]\nASSISTANT:\n "

question = DEFAULTIMAGETOKEN + "\n" + criticprompt conv = copy.deepcopy(convtemplates[convtemplate]) conv.appendmessage(conv.roles[0], question) conv.appendmessage(conv.roles[1], None) promptquestion = conv.get_prompt()

inputids = tokenizerimagetoken(promptquestion, tokenizer, IMAGETOKENINDEX, returntensors="pt").unsqueeze(0).to(device) imagesizes = [image.size]

cont = model.generate( inputids, images=imagetensor, imagesizes=imagesizes, dosample=False, temperature=0, maxnewtokens=4096, ) textoutputs = tokenizer.batchdecode(cont, skipspecialtokens=True) print(textoutputs[0]) ~~~

Citation

@article{xiong2024llavacritic,
  title={LLaVA-Critic: Learning to Evaluate Multimodal Models},
  author={Xiong, Tianyi and Wang, Xiyao and Guo, Dong and Ye, Qinghao and Fan, Haoqi and Gu, Quanquan and Huang, Heng and Li, Chunyuan},
  year={2024},
  eprint={2410.02712},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2410.02712},
}