Markr-AI/Gukbap-Ovis2-34B-VL
HumanF-MarkrAI/Gukbap-Ovis2-34B-VL🍚
Model Details🍚
Model Description
- Developed by: HumanF-MarkrAI
- Model type: Korean-VL-Ovis2-34B
- Language(s): Korean + English
- Context Length: 2048
- License: cc-by-nc-4.0
- Finetuned from model: AIDC-AI/Ovis2-34B.
Model Sources
When training, we used H100 80GB GPUx6.
Implications🍚
If you want to know our model's details, please see 🔥Gukbap-LMM Blog🔥. And also, we provided the Korean-LMM training code based Ovis!! 🔥Github🔥. Please star⭐⭐!!
Training Method (SFT)🧐
The following papers contain the foundational methodologies for the dataset and training methods we are currently proceeding.
SFT Text-Datasets (Private)
When we made the Open-Source based dataset, we use microsoft/WizardLM-2-8x22B through DeepInfra. Our datasets are made by Evolving system, which is propsed by WizardLM. In training, we used 1849 training dataset, and 200 validation dataset.
- Wizard-Korea-Datasets: MarkrAI/Markr_WizardLM_train_ver4.
Learning rate: 2e-5; Epoch: 3
Benchmakrs🤗
Global MM Benchmark Score (Zero-shot)
We internally evaluated VLMEvalKit. We utilized chatgpt-0125, gpt-4o-mini and gpt-4-turbo in MMBench, MathVista and MMVet, respectively.
HallusionBench score: (aAcc + fAcc + qAcc) / 3
Korean MM Benchmark Score (Zero-shot)
We internally evaluated 🔥our code🔥. We utilized gpt-4o-2024-08-06 in K-LLAVA-W evaluation.
MM Benchmarks
- Global MM Bench dataset: OpenCampass MM leaderboard
- Korean MM Bench dataset: NCSOFT.
Inference
import torch
from PIL import Image
from transformers import AutoModelForCausalLM
#import os
#os.environ["cuda_visible_devices"]="0"
# load model
if __name__ == '__main__':
# HumanF-MarkrAI/Gukbap-Ovis2-34B-VL
# AIDC-AI/Ovis2-34B
model = AutoModelForCausalLM.from_pretrained("HumanF-MarkrAI/Gukbap-Ovis2-34B-VL",
torch_dtype=torch.bfloat16,
multimodal_max_length=2048,
cache_dir="/data/cache/",
trust_remote_code=True).cuda()
text_tokenizer = model.get_text_tokenizer()
visual_tokenizer = model.get_visual_tokenizer()
# single-image input (K-LLAVA-W)
image_path = './images/ex_4.jpg'
images = [Image.open(image_path)]
max_partition = 9
text = '이미지에서 잘리지 않은 과일은 몇 개인가요?'
query = f'<image>\n{text}'
# format conversation
prompt, input_ids, pixel_values = model.preprocess_inputs(query, images, max_partition=max_partition)
attention_mask = torch.ne(input_ids, text_tokenizer.pad_token_id)
input_ids = input_ids.unsqueeze(0).to(device=model.device)
attention_mask = attention_mask.unsqueeze(0).to(device=model.device)
if pixel_values is not None:
pixel_values = pixel_values.to(dtype=visual_tokenizer.dtype, device=visual_tokenizer.device)
pixel_values = [pixel_values]
# generate output
with torch.inference_mode():
gen_kwargs = dict(
max_new_tokens=2048,
do_sample=False,
top_p=None,
top_k=None,
temperature=None,
repetition_penalty=None,
eos_token_id=model.generation_config.eos_token_id,
pad_token_id=text_tokenizer.pad_token_id,
use_cache=True
)
output_ids = model.generate(input_ids, pixel_values=pixel_values, attention_mask=attention_mask, **gen_kwargs)[0]
output = text_tokenizer.decode(output_ids, skip_special_tokens=True)
print(f'Output:\n{output}')Chat Prompt😶🌫️
<|im_start|>user<image>
Hello! My favorite food is Gukbap🍚!<|im_end|>
<|im_start|>assistant
(model answer)Gukbap-VL Series models🍚🍚
BibTeX
@article{HumanF-MarkrAI,
title={Gukbap-Ovis2-34B-VL},
author={MarkrAI},
year={2025},
url={https://huggingface.co/HumanF-MarkrAI}
}