OpenGVLab/InternVL2-1B
83696k
1---2license: mit3pipeline_tag: image-text-to-text4library_name: transformers5base_model:6 - OpenGVLab/InternViT-300M-448px7 - Qwen/Qwen2-0.5B-Instruct8new_version: OpenGVLab/InternVL2_5-1B9base_model_relation: merge10language:11 - multilingual12tags:13 - internvl14 - custom_code15---16 17# InternVL2-1B18 19[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271)20 21[\[🆕 Blog\]](https://internvl.github.io/blog/) [\[🗨️ Chat Demo\]](https://internvl.opengvlab.com/) [\[🤗 HF Demo\]](https://huggingface.co/spaces/OpenGVLab/InternVL) [\[🚀 Quick Start\]](#quick-start) [\[📖 Documents\]](https://internvl.readthedocs.io/en/latest/)22 23<div align="center">24 <img width="500" alt="image" src="https://cdn-uploads.huggingface.co/production/uploads/64006c09330a45b03605bba3/zJsd2hqd3EevgXo6fNgC-.png">25</div>26 27## Introduction28 29We are excited to announce the release of InternVL 2.0, the latest addition to the InternVL series of multimodal large language models. InternVL 2.0 features a variety of **instruction-tuned models**, ranging from 1 billion to 108 billion parameters. This repository contains the instruction-tuned InternVL2-1B model.30 31Compared to the state-of-the-art open-source multimodal large language models, InternVL 2.0 surpasses most open-source models. It demonstrates competitive performance on par with proprietary commercial models across various capabilities, including document and chart comprehension, infographics QA, scene text understanding and OCR tasks, scientific and mathematical problem solving, as well as cultural understanding and integrated multimodal capabilities.32 33InternVL 2.0 is trained with an 8k context window and utilizes training data consisting of long texts, multiple images, and videos, significantly improving its ability to handle these types of inputs compared to InternVL 1.5. For more details, please refer to our [blog](https://internvl.github.io/blog/2024-07-02-InternVL-2.0/) and [GitHub](https://github.com/OpenGVLab/InternVL).34 35| Model Name | Vision Part | Language Part | HF Link | MS Link |36| :------------------: | :---------------------------------------------------------------------------------: | :------------------------------------------------------------------------------------------: | :--------------------------------------------------------------: | :--------------------------------------------------------------------: |37| InternVL2-1B | [InternViT-300M-448px](https://huggingface.co/OpenGVLab/InternViT-300M-448px) | [Qwen2-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2-0.5B-Instruct) | [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-1B) | [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-1B) |38| InternVL2-2B | [InternViT-300M-448px](https://huggingface.co/OpenGVLab/InternViT-300M-448px) | [internlm2-chat-1_8b](https://huggingface.co/internlm/internlm2-chat-1_8b) | [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-2B) | [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-2B) |39| InternVL2-4B | [InternViT-300M-448px](https://huggingface.co/OpenGVLab/InternViT-300M-448px) | [Phi-3-mini-128k-instruct](https://huggingface.co/microsoft/Phi-3-mini-128k-instruct) | [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-4B) | [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-4B) |40| InternVL2-8B | [InternViT-300M-448px](https://huggingface.co/OpenGVLab/InternViT-300M-448px) | [internlm2_5-7b-chat](https://huggingface.co/internlm/internlm2_5-7b-chat) | [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-8B) | [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-8B) |41| InternVL2-26B | [InternViT-6B-448px-V1-5](https://huggingface.co/OpenGVLab/InternViT-6B-448px-V1-5) | [internlm2-chat-20b](https://huggingface.co/internlm/internlm2-chat-20b) | [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-26B) | [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-26B) |42| InternVL2-40B | [InternViT-6B-448px-V1-5](https://huggingface.co/OpenGVLab/InternViT-6B-448px-V1-5) | [Nous-Hermes-2-Yi-34B](https://huggingface.co/NousResearch/Nous-Hermes-2-Yi-34B) | [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-40B) | [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-40B) |43| InternVL2-Llama3-76B | [InternViT-6B-448px-V1-5](https://huggingface.co/OpenGVLab/InternViT-6B-448px-V1-5) | [Hermes-2-Theta-Llama-3-70B](https://huggingface.co/NousResearch/Hermes-2-Theta-Llama-3-70B) | [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-Llama3-76B) | [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-Llama3-76B) |44 45## Model Details46 47InternVL 2.0 is a multimodal large language model series, featuring models of various sizes. For each size, we release instruction-tuned models optimized for multimodal tasks. InternVL2-1B consists of [InternViT-300M-448px](https://huggingface.co/OpenGVLab/InternViT-300M-448px), an MLP projector, and [Qwen2-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2-0.5B-Instruct).48 49## Performance50 51### Image Benchmarks52 53| Benchmark | PaliGemma-3B | Mini-InternVL-2B-1-5 | InternVL2-1B |54| :--------------------------: | :----------: | :------------------: | :----------: |55| Model Size | 2.9B | 2.2B | 0.9B |56| | | | |57| DocVQA<sub>test</sub> | - | 85.0 | 81.7 |58| ChartQA<sub>test</sub> | - | 74.8 | 72.9 |59| InfoVQA<sub>test</sub> | - | 55.4 | 50.9 |60| TextVQA<sub>val</sub> | 68.1 | 70.5 | 70.5 |61| OCRBench | 614 | 654 | 754 |62| MME<sub>sum</sub> | 1686.1 | 1901.5 | 1794.4 |63| RealWorldQA | 55.2 | 57.9 | 50.3 |64| AI2D<sub>test</sub> | 68.3 | 69.8 | 64.1 |65| MMMU<sub>val</sub> | 34.9 | 37.4 | 36.7 |66| MMBench-EN<sub>test</sub> | 71.0 | 70.9 | 65.4 |67| MMBench-CN<sub>test</sub> | 63.6 | 66.2 | 60.7 |68| CCBench<sub>dev</sub> | 29.6 | 63.5 | 75.7 |69| MMVet<sub>GPT-4-0613</sub> | - | 39.3 | 37.8 |70| MMVet<sub>GPT-4-Turbo</sub> | 33.1 | 35.5 | 32.7 |71| SEED-Image | 69.6 | 69.8 | 65.6 |72| HallBench<sub>avg</sub> | 32.2 | 37.5 | 34.0 |73| MathVista<sub>testmini</sub> | 28.7 | 41.1 | 37.7 |74| OpenCompass<sub>avg</sub> | 46.6 | 49.8 | 48.3 |75 76- For more details and evaluation reproduction, please refer to our [Evaluation Guide](https://internvl.readthedocs.io/en/latest/internvl2.0/evaluation.html).77 78- We simultaneously use [InternVL](https://github.com/OpenGVLab/InternVL) and [VLMEvalKit](https://github.com/open-compass/VLMEvalKit) repositories for model evaluation. Specifically, the results reported for DocVQA, ChartQA, InfoVQA, TextVQA, MME, AI2D, MMBench, CCBench, MMVet (GPT-4-0613), and SEED-Image were tested using the InternVL repository. MMMU, OCRBench, RealWorldQA, HallBench, MMVet (GPT-4-Turbo), and MathVista were evaluated using the VLMEvalKit.79 80### Video Benchmarks81 82| Benchmark | VideoChat2-Phi3 | Mini-InternVL-2B-1-5 | InternVL2-1B |83| :-------------------------: | :-------------: | :------------------: | :----------: |84| Model Size | 4B | 2.2B | 0.9B |85| | | | |86| MVBench | 55.1 | 37.0 | 57.5 |87| MMBench-Video<sub>8f</sub> | - | 0.99 | 0.95 |88| MMBench-Video<sub>16f</sub> | - | 1.04 | 0.98 |89| Video-MME<br>w/o subs | - | 42.9 | 42.6 |90| Video-MME<br>w subs | - | 44.7 | 44.7 |91 92- We evaluate our models on MVBench and Video-MME by extracting 16 frames from each video, and each frame was resized to a 448x448 image.93 94### Grounding Benchmarks95 96| Model | avg. | RefCOCO<br>(val) | RefCOCO<br>(testA) | RefCOCO<br>(testB) | RefCOCO+<br>(val) | RefCOCO+<br>(testA) | RefCOCO+<br>(testB) | RefCOCO‑g<br>(val) | RefCOCO‑g<br>(test) |97| :----------------------------: | :--: | :--------------: | :----------------: | :----------------: | :---------------: | :-----------------: | :-----------------: | :----------------: | :-----------------: |98| UNINEXT-H<br>(Specialist SOTA) | 88.9 | 92.6 | 94.3 | 91.5 | 85.2 | 89.6 | 79.8 | 88.7 | 89.4 |99| | | | | | | | | | |100| Mini-InternVL-<br>Chat-2B-V1-5 | 75.8 | 80.7 | 86.7 | 72.9 | 72.5 | 82.3 | 60.8 | 75.6 | 74.9 |101| Mini-InternVL-<br>Chat-4B-V1-5 | 84.4 | 88.0 | 91.4 | 83.5 | 81.5 | 87.4 | 73.8 | 84.7 | 84.6 |102| InternVL‑Chat‑V1‑5 | 88.8 | 91.4 | 93.7 | 87.1 | 87.0 | 92.3 | 80.9 | 88.5 | 89.3 |103| | | | | | | | | | |104| InternVL2‑1B | 79.9 | 83.6 | 88.7 | 79.8 | 76.0 | 83.6 | 67.7 | 80.2 | 79.9 |105| InternVL2‑2B | 77.7 | 82.3 | 88.2 | 75.9 | 73.5 | 82.8 | 63.3 | 77.6 | 78.3 |106| InternVL2‑4B | 84.4 | 88.5 | 91.2 | 83.9 | 81.2 | 87.2 | 73.8 | 84.6 | 84.6 |107| InternVL2‑8B | 82.9 | 87.1 | 91.1 | 80.7 | 79.8 | 87.9 | 71.4 | 82.7 | 82.7 |108| InternVL2‑26B | 88.5 | 91.2 | 93.3 | 87.4 | 86.8 | 91.0 | 81.2 | 88.5 | 88.6 |109| InternVL2‑40B | 90.3 | 93.0 | 94.7 | 89.2 | 88.5 | 92.8 | 83.6 | 90.3 | 90.6 |110| InternVL2-<br>Llama3‑76B | 90.0 | 92.2 | 94.8 | 88.4 | 88.8 | 93.1 | 82.8 | 89.5 | 90.3 |111 112- We use the following prompt to evaluate InternVL's grounding ability: `Please provide the bounding box coordinates of the region this sentence describes: <ref>{}</ref>`113 114Limitations: Although we have made efforts to ensure the safety of the model during the training process and to encourage the model to generate text that complies with ethical and legal requirements, the model may still produce unexpected outputs due to its size and probabilistic generation paradigm. For example, the generated responses may contain biases, discrimination, or other harmful content. Please do not propagate such content. We are not responsible for any consequences resulting from the dissemination of harmful information.115 116## Quick Start117 118We provide an example code to run `InternVL2-1B` using `transformers`.119 120> Please use transformers>=4.37.2 to ensure the model works normally.121 122### Model Loading123 124#### 16-bit (bf16 / fp16)125 126```python127import torch128from transformers import AutoTokenizer, AutoModel129path = "OpenGVLab/InternVL2-1B"130model = AutoModel.from_pretrained(131 path,132 torch_dtype=torch.bfloat16,133 low_cpu_mem_usage=True,134 use_flash_attn=True,135 trust_remote_code=True).eval().cuda()136```137 138#### BNB 8-bit Quantization139 140```python141import torch142from transformers import AutoTokenizer, AutoModel143path = "OpenGVLab/InternVL2-1B"144model = AutoModel.from_pretrained(145 path,146 torch_dtype=torch.bfloat16,147 load_in_8bit=True,148 low_cpu_mem_usage=True,149 use_flash_attn=True,150 trust_remote_code=True).eval()151```152 153#### Multiple GPUs154 155The reason for writing the code this way is to avoid errors that occur during multi-GPU inference due to tensors not being on the same device. By ensuring that the first and last layers of the large language model (LLM) are on the same device, we prevent such errors.156 157```python158import math159import torch160from transformers import AutoTokenizer, AutoModel161 162def split_model(model_name):163 device_map = {}164 world_size = torch.cuda.device_count()165 num_layers = {166 'InternVL2-1B': 24, 'InternVL2-2B': 24, 'InternVL2-4B': 32, 'InternVL2-8B': 32,167 'InternVL2-26B': 48, 'InternVL2-40B': 60, 'InternVL2-Llama3-76B': 80}[model_name]168 # Since the first GPU will be used for ViT, treat it as half a GPU.169 num_layers_per_gpu = math.ceil(num_layers / (world_size - 0.5))170 num_layers_per_gpu = [num_layers_per_gpu] * world_size171 num_layers_per_gpu[0] = math.ceil(num_layers_per_gpu[0] * 0.5)172 layer_cnt = 0173 for i, num_layer in enumerate(num_layers_per_gpu):174 for j in range(num_layer):175 device_map[f'language_model.model.layers.{layer_cnt}'] = i176 layer_cnt += 1177 device_map['vision_model'] = 0178 device_map['mlp1'] = 0179 device_map['language_model.model.tok_embeddings'] = 0180 device_map['language_model.model.embed_tokens'] = 0181 device_map['language_model.output'] = 0182 device_map['language_model.model.norm'] = 0183 device_map['language_model.model.rotary_emb'] = 0184 device_map['language_model.lm_head'] = 0185 device_map[f'language_model.model.layers.{num_layers - 1}'] = 0186 187 return device_map188 189path = "OpenGVLab/InternVL2-1B"190device_map = split_model('InternVL2-1B')191model = AutoModel.from_pretrained(192 path,193 torch_dtype=torch.bfloat16,194 low_cpu_mem_usage=True,195 use_flash_attn=True,196 trust_remote_code=True,197 device_map=device_map).eval()198```199 200### Inference with Transformers201 202```python203import numpy as np204import torch205import torchvision.transforms as T206from decord import VideoReader, cpu207from PIL import Image208from torchvision.transforms.functional import InterpolationMode209from transformers import AutoModel, AutoTokenizer210 211IMAGENET_MEAN = (0.485, 0.456, 0.406)212IMAGENET_STD = (0.229, 0.224, 0.225)213 214def build_transform(input_size):215 MEAN, STD = IMAGENET_MEAN, IMAGENET_STD216 transform = T.Compose([217 T.Lambda(lambda img: img.convert('RGB') if img.mode != 'RGB' else img),218 T.Resize((input_size, input_size), interpolation=InterpolationMode.BICUBIC),219 T.ToTensor(),220 T.Normalize(mean=MEAN, std=STD)221 ])222 return transform223 224def find_closest_aspect_ratio(aspect_ratio, target_ratios, width, height, image_size):225 best_ratio_diff = float('inf')226 best_ratio = (1, 1)227 area = width * height228 for ratio in target_ratios:229 target_aspect_ratio = ratio[0] / ratio[1]230 ratio_diff = abs(aspect_ratio - target_aspect_ratio)231 if ratio_diff < best_ratio_diff:232 best_ratio_diff = ratio_diff233 best_ratio = ratio234 elif ratio_diff == best_ratio_diff:235 if area > 0.5 * image_size * image_size * ratio[0] * ratio[1]:236 best_ratio = ratio237 return best_ratio238 239def dynamic_preprocess(image, min_num=1, max_num=12, image_size=448, use_thumbnail=False):240 orig_width, orig_height = image.size241 aspect_ratio = orig_width / orig_height242 243 # calculate the existing image aspect ratio244 target_ratios = set(245 (i, j) for n in range(min_num, max_num + 1) for i in range(1, n + 1) for j in range(1, n + 1) if246 i * j <= max_num and i * j >= min_num)247 target_ratios = sorted(target_ratios, key=lambda x: x[0] * x[1])248 249 # find the closest aspect ratio to the target250 target_aspect_ratio = find_closest_aspect_ratio(251 aspect_ratio, target_ratios, orig_width, orig_height, image_size)252 253 # calculate the target width and height254 target_width = image_size * target_aspect_ratio[0]255 target_height = image_size * target_aspect_ratio[1]256 blocks = target_aspect_ratio[0] * target_aspect_ratio[1]257 258 # resize the image259 resized_img = image.resize((target_width, target_height))260 processed_images = []261 for i in range(blocks):262 box = (263 (i % (target_width // image_size)) * image_size,264 (i // (target_width // image_size)) * image_size,265 ((i % (target_width // image_size)) + 1) * image_size,266 ((i // (target_width // image_size)) + 1) * image_size267 )268 # split the image269 split_img = resized_img.crop(box)270 processed_images.append(split_img)271 assert len(processed_images) == blocks272 if use_thumbnail and len(processed_images) != 1:273 thumbnail_img = image.resize((image_size, image_size))274 processed_images.append(thumbnail_img)275 return processed_images276 277def load_image(image_file, input_size=448, max_num=12):278 image = Image.open(image_file).convert('RGB')279 transform = build_transform(input_size=input_size)280 images = dynamic_preprocess(image, image_size=input_size, use_thumbnail=True, max_num=max_num)281 pixel_values = [transform(image) for image in images]282 pixel_values = torch.stack(pixel_values)283 return pixel_values284 285# If you want to load a model using multiple GPUs, please refer to the `Multiple GPUs` section.286path = 'OpenGVLab/InternVL2-1B'287model = AutoModel.from_pretrained(288 path,289 torch_dtype=torch.bfloat16,290 low_cpu_mem_usage=True,291 use_flash_attn=True,292 trust_remote_code=True).eval().cuda()293tokenizer = AutoTokenizer.from_pretrained(path, trust_remote_code=True, use_fast=False)294 295# set the max number of tiles in `max_num`296pixel_values = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()297generation_config = dict(max_new_tokens=1024, do_sample=True)298 299# pure-text conversation (纯文本对话)300question = 'Hello, who are you?'301response, history = model.chat(tokenizer, None, question, generation_config, history=None, return_history=True)302print(f'User: {question}\nAssistant: {response}')303 304question = 'Can you tell me a story?'305response, history = model.chat(tokenizer, None, question, generation_config, history=history, return_history=True)306print(f'User: {question}\nAssistant: {response}')307 308# single-image single-round conversation (单图单轮对话)309question = '<image>\nPlease describe the image shortly.'310response = model.chat(tokenizer, pixel_values, question, generation_config)311print(f'User: {question}\nAssistant: {response}')312 313# single-image multi-round conversation (单图多轮对话)314question = '<image>\nPlease describe the image in detail.'315response, history = model.chat(tokenizer, pixel_values, question, generation_config, history=None, return_history=True)316print(f'User: {question}\nAssistant: {response}')317 318question = 'Please write a poem according to the image.'319response, history = model.chat(tokenizer, pixel_values, question, generation_config, history=history, return_history=True)320print(f'User: {question}\nAssistant: {response}')321 322# multi-image multi-round conversation, combined images (多图多轮对话,拼接图像)323pixel_values1 = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()324pixel_values2 = load_image('./examples/image2.jpg', max_num=12).to(torch.bfloat16).cuda()325pixel_values = torch.cat((pixel_values1, pixel_values2), dim=0)326 327question = '<image>\nDescribe the two images in detail.'328response, history = model.chat(tokenizer, pixel_values, question, generation_config,329 history=None, return_history=True)330print(f'User: {question}\nAssistant: {response}')331 332question = 'What are the similarities and differences between these two images.'333response, history = model.chat(tokenizer, pixel_values, question, generation_config,334 history=history, return_history=True)335print(f'User: {question}\nAssistant: {response}')336 337# multi-image multi-round conversation, separate images (多图多轮对话,独立图像)338pixel_values1 = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()339pixel_values2 = load_image('./examples/image2.jpg', max_num=12).to(torch.bfloat16).cuda()340pixel_values = torch.cat((pixel_values1, pixel_values2), dim=0)341num_patches_list = [pixel_values1.size(0), pixel_values2.size(0)]342 343question = 'Image-1: <image>\nImage-2: <image>\nDescribe the two images in detail.'344response, history = model.chat(tokenizer, pixel_values, question, generation_config,345 num_patches_list=num_patches_list,346 history=None, return_history=True)347print(f'User: {question}\nAssistant: {response}')348 349question = 'What are the similarities and differences between these two images.'350response, history = model.chat(tokenizer, pixel_values, question, generation_config,351 num_patches_list=num_patches_list,352 history=history, return_history=True)353print(f'User: {question}\nAssistant: {response}')354 355# batch inference, single image per sample (单图批处理)356pixel_values1 = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()357pixel_values2 = load_image('./examples/image2.jpg', max_num=12).to(torch.bfloat16).cuda()358num_patches_list = [pixel_values1.size(0), pixel_values2.size(0)]359pixel_values = torch.cat((pixel_values1, pixel_values2), dim=0)360 361questions = ['<image>\nDescribe the image in detail.'] * len(num_patches_list)362responses = model.batch_chat(tokenizer, pixel_values,363 num_patches_list=num_patches_list,364 questions=questions,365 generation_config=generation_config)366for question, response in zip(questions, responses):367 print(f'User: {question}\nAssistant: {response}')368 369# video multi-round conversation (视频多轮对话)370def get_index(bound, fps, max_frame, first_idx=0, num_segments=32):371 if bound:372 start, end = bound[0], bound[1]373 else:374 start, end = -100000, 100000375 start_idx = max(first_idx, round(start * fps))376 end_idx = min(round(end * fps), max_frame)377 seg_size = float(end_idx - start_idx) / num_segments378 frame_indices = np.array([379 int(start_idx + (seg_size / 2) + np.round(seg_size * idx))380 for idx in range(num_segments)381 ])382 return frame_indices383 384def load_video(video_path, bound=None, input_size=448, max_num=1, num_segments=32):385 vr = VideoReader(video_path, ctx=cpu(0), num_threads=1)386 max_frame = len(vr) - 1387 fps = float(vr.get_avg_fps())388 389 pixel_values_list, num_patches_list = [], []390 transform = build_transform(input_size=input_size)391 frame_indices = get_index(bound, fps, max_frame, first_idx=0, num_segments=num_segments)392 for frame_index in frame_indices:393 img = Image.fromarray(vr[frame_index].asnumpy()).convert('RGB')394 img = dynamic_preprocess(img, image_size=input_size, use_thumbnail=True, max_num=max_num)395 pixel_values = [transform(tile) for tile in img]396 pixel_values = torch.stack(pixel_values)397 num_patches_list.append(pixel_values.shape[0])398 pixel_values_list.append(pixel_values)399 pixel_values = torch.cat(pixel_values_list)400 return pixel_values, num_patches_list401 402video_path = './examples/red-panda.mp4'403pixel_values, num_patches_list = load_video(video_path, num_segments=8, max_num=1)404pixel_values = pixel_values.to(torch.bfloat16).cuda()405video_prefix = ''.join([f'Frame{i+1}: <image>\n' for i in range(len(num_patches_list))])406question = video_prefix + 'What is the red panda doing?'407# Frame1: <image>\nFrame2: <image>\n...\nFrame8: <image>\n{question}408response, history = model.chat(tokenizer, pixel_values, question, generation_config,409 num_patches_list=num_patches_list, history=None, return_history=True)410print(f'User: {question}\nAssistant: {response}')411 412question = 'Describe this video in detail.'413response, history = model.chat(tokenizer, pixel_values, question, generation_config,414 num_patches_list=num_patches_list, history=history, return_history=True)415print(f'User: {question}\nAssistant: {response}')416```417 418#### Streaming Output419 420Besides this method, you can also use the following code to get streamed output.421 422```python423from transformers import TextIteratorStreamer424from threading import Thread425 426# Initialize the streamer427streamer = TextIteratorStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True, timeout=10)428# Define the generation configuration429generation_config = dict(max_new_tokens=1024, do_sample=False, streamer=streamer)430# Start the model chat in a separate thread431thread = Thread(target=model.chat, kwargs=dict(432 tokenizer=tokenizer, pixel_values=pixel_values, question=question,433 history=None, return_history=False, generation_config=generation_config,434))435thread.start()436 437# Initialize an empty string to store the generated text438generated_text = ''439# Loop through the streamer to get the new text as it is generated440for new_text in streamer:441 if new_text == model.conv_template.sep:442 break443 generated_text += new_text444 print(new_text, end='', flush=True) # Print each new chunk of generated text on the same line445```446 447## Finetune448 449Many repositories now support fine-tuning of the InternVL series models, including [InternVL](https://github.com/OpenGVLab/InternVL), [SWIFT](https://github.com/modelscope/ms-swift), [XTurner](https://github.com/InternLM/xtuner), and others. Please refer to their documentation for more details on fine-tuning.450 451## Deployment452 453### LMDeploy454 455LMDeploy is a toolkit for compressing, deploying, and serving LLMs & VLMs.456 457```sh458pip install lmdeploy>=0.5.3459```460 461LMDeploy abstracts the complex inference process of multi-modal Vision-Language Models (VLM) into an easy-to-use pipeline, similar to the Large Language Model (LLM) inference pipeline.462 463#### A 'Hello, world' Example464 465```python466from lmdeploy import pipeline, TurbomindEngineConfig467from lmdeploy.vl import load_image468 469model = 'OpenGVLab/InternVL2-1B'470image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/tests/data/tiger.jpeg')471pipe = pipeline(model, backend_config=TurbomindEngineConfig(session_len=8192))472response = pipe(('describe this image', image))473print(response.text)474```475 476If `ImportError` occurs while executing this case, please install the required dependency packages as prompted.477 478#### Multi-images Inference479 480When dealing with multiple images, you can put them all in one list. Keep in mind that multiple images will lead to a higher number of input tokens, and as a result, the size of the context window typically needs to be increased.481 482> Warning: Due to the scarcity of multi-image conversation data, the performance on multi-image tasks may be unstable, and it may require multiple attempts to achieve satisfactory results.483 484```python485from lmdeploy import pipeline, TurbomindEngineConfig486from lmdeploy.vl import load_image487from lmdeploy.vl.constants import IMAGE_TOKEN488 489model = 'OpenGVLab/InternVL2-1B'490pipe = pipeline(model, backend_config=TurbomindEngineConfig(session_len=8192))491 492image_urls=[493 'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg',494 'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/det.jpg'495]496 497images = [load_image(img_url) for img_url in image_urls]498# Numbering images improves multi-image conversations499response = pipe((f'Image-1: {IMAGE_TOKEN}\nImage-2: {IMAGE_TOKEN}\ndescribe these two images', images))500print(response.text)501```502 503#### Batch Prompts Inference504 505Conducting inference with batch prompts is quite straightforward; just place them within a list structure:506 507```python508from lmdeploy import pipeline, TurbomindEngineConfig509from lmdeploy.vl import load_image510 511model = 'OpenGVLab/InternVL2-1B'512pipe = pipeline(model, backend_config=TurbomindEngineConfig(session_len=8192))513 514image_urls=[515 "https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg",516 "https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/det.jpg"517]518prompts = [('describe this image', load_image(img_url)) for img_url in image_urls]519response = pipe(prompts)520print(response)521```522 523#### Multi-turn Conversation524 525There are two ways to do the multi-turn conversations with the pipeline. One is to construct messages according to the format of OpenAI and use above introduced method, the other is to use the `pipeline.chat` interface.526 527```python528from lmdeploy import pipeline, TurbomindEngineConfig, GenerationConfig529from lmdeploy.vl import load_image530 531model = 'OpenGVLab/InternVL2-1B'532pipe = pipeline(model, backend_config=TurbomindEngineConfig(session_len=8192))533 534image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg')535gen_config = GenerationConfig(top_k=40, top_p=0.8, temperature=0.8)536sess = pipe.chat(('describe this image', image), gen_config=gen_config)537print(sess.response.text)538sess = pipe.chat('What is the woman doing?', session=sess, gen_config=gen_config)539print(sess.response.text)540```541 542#### Service543 544LMDeploy's `api_server` enables models to be easily packed into services with a single command. The provided RESTful APIs are compatible with OpenAI's interfaces. Below are an example of service startup:545 546```shell547lmdeploy serve api_server OpenGVLab/InternVL2-1B --server-port 23333548```549 550To use the OpenAI-style interface, you need to install OpenAI:551 552```shell553pip install openai554```555 556Then, use the code below to make the API call:557 558```python559from openai import OpenAI560 561client = OpenAI(api_key='YOUR_API_KEY', base_url='http://0.0.0.0:23333/v1')562model_name = client.models.list().data[0].id563response = client.chat.completions.create(564 model=model_name,565 messages=[{566 'role':567 'user',568 'content': [{569 'type': 'text',570 'text': 'describe this image',571 }, {572 'type': 'image_url',573 'image_url': {574 'url':575 'https://modelscope.oss-cn-beijing.aliyuncs.com/resource/tiger.jpeg',576 },577 }],578 }],579 temperature=0.8,580 top_p=0.8)581print(response)582```583 584## License585 586This project is released under the MIT License. This project uses the pre-trained Qwen2-0.5B-Instruct as a component, which is licensed under the Apache License 2.0.587 588## Citation589 590If you find this project useful in your research, please consider citing:591 592```BibTeX593@article{chen2024expanding,594 title={Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling},595 author={Chen, Zhe and Wang, Weiyun and Cao, Yue and Liu, Yangzhou and Gao, Zhangwei and Cui, Erfei and Zhu, Jinguo and Ye, Shenglong and Tian, Hao and Liu, Zhaoyang and others},596 journal={arXiv preprint arXiv:2412.05271},597 year={2024}598}599@article{gao2024mini,600 title={Mini-internvl: A flexible-transfer pocket multimodal model with 5\% parameters and 90\% performance},601 author={Gao, Zhangwei and Chen, Zhe and Cui, Erfei and Ren, Yiming and Wang, Weiyun and Zhu, Jinguo and Tian, Hao and Ye, Shenglong and He, Junjun and Zhu, Xizhou and others},602 journal={arXiv preprint arXiv:2410.16261},603 year={2024}604}605@article{chen2024far,606 title={How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites},607 author={Chen, Zhe and Wang, Weiyun and Tian, Hao and Ye, Shenglong and Gao, Zhangwei and Cui, Erfei and Tong, Wenwen and Hu, Kongzhi and Luo, Jiapeng and Ma, Zheng and others},608 journal={arXiv preprint arXiv:2404.16821},609 year={2024}610}611@inproceedings{chen2024internvl,612 title={Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks},613 author={Chen, Zhe and Wu, Jiannan and Wang, Wenhai and Su, Weijie and Chen, Guo and Xing, Sen and Zhong, Muyan and Zhang, Qinglong and Zhu, Xizhou and Lu, Lewei and others},614 booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},615 pages={24185--24198},616 year={2024}617}618```619 