CoolFace
Modelpublic

OpenGVLab/InternVL2-1B

sourceHugging Facemitupdated 1y agoView on Hugging Face
83likes696kdownloads
README.md619 linesDownload Raw Back to root
1---2license: mit3pipeline_tag: image-text-to-text4library_name: transformers5base_model:6  - OpenGVLab/InternViT-300M-448px7  - Qwen/Qwen2-0.5B-Instruct8new_version: OpenGVLab/InternVL2_5-1B9base_model_relation: merge10language:11  - multilingual12tags:13  - internvl14  - custom_code15---16 17# InternVL2-1B18 19[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL)  [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238)  [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821)  [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261)  [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271)20 21[\[🆕 Blog\]](https://internvl.github.io/blog/)  [\[🗨️ Chat Demo\]](https://internvl.opengvlab.com/)  [\[🤗 HF Demo\]](https://huggingface.co/spaces/OpenGVLab/InternVL)  [\[🚀 Quick Start\]](#quick-start)  [\[📖 Documents\]](https://internvl.readthedocs.io/en/latest/)22 23<div align="center">24  <img width="500" alt="image" src="https://cdn-uploads.huggingface.co/production/uploads/64006c09330a45b03605bba3/zJsd2hqd3EevgXo6fNgC-.png">25</div>26 27## Introduction28 29We are excited to announce the release of InternVL 2.0, the latest addition to the InternVL series of multimodal large language models. InternVL 2.0 features a variety of **instruction-tuned models**, ranging from 1 billion to 108 billion parameters. This repository contains the instruction-tuned InternVL2-1B model.30 31Compared to the state-of-the-art open-source multimodal large language models, InternVL 2.0 surpasses most open-source models. It demonstrates competitive performance on par with proprietary commercial models across various capabilities, including document and chart comprehension, infographics QA, scene text understanding and OCR tasks, scientific and mathematical problem solving, as well as cultural understanding and integrated multimodal capabilities.32 33InternVL 2.0 is trained with an 8k context window and utilizes training data consisting of long texts, multiple images, and videos, significantly improving its ability to handle these types of inputs compared to InternVL 1.5. For more details, please refer to our [blog](https://internvl.github.io/blog/2024-07-02-InternVL-2.0/) and [GitHub](https://github.com/OpenGVLab/InternVL).34 35|      Model Name      |                                     Vision Part                                     |                                        Language Part                                         |                             HF Link                              |                                MS Link                                 |36| :------------------: | :---------------------------------------------------------------------------------: | :------------------------------------------------------------------------------------------: | :--------------------------------------------------------------: | :--------------------------------------------------------------------: |37|     InternVL2-1B     |    [InternViT-300M-448px](https://huggingface.co/OpenGVLab/InternViT-300M-448px)    |            [Qwen2-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2-0.5B-Instruct)            |     [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-1B)     |     [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-1B)     |38|     InternVL2-2B     |    [InternViT-300M-448px](https://huggingface.co/OpenGVLab/InternViT-300M-448px)    |          [internlm2-chat-1_8b](https://huggingface.co/internlm/internlm2-chat-1_8b)          |     [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-2B)     |     [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-2B)     |39|     InternVL2-4B     |    [InternViT-300M-448px](https://huggingface.co/OpenGVLab/InternViT-300M-448px)    |    [Phi-3-mini-128k-instruct](https://huggingface.co/microsoft/Phi-3-mini-128k-instruct)     |     [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-4B)     |     [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-4B)     |40|     InternVL2-8B     |    [InternViT-300M-448px](https://huggingface.co/OpenGVLab/InternViT-300M-448px)    |          [internlm2_5-7b-chat](https://huggingface.co/internlm/internlm2_5-7b-chat)          |     [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-8B)     |     [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-8B)     |41|    InternVL2-26B     | [InternViT-6B-448px-V1-5](https://huggingface.co/OpenGVLab/InternViT-6B-448px-V1-5) |           [internlm2-chat-20b](https://huggingface.co/internlm/internlm2-chat-20b)           |    [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-26B)     |    [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-26B)     |42|    InternVL2-40B     | [InternViT-6B-448px-V1-5](https://huggingface.co/OpenGVLab/InternViT-6B-448px-V1-5) |       [Nous-Hermes-2-Yi-34B](https://huggingface.co/NousResearch/Nous-Hermes-2-Yi-34B)       |    [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-40B)     |    [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-40B)     |43| InternVL2-Llama3-76B | [InternViT-6B-448px-V1-5](https://huggingface.co/OpenGVLab/InternViT-6B-448px-V1-5) | [Hermes-2-Theta-Llama-3-70B](https://huggingface.co/NousResearch/Hermes-2-Theta-Llama-3-70B) | [🤗 link](https://huggingface.co/OpenGVLab/InternVL2-Llama3-76B) | [🤖 link](https://modelscope.cn/models/OpenGVLab/InternVL2-Llama3-76B) |44 45## Model Details46 47InternVL 2.0 is a multimodal large language model series, featuring models of various sizes. For each size, we release instruction-tuned models optimized for multimodal tasks. InternVL2-1B consists of [InternViT-300M-448px](https://huggingface.co/OpenGVLab/InternViT-300M-448px), an MLP projector, and [Qwen2-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2-0.5B-Instruct).48 49## Performance50 51### Image Benchmarks52 53|          Benchmark           | PaliGemma-3B | Mini-InternVL-2B-1-5 | InternVL2-1B |54| :--------------------------: | :----------: | :------------------: | :----------: |55|          Model Size          |     2.9B     |         2.2B         |     0.9B     |56|                              |              |                      |              |57|    DocVQA<sub>test</sub>     |      -       |         85.0         |     81.7     |58|    ChartQA<sub>test</sub>    |      -       |         74.8         |     72.9     |59|    InfoVQA<sub>test</sub>    |      -       |         55.4         |     50.9     |60|    TextVQA<sub>val</sub>     |     68.1     |         70.5         |     70.5     |61|           OCRBench           |     614      |         654          |     754      |62|      MME<sub>sum</sub>       |    1686.1    |        1901.5        |    1794.4    |63|         RealWorldQA          |     55.2     |         57.9         |     50.3     |64|     AI2D<sub>test</sub>      |     68.3     |         69.8         |     64.1     |65|      MMMU<sub>val</sub>      |     34.9     |         37.4         |     36.7     |66|  MMBench-EN<sub>test</sub>   |     71.0     |         70.9         |     65.4     |67|  MMBench-CN<sub>test</sub>   |     63.6     |         66.2         |     60.7     |68|    CCBench<sub>dev</sub>     |     29.6     |         63.5         |     75.7     |69|  MMVet<sub>GPT-4-0613</sub>  |      -       |         39.3         |     37.8     |70| MMVet<sub>GPT-4-Turbo</sub>  |     33.1     |         35.5         |     32.7     |71|          SEED-Image          |     69.6     |         69.8         |     65.6     |72|   HallBench<sub>avg</sub>    |     32.2     |         37.5         |     34.0     |73| MathVista<sub>testmini</sub> |     28.7     |         41.1         |     37.7     |74|  OpenCompass<sub>avg</sub>   |     46.6     |         49.8         |     48.3     |75 76- For more details and evaluation reproduction, please refer to our [Evaluation Guide](https://internvl.readthedocs.io/en/latest/internvl2.0/evaluation.html).77 78- We simultaneously use [InternVL](https://github.com/OpenGVLab/InternVL) and [VLMEvalKit](https://github.com/open-compass/VLMEvalKit) repositories for model evaluation. Specifically, the results reported for DocVQA, ChartQA, InfoVQA, TextVQA, MME, AI2D, MMBench, CCBench, MMVet (GPT-4-0613), and SEED-Image were tested using the InternVL repository. MMMU, OCRBench, RealWorldQA, HallBench, MMVet (GPT-4-Turbo), and MathVista were evaluated using the VLMEvalKit.79 80### Video Benchmarks81 82|          Benchmark          | VideoChat2-Phi3 | Mini-InternVL-2B-1-5 | InternVL2-1B |83| :-------------------------: | :-------------: | :------------------: | :----------: |84|         Model Size          |       4B        |         2.2B         |     0.9B     |85|                             |                 |                      |              |86|           MVBench           |      55.1       |         37.0         |     57.5     |87| MMBench-Video<sub>8f</sub>  |        -        |         0.99         |     0.95     |88| MMBench-Video<sub>16f</sub> |        -        |         1.04         |     0.98     |89|    Video-MME<br>w/o subs    |        -        |         42.9         |     42.6     |90|     Video-MME<br>w subs     |        -        |         44.7         |     44.7     |91 92- We evaluate our models on MVBench and Video-MME by extracting 16 frames from each video, and each frame was resized to a 448x448 image.93 94### Grounding Benchmarks95 96|             Model              | avg. | RefCOCO<br>(val) | RefCOCO<br>(testA) | RefCOCO<br>(testB) | RefCOCO+<br>(val) | RefCOCO+<br>(testA) | RefCOCO+<br>(testB) | RefCOCO‑g<br>(val) | RefCOCO‑g<br>(test) |97| :----------------------------: | :--: | :--------------: | :----------------: | :----------------: | :---------------: | :-----------------: | :-----------------: | :----------------: | :-----------------: |98| UNINEXT-H<br>(Specialist SOTA) | 88.9 |       92.6       |        94.3        |        91.5        |       85.2        |        89.6         |        79.8         |        88.7        |        89.4         |99|                                |      |                  |                    |                    |                   |                     |                     |                    |                     |100| Mini-InternVL-<br>Chat-2B-V1-5 | 75.8 |       80.7       |        86.7        |        72.9        |       72.5        |        82.3         |        60.8         |        75.6        |        74.9         |101| Mini-InternVL-<br>Chat-4B-V1-5 | 84.4 |       88.0       |        91.4        |        83.5        |       81.5        |        87.4         |        73.8         |        84.7        |        84.6         |102|       InternVL‑Chat‑V1‑5       | 88.8 |       91.4       |        93.7        |        87.1        |       87.0        |        92.3         |        80.9         |        88.5        |        89.3         |103|                                |      |                  |                    |                    |                   |                     |                     |                    |                     |104|          InternVL2‑1B          | 79.9 |       83.6       |        88.7        |        79.8        |       76.0        |        83.6         |        67.7         |        80.2        |        79.9         |105|          InternVL2‑2B          | 77.7 |       82.3       |        88.2        |        75.9        |       73.5        |        82.8         |        63.3         |        77.6        |        78.3         |106|          InternVL2‑4B          | 84.4 |       88.5       |        91.2        |        83.9        |       81.2        |        87.2         |        73.8         |        84.6        |        84.6         |107|          InternVL2‑8B          | 82.9 |       87.1       |        91.1        |        80.7        |       79.8        |        87.9         |        71.4         |        82.7        |        82.7         |108|         InternVL2‑26B          | 88.5 |       91.2       |        93.3        |        87.4        |       86.8        |        91.0         |        81.2         |        88.5        |        88.6         |109|         InternVL2‑40B          | 90.3 |       93.0       |        94.7        |        89.2        |       88.5        |        92.8         |        83.6         |        90.3        |        90.6         |110|    InternVL2-<br>Llama3‑76B    | 90.0 |       92.2       |        94.8        |        88.4        |       88.8        |        93.1         |        82.8         |        89.5        |        90.3         |111 112- We use the following prompt to evaluate InternVL's grounding ability: `Please provide the bounding box coordinates of the region this sentence describes: <ref>{}</ref>`113 114Limitations: Although we have made efforts to ensure the safety of the model during the training process and to encourage the model to generate text that complies with ethical and legal requirements, the model may still produce unexpected outputs due to its size and probabilistic generation paradigm. For example, the generated responses may contain biases, discrimination, or other harmful content. Please do not propagate such content. We are not responsible for any consequences resulting from the dissemination of harmful information.115 116## Quick Start117 118We provide an example code to run `InternVL2-1B` using `transformers`.119 120> Please use transformers>=4.37.2 to ensure the model works normally.121 122### Model Loading123 124#### 16-bit (bf16 / fp16)125 126```python127import torch128from transformers import AutoTokenizer, AutoModel129path = "OpenGVLab/InternVL2-1B"130model = AutoModel.from_pretrained(131    path,132    torch_dtype=torch.bfloat16,133    low_cpu_mem_usage=True,134    use_flash_attn=True,135    trust_remote_code=True).eval().cuda()136```137 138#### BNB 8-bit Quantization139 140```python141import torch142from transformers import AutoTokenizer, AutoModel143path = "OpenGVLab/InternVL2-1B"144model = AutoModel.from_pretrained(145    path,146    torch_dtype=torch.bfloat16,147    load_in_8bit=True,148    low_cpu_mem_usage=True,149    use_flash_attn=True,150    trust_remote_code=True).eval()151```152 153#### Multiple GPUs154 155The reason for writing the code this way is to avoid errors that occur during multi-GPU inference due to tensors not being on the same device. By ensuring that the first and last layers of the large language model (LLM) are on the same device, we prevent such errors.156 157```python158import math159import torch160from transformers import AutoTokenizer, AutoModel161 162def split_model(model_name):163    device_map = {}164    world_size = torch.cuda.device_count()165    num_layers = {166        'InternVL2-1B': 24, 'InternVL2-2B': 24, 'InternVL2-4B': 32, 'InternVL2-8B': 32,167        'InternVL2-26B': 48, 'InternVL2-40B': 60, 'InternVL2-Llama3-76B': 80}[model_name]168    # Since the first GPU will be used for ViT, treat it as half a GPU.169    num_layers_per_gpu = math.ceil(num_layers / (world_size - 0.5))170    num_layers_per_gpu = [num_layers_per_gpu] * world_size171    num_layers_per_gpu[0] = math.ceil(num_layers_per_gpu[0] * 0.5)172    layer_cnt = 0173    for i, num_layer in enumerate(num_layers_per_gpu):174        for j in range(num_layer):175            device_map[f'language_model.model.layers.{layer_cnt}'] = i176            layer_cnt += 1177    device_map['vision_model'] = 0178    device_map['mlp1'] = 0179    device_map['language_model.model.tok_embeddings'] = 0180    device_map['language_model.model.embed_tokens'] = 0181    device_map['language_model.output'] = 0182    device_map['language_model.model.norm'] = 0183    device_map['language_model.model.rotary_emb'] = 0184    device_map['language_model.lm_head'] = 0185    device_map[f'language_model.model.layers.{num_layers - 1}'] = 0186 187    return device_map188 189path = "OpenGVLab/InternVL2-1B"190device_map = split_model('InternVL2-1B')191model = AutoModel.from_pretrained(192    path,193    torch_dtype=torch.bfloat16,194    low_cpu_mem_usage=True,195    use_flash_attn=True,196    trust_remote_code=True,197    device_map=device_map).eval()198```199 200### Inference with Transformers201 202```python203import numpy as np204import torch205import torchvision.transforms as T206from decord import VideoReader, cpu207from PIL import Image208from torchvision.transforms.functional import InterpolationMode209from transformers import AutoModel, AutoTokenizer210 211IMAGENET_MEAN = (0.485, 0.456, 0.406)212IMAGENET_STD = (0.229, 0.224, 0.225)213 214def build_transform(input_size):215    MEAN, STD = IMAGENET_MEAN, IMAGENET_STD216    transform = T.Compose([217        T.Lambda(lambda img: img.convert('RGB') if img.mode != 'RGB' else img),218        T.Resize((input_size, input_size), interpolation=InterpolationMode.BICUBIC),219        T.ToTensor(),220        T.Normalize(mean=MEAN, std=STD)221    ])222    return transform223 224def find_closest_aspect_ratio(aspect_ratio, target_ratios, width, height, image_size):225    best_ratio_diff = float('inf')226    best_ratio = (1, 1)227    area = width * height228    for ratio in target_ratios:229        target_aspect_ratio = ratio[0] / ratio[1]230        ratio_diff = abs(aspect_ratio - target_aspect_ratio)231        if ratio_diff < best_ratio_diff:232            best_ratio_diff = ratio_diff233            best_ratio = ratio234        elif ratio_diff == best_ratio_diff:235            if area > 0.5 * image_size * image_size * ratio[0] * ratio[1]:236                best_ratio = ratio237    return best_ratio238 239def dynamic_preprocess(image, min_num=1, max_num=12, image_size=448, use_thumbnail=False):240    orig_width, orig_height = image.size241    aspect_ratio = orig_width / orig_height242 243    # calculate the existing image aspect ratio244    target_ratios = set(245        (i, j) for n in range(min_num, max_num + 1) for i in range(1, n + 1) for j in range(1, n + 1) if246        i * j <= max_num and i * j >= min_num)247    target_ratios = sorted(target_ratios, key=lambda x: x[0] * x[1])248 249    # find the closest aspect ratio to the target250    target_aspect_ratio = find_closest_aspect_ratio(251        aspect_ratio, target_ratios, orig_width, orig_height, image_size)252 253    # calculate the target width and height254    target_width = image_size * target_aspect_ratio[0]255    target_height = image_size * target_aspect_ratio[1]256    blocks = target_aspect_ratio[0] * target_aspect_ratio[1]257 258    # resize the image259    resized_img = image.resize((target_width, target_height))260    processed_images = []261    for i in range(blocks):262        box = (263            (i % (target_width // image_size)) * image_size,264            (i // (target_width // image_size)) * image_size,265            ((i % (target_width // image_size)) + 1) * image_size,266            ((i // (target_width // image_size)) + 1) * image_size267        )268        # split the image269        split_img = resized_img.crop(box)270        processed_images.append(split_img)271    assert len(processed_images) == blocks272    if use_thumbnail and len(processed_images) != 1:273        thumbnail_img = image.resize((image_size, image_size))274        processed_images.append(thumbnail_img)275    return processed_images276 277def load_image(image_file, input_size=448, max_num=12):278    image = Image.open(image_file).convert('RGB')279    transform = build_transform(input_size=input_size)280    images = dynamic_preprocess(image, image_size=input_size, use_thumbnail=True, max_num=max_num)281    pixel_values = [transform(image) for image in images]282    pixel_values = torch.stack(pixel_values)283    return pixel_values284 285# If you want to load a model using multiple GPUs, please refer to the `Multiple GPUs` section.286path = 'OpenGVLab/InternVL2-1B'287model = AutoModel.from_pretrained(288    path,289    torch_dtype=torch.bfloat16,290    low_cpu_mem_usage=True,291    use_flash_attn=True,292    trust_remote_code=True).eval().cuda()293tokenizer = AutoTokenizer.from_pretrained(path, trust_remote_code=True, use_fast=False)294 295# set the max number of tiles in `max_num`296pixel_values = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()297generation_config = dict(max_new_tokens=1024, do_sample=True)298 299# pure-text conversation (纯文本对话)300question = 'Hello, who are you?'301response, history = model.chat(tokenizer, None, question, generation_config, history=None, return_history=True)302print(f'User: {question}\nAssistant: {response}')303 304question = 'Can you tell me a story?'305response, history = model.chat(tokenizer, None, question, generation_config, history=history, return_history=True)306print(f'User: {question}\nAssistant: {response}')307 308# single-image single-round conversation (单图单轮对话)309question = '<image>\nPlease describe the image shortly.'310response = model.chat(tokenizer, pixel_values, question, generation_config)311print(f'User: {question}\nAssistant: {response}')312 313# single-image multi-round conversation (单图多轮对话)314question = '<image>\nPlease describe the image in detail.'315response, history = model.chat(tokenizer, pixel_values, question, generation_config, history=None, return_history=True)316print(f'User: {question}\nAssistant: {response}')317 318question = 'Please write a poem according to the image.'319response, history = model.chat(tokenizer, pixel_values, question, generation_config, history=history, return_history=True)320print(f'User: {question}\nAssistant: {response}')321 322# multi-image multi-round conversation, combined images (多图多轮对话,拼接图像)323pixel_values1 = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()324pixel_values2 = load_image('./examples/image2.jpg', max_num=12).to(torch.bfloat16).cuda()325pixel_values = torch.cat((pixel_values1, pixel_values2), dim=0)326 327question = '<image>\nDescribe the two images in detail.'328response, history = model.chat(tokenizer, pixel_values, question, generation_config,329                               history=None, return_history=True)330print(f'User: {question}\nAssistant: {response}')331 332question = 'What are the similarities and differences between these two images.'333response, history = model.chat(tokenizer, pixel_values, question, generation_config,334                               history=history, return_history=True)335print(f'User: {question}\nAssistant: {response}')336 337# multi-image multi-round conversation, separate images (多图多轮对话,独立图像)338pixel_values1 = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()339pixel_values2 = load_image('./examples/image2.jpg', max_num=12).to(torch.bfloat16).cuda()340pixel_values = torch.cat((pixel_values1, pixel_values2), dim=0)341num_patches_list = [pixel_values1.size(0), pixel_values2.size(0)]342 343question = 'Image-1: <image>\nImage-2: <image>\nDescribe the two images in detail.'344response, history = model.chat(tokenizer, pixel_values, question, generation_config,345                               num_patches_list=num_patches_list,346                               history=None, return_history=True)347print(f'User: {question}\nAssistant: {response}')348 349question = 'What are the similarities and differences between these two images.'350response, history = model.chat(tokenizer, pixel_values, question, generation_config,351                               num_patches_list=num_patches_list,352                               history=history, return_history=True)353print(f'User: {question}\nAssistant: {response}')354 355# batch inference, single image per sample (单图批处理)356pixel_values1 = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()357pixel_values2 = load_image('./examples/image2.jpg', max_num=12).to(torch.bfloat16).cuda()358num_patches_list = [pixel_values1.size(0), pixel_values2.size(0)]359pixel_values = torch.cat((pixel_values1, pixel_values2), dim=0)360 361questions = ['<image>\nDescribe the image in detail.'] * len(num_patches_list)362responses = model.batch_chat(tokenizer, pixel_values,363                             num_patches_list=num_patches_list,364                             questions=questions,365                             generation_config=generation_config)366for question, response in zip(questions, responses):367    print(f'User: {question}\nAssistant: {response}')368 369# video multi-round conversation (视频多轮对话)370def get_index(bound, fps, max_frame, first_idx=0, num_segments=32):371    if bound:372        start, end = bound[0], bound[1]373    else:374        start, end = -100000, 100000375    start_idx = max(first_idx, round(start * fps))376    end_idx = min(round(end * fps), max_frame)377    seg_size = float(end_idx - start_idx) / num_segments378    frame_indices = np.array([379        int(start_idx + (seg_size / 2) + np.round(seg_size * idx))380        for idx in range(num_segments)381    ])382    return frame_indices383 384def load_video(video_path, bound=None, input_size=448, max_num=1, num_segments=32):385    vr = VideoReader(video_path, ctx=cpu(0), num_threads=1)386    max_frame = len(vr) - 1387    fps = float(vr.get_avg_fps())388 389    pixel_values_list, num_patches_list = [], []390    transform = build_transform(input_size=input_size)391    frame_indices = get_index(bound, fps, max_frame, first_idx=0, num_segments=num_segments)392    for frame_index in frame_indices:393        img = Image.fromarray(vr[frame_index].asnumpy()).convert('RGB')394        img = dynamic_preprocess(img, image_size=input_size, use_thumbnail=True, max_num=max_num)395        pixel_values = [transform(tile) for tile in img]396        pixel_values = torch.stack(pixel_values)397        num_patches_list.append(pixel_values.shape[0])398        pixel_values_list.append(pixel_values)399    pixel_values = torch.cat(pixel_values_list)400    return pixel_values, num_patches_list401 402video_path = './examples/red-panda.mp4'403pixel_values, num_patches_list = load_video(video_path, num_segments=8, max_num=1)404pixel_values = pixel_values.to(torch.bfloat16).cuda()405video_prefix = ''.join([f'Frame{i+1}: <image>\n' for i in range(len(num_patches_list))])406question = video_prefix + 'What is the red panda doing?'407# Frame1: <image>\nFrame2: <image>\n...\nFrame8: <image>\n{question}408response, history = model.chat(tokenizer, pixel_values, question, generation_config,409                               num_patches_list=num_patches_list, history=None, return_history=True)410print(f'User: {question}\nAssistant: {response}')411 412question = 'Describe this video in detail.'413response, history = model.chat(tokenizer, pixel_values, question, generation_config,414                               num_patches_list=num_patches_list, history=history, return_history=True)415print(f'User: {question}\nAssistant: {response}')416```417 418#### Streaming Output419 420Besides this method, you can also use the following code to get streamed output.421 422```python423from transformers import TextIteratorStreamer424from threading import Thread425 426# Initialize the streamer427streamer = TextIteratorStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True, timeout=10)428# Define the generation configuration429generation_config = dict(max_new_tokens=1024, do_sample=False, streamer=streamer)430# Start the model chat in a separate thread431thread = Thread(target=model.chat, kwargs=dict(432    tokenizer=tokenizer, pixel_values=pixel_values, question=question,433    history=None, return_history=False, generation_config=generation_config,434))435thread.start()436 437# Initialize an empty string to store the generated text438generated_text = ''439# Loop through the streamer to get the new text as it is generated440for new_text in streamer:441    if new_text == model.conv_template.sep:442        break443    generated_text += new_text444    print(new_text, end='', flush=True)  # Print each new chunk of generated text on the same line445```446 447## Finetune448 449Many repositories now support fine-tuning of the InternVL series models, including [InternVL](https://github.com/OpenGVLab/InternVL), [SWIFT](https://github.com/modelscope/ms-swift), [XTurner](https://github.com/InternLM/xtuner), and others. Please refer to their documentation for more details on fine-tuning.450 451## Deployment452 453### LMDeploy454 455LMDeploy is a toolkit for compressing, deploying, and serving LLMs & VLMs.456 457```sh458pip install lmdeploy>=0.5.3459```460 461LMDeploy abstracts the complex inference process of multi-modal Vision-Language Models (VLM) into an easy-to-use pipeline, similar to the Large Language Model (LLM) inference pipeline.462 463#### A 'Hello, world' Example464 465```python466from lmdeploy import pipeline, TurbomindEngineConfig467from lmdeploy.vl import load_image468 469model = 'OpenGVLab/InternVL2-1B'470image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/tests/data/tiger.jpeg')471pipe = pipeline(model, backend_config=TurbomindEngineConfig(session_len=8192))472response = pipe(('describe this image', image))473print(response.text)474```475 476If `ImportError` occurs while executing this case, please install the required dependency packages as prompted.477 478#### Multi-images Inference479 480When dealing with multiple images, you can put them all in one list. Keep in mind that multiple images will lead to a higher number of input tokens, and as a result, the size of the context window typically needs to be increased.481 482> Warning: Due to the scarcity of multi-image conversation data, the performance on multi-image tasks may be unstable, and it may require multiple attempts to achieve satisfactory results.483 484```python485from lmdeploy import pipeline, TurbomindEngineConfig486from lmdeploy.vl import load_image487from lmdeploy.vl.constants import IMAGE_TOKEN488 489model = 'OpenGVLab/InternVL2-1B'490pipe = pipeline(model, backend_config=TurbomindEngineConfig(session_len=8192))491 492image_urls=[493    'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg',494    'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/det.jpg'495]496 497images = [load_image(img_url) for img_url in image_urls]498# Numbering images improves multi-image conversations499response = pipe((f'Image-1: {IMAGE_TOKEN}\nImage-2: {IMAGE_TOKEN}\ndescribe these two images', images))500print(response.text)501```502 503#### Batch Prompts Inference504 505Conducting inference with batch prompts is quite straightforward; just place them within a list structure:506 507```python508from lmdeploy import pipeline, TurbomindEngineConfig509from lmdeploy.vl import load_image510 511model = 'OpenGVLab/InternVL2-1B'512pipe = pipeline(model, backend_config=TurbomindEngineConfig(session_len=8192))513 514image_urls=[515    "https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg",516    "https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/det.jpg"517]518prompts = [('describe this image', load_image(img_url)) for img_url in image_urls]519response = pipe(prompts)520print(response)521```522 523#### Multi-turn Conversation524 525There are two ways to do the multi-turn conversations with the pipeline. One is to construct messages according to the format of OpenAI and use above introduced method, the other is to use the `pipeline.chat` interface.526 527```python528from lmdeploy import pipeline, TurbomindEngineConfig, GenerationConfig529from lmdeploy.vl import load_image530 531model = 'OpenGVLab/InternVL2-1B'532pipe = pipeline(model, backend_config=TurbomindEngineConfig(session_len=8192))533 534image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg')535gen_config = GenerationConfig(top_k=40, top_p=0.8, temperature=0.8)536sess = pipe.chat(('describe this image', image), gen_config=gen_config)537print(sess.response.text)538sess = pipe.chat('What is the woman doing?', session=sess, gen_config=gen_config)539print(sess.response.text)540```541 542#### Service543 544LMDeploy's `api_server` enables models to be easily packed into services with a single command. The provided RESTful APIs are compatible with OpenAI's interfaces. Below are an example of service startup:545 546```shell547lmdeploy serve api_server OpenGVLab/InternVL2-1B --server-port 23333548```549 550To use the OpenAI-style interface, you need to install OpenAI:551 552```shell553pip install openai554```555 556Then, use the code below to make the API call:557 558```python559from openai import OpenAI560 561client = OpenAI(api_key='YOUR_API_KEY', base_url='http://0.0.0.0:23333/v1')562model_name = client.models.list().data[0].id563response = client.chat.completions.create(564    model=model_name,565    messages=[{566        'role':567        'user',568        'content': [{569            'type': 'text',570            'text': 'describe this image',571        }, {572            'type': 'image_url',573            'image_url': {574                'url':575                'https://modelscope.oss-cn-beijing.aliyuncs.com/resource/tiger.jpeg',576            },577        }],578    }],579    temperature=0.8,580    top_p=0.8)581print(response)582```583 584## License585 586This project is released under the MIT License. This project uses the pre-trained Qwen2-0.5B-Instruct as a component, which is licensed under the Apache License 2.0.587 588## Citation589 590If you find this project useful in your research, please consider citing:591 592```BibTeX593@article{chen2024expanding,594  title={Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling},595  author={Chen, Zhe and Wang, Weiyun and Cao, Yue and Liu, Yangzhou and Gao, Zhangwei and Cui, Erfei and Zhu, Jinguo and Ye, Shenglong and Tian, Hao and Liu, Zhaoyang and others},596  journal={arXiv preprint arXiv:2412.05271},597  year={2024}598}599@article{gao2024mini,600  title={Mini-internvl: A flexible-transfer pocket multimodal model with 5\% parameters and 90\% performance},601  author={Gao, Zhangwei and Chen, Zhe and Cui, Erfei and Ren, Yiming and Wang, Weiyun and Zhu, Jinguo and Tian, Hao and Ye, Shenglong and He, Junjun and Zhu, Xizhou and others},602  journal={arXiv preprint arXiv:2410.16261},603  year={2024}604}605@article{chen2024far,606  title={How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites},607  author={Chen, Zhe and Wang, Weiyun and Tian, Hao and Ye, Shenglong and Gao, Zhangwei and Cui, Erfei and Tong, Wenwen and Hu, Kongzhi and Luo, Jiapeng and Ma, Zheng and others},608  journal={arXiv preprint arXiv:2404.16821},609  year={2024}610}611@inproceedings{chen2024internvl,612  title={Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks},613  author={Chen, Zhe and Wu, Jiannan and Wang, Wenhai and Su, Weijie and Chen, Guo and Xing, Sen and Zhong, Muyan and Zhang, Qinglong and Zhu, Xizhou and Lu, Lewei and others},614  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},615  pages={24185--24198},616  year={2024}617}618```619