CoolFace
Modelpublic

OpenGVLab/InternVL3_5-8B-Flash

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
8likes2.1kdownloads
README.md832 linesDownload Raw Back to root
1---2license: apache-2.03pipeline_tag: image-text-to-text4library_name: transformers5base_model:6  - OpenGVLab/InternVL3_5-8B7base_model_relation: finetune8datasets:9  - OpenGVLab/MMPR-v1.210  - OpenGVLab/MMPR-Tiny11language:12  - multilingual13tags:14  - internvl15  - custom_code16---17 18# InternVL3_5-8B-Flash19 20[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL)  [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238)  [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821)  [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271)  [\[📜 InternVL2.5-MPO\]](https://huggingface.co/papers/2411.10442)  [\[📜 InternVL3\]](https://huggingface.co/papers/2504.10479) [\[📜 InternVL3.5\]](https://huggingface.co/papers/2508.18265)21 22[\[🆕 Blog\]](https://internvl.github.io/blog/)  [\[🗨️ Chat Demo\]](https://chat.intern-ai.org.cn/)  [\[🚀 Quick Start\]](#quick-start)  [\[📖 Documents\]](https://internvl.readthedocs.io/en/latest/)23 24<div align="center">25  <img width="500" alt="image" src="https://cdn-uploads.huggingface.co/production/uploads/64006c09330a45b03605bba3/zJsd2hqd3EevgXo6fNgC-.png">26</div>27 28## Introduction29 30We introduce *InternVL3.5*, a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL series. A key innovation is the *Cascade Reinforcement Learning (Cascade RL)* framework, which enhances reasoning through a two-stage process: offline RL for stable convergence and online RL for refined alignment. This coarse-to-fine training strategy leads to substantial improvements on downstream reasoning tasks, e.g., MMMU and MathVista. To optimize efficiency, we propose a *Visual Resolution Router (ViR)* that dynamically adjusts the resolution of visual tokens without compromising performance. Coupled with ViR, our Decoupled *Vision-Language Deployment (DvD)* strategy separates the vision encoder and language model across different GPUs, effectively balancing computational load. These contributions collectively enable InternVL3.5 to achieve up to a +16.0\% gain in overall reasoning performance and a 4.05 \\(\times\\) inference speedup compared to its predecessor, i.e., InternVL3. In addition, InternVL3.5 supports novel capabilities such as GUI interaction and embodied agency. Notably, our largest model, i.e.,  InternVL3.5-241B-A28B, attains state-of-the-art results among open-source MLLMs across general multimodal, reasoning, text, and agentic tasks—narrowing the performance gap with leading commercial models like GPT-5. All models and code are publicly released.31 32![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance.jpg)33 34> Hatched bars represent closed-source commercial models. We report average scores on a set of multimodal general, reasoning, text, and agentic benchmarks: MMBench v1.1 (en), MMStar,BLINK, HallusionBench, AI2D, OCRBench, MMVet, MME-RealWorld (en), MVBench, VideoMME, MMMU, MathVista, MathVision, MathVerse, DynaMath, WeMath, LogicVista, MATH500, AIME24, AIME25, GPQA, MMLU-Pro, GAOKAO, IFEval, SGP-Bench, VSI-Bench, ERQA, SpaCE-10, and OmniSpatial.35 36See [quick start](#quick-start) for how to use our model.37 38## InternVL3.5 Family39 40In the following table, we provide an overview of the InternVL3.5 series.41To maintain consistency with earlier generations, we provide two model formats: [the GitHub format](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B), consistent with prior releases, and [the HF format](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B-HF), aligned with the official Transformers standard.42 43> If you want to convert the checkpoint between these two formats, please refer to the scripts about [custom2hf](https://github.com/OpenGVLab/InternVL/blob/main/internvl_chat/tools/internvl_custom2hf.py) and [hf2custom](https://github.com/OpenGVLab/InternVL/blob/main/internvl_chat/tools/internvl_hf2custom.py).44 45 46### Github Format47 48 49| Model                 | #Vision Param | #Language Param | #Total Param | HF Link                                                                        | ModelScope Link                                                                          |50| --------------------- | ------------- | --------------- | ------------ | ------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------- |51| InternVL3.5-1B        | 0.3B          | 0.8B            | 1.1B         | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-1B)                      | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-1B)                      |52| InternVL3.5-2B        | 0.3B          | 2.0B            | 2.3B         | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-2B)                      | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-2B)                      |53| InternVL3.5-4B        | 0.3B          | 4.4B            | 4.7B         | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-4B)                      | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-4B)                      |54| InternVL3.5-8B        | 0.3B          | 8.2B            | 8.5B         | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-8B)                      | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-8B)                      |55| InternVL3.5-14B       | 0.3B          | 14.8B           | 15.1B        | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-14B)                     | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-14B)                     |56| InternVL3.5-38B       | 5.5B          | 32.8B           | 38.4B        | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-38B)                     | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-38B)                     |57| InternVL3.5-20B-A4B   | 0.3B          | 20.9B           | 21.2B-A4B    | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-GPT-OSS-20B-A4B-Preview) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-GPT-OSS-20B-A4B-Preview) |58| InternVL3.5-30B-A3B   | 0.3B          | 30.5B           | 30.8B-A3B    | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-30B-A3B)                 | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-30B-A3B)                 |59| InternVL3.5-241B-A28B | 5.5B          | 235.1B          | 240.7B-A28B  | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B)               | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-241B-A28B)               |60 61 62### HuggingFace Format63 64 65| Model                    | #Vision Param | #Language Param | #Total Param | HF Link                                                                           | ModelScope Link                                                                             |66| ------------------------ | ------------- | --------------- | ------------ | --------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- |67| InternVL3.5-1B-HF        | 0.3B          | 0.8B            | 1.1B         | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-1B-HF)                      | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-1B-HF)                      |68| InternVL3.5-2B-HF        | 0.3B          | 2.0B            | 2.3B         | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-2B-HF)                      | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-2B-HF)                      |69| InternVL3.5-4B-HF        | 0.3B          | 4.4B            | 4.7B         | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-4B-HF)                      | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-4B-HF)                      |70| InternVL3.5-8B-HF        | 0.3B          | 8.2B            | 8.5B         | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-8B-HF)                      | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-8B-HF)                      |71| InternVL3.5-14B-HF       | 0.3B          | 14.8B           | 15.1B        | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-14B-HF)                     | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-14B-HF)                     |72| InternVL3.5-38B-HF       | 5.5B          | 32.8B           | 38.4B        | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-38B-HF)                     | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-38B-HF)                     |73| InternVL3.5-20B-A4B-HF   | 0.3B          | 20.9B           | 21.2B-A4B    | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-GPT-OSS-20B-A4B-Preview-HF) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-GPT-OSS-20B-A4B-Preview-HF) |74| InternVL3.5-30B-A3B-HF   | 0.3B          | 30.5B           | 30.8B-A3B    | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-30B-A3B-HF)                 | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-30B-A3B-HF)                 |75| InternVL3.5-241B-A28B-HF | 5.5B          | 235.1B          | 240.7B-A28B  | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B-HF)               | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-241B-A28B-HF)               |76 77 78![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_overall.jpg)79 80> We conduct the evaluation with [VLMEvalkit](https://github.com/open-compass/VLMEvalKit). ***To enable the Thinking mode of our model, please set the system prompt to [R1_SYSTEM_PROMPT](https://github.com/open-compass/VLMEvalKit/blob/main/vlmeval/vlm/internvl/internvl_chat.py#L38).*** When enabling Thinking mode, we recommend setting `do_sample=True` and `temperature=0.6` to mitigate undesired repetition.81 82Our training pipeline comprises four stages: Multimodal Continual Pre-Training (**CPT**), Supervised Fine-Tuning (**SFT**), and Cascade Reinforcement Learning (**CascadeRL**). In CascadeRL, we first fine-tune the model using Mixed Preference Optimization (**MPO**) under an offline RL setting, followed by **GSPO** under an oneline RL setting.83For the Flash version of InternVL3.5, we additionally introduce a lightweight training stage, termed Visual Consistency Learning (**ViCO**), which reduces the token cost required to represent an image patch.84 85![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/training_pipeline.jpg)86 87Here, we also open-source the model weights after different training stages for potential research usage.88***If you're unsure which version to use, please select the one without any suffix, as it has completed the full training pipeline.***89 90 91| Model                            | Training Pipeline     | HF Link                                                                     | ModelScope Link                                                                       |92| -------------------------------- | --------------------- | --------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- |93| InternVL3.5-1B-Pretrained        | CPT                   | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-1B-Pretrained)        | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-1B-Pretrained)        |94| InternVL3.5-1B-Instruct          | CPT + SFT             | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-1B-Instruct)          | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-1B-Instruct)          |95| InternVL3.5-1B-MPO               | CPT + SFT + MPO       | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-1B-MPO)               | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-1B-MPO)               |96| InternVL3.5-1B                   | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-1B)                   | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-1B)                   |97| InternVL3.5-2B-Pretrained        | CPT                   | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-2B-Pretrained)        | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-2B-Pretrained)        |98| InternVL3.5-2B-Instruct          | CPT + SFT             | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-2B-Instruct)          | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-2B-Instruct)          |99| InternVL3.5-2B-MPO               | CPT + SFT + MPO       | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-2B-MPO)               | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-2B-MPO)               |100| InternVL3.5-2B                   | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-2B)                   | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-2B)                   |101| InternVL3.5-4B-Pretrained        | CPT                   | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-4B-Pretrained)        | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-4B-Pretrained)        |102| InternVL3.5-4B-Instruct          | CPT + SFT             | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-4B-Instruct)          | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-4B-Instruct)          |103| InternVL3.5-4B-MPO               | CPT + SFT + MPO       | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-4B-MPO)               | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-4B-MPO)               |104| InternVL3.5-4B                   | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-4B)                   | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-4B)                   |105| InternVL3.5-8B-Pretrained        | CPT                   | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-8B-Pretrained)        | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-8B-Pretrained)        |106| InternVL3.5-8B-Instruct          | CPT + SFT             | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-8B-Instruct)          | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-8B-Instruct)          |107| InternVL3.5-8B-MPO               | CPT + SFT + MPO       | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-8B-MPO)               | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-8B-MPO)               |108| InternVL3.5-8B                   | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-8B)                   | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-8B)                   |109| InternVL3.5-14B-Pretrained       | CPT                   | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-14B-Pretrained)       | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-14B-Pretrained)       |110| InternVL3.5-14B-Instruct         | CPT + SFT             | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-14B-Instruct)         | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-14B-Instruct)         |111| InternVL3.5-14B-MPO              | CPT + SFT + MPO       | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-14B-MPO)              | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-14B-MPO)              |112| InternVL3.5-14B                  | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-14B)                  | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-14B)                  |113| InternVL3.5-30B-A3B-Pretrained   | CPT                   | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-30B-A3B-Pretrained)   | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-30B-A3B-Pretrained)   |114| InternVL3.5-30B-A3B-Instruct     | CPT + SFT             | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-30B-A3B-Instruct)     | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-30B-A3B-Instruct)     |115| InternVL3.5-30B-A3B-MPO          | CPT + SFT + MPO       | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-30B-A3B-MPO)          | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-30B-A3B-MPO)          |116| InternVL3.5-30B-A3B              | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-30B-A3B)              | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-30B-A3B)              |117| InternVL3.5-38B-Pretrained       | CPT                   | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-38B-Pretrained)       | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-38B-Pretrained)       |118| InternVL3.5-38B-Instruct         | CPT + SFT             | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-38B-Instruct)         | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-38B-Instruct)         |119| InternVL3.5-38B-MPO              | CPT + SFT + MPO       | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-38B-MPO)              | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-38B-MPO)              |120| InternVL3.5-38B                  | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-38B)                  | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-38B)                  |121| InternVL3.5-241B-A28B-Pretrained | CPT                   | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B-Pretrained) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-241B-A28B-Pretrained) |122| InternVL3.5-241B-A28B-Instruct   | CPT + SFT             | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B-Instruct)   | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-241B-A28B-Instruct)   |123| InternVL3.5-241B-A28B-MPO        | CPT + SFT + MPO       | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B-MPO)        | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-241B-A28B-MPO)        |124| InternVL3.5-241B-A28B            | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B)            | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-241B-A28B)            |125 126 127The Flash version of our model will be released as soon as possible.128 129 130 131## Model Architecture132 133`InternVL3.5`:134This series of models follow the "ViT–MLP–LLM" paradigm adopted in previous versions of InternVL.135We initialize the language model using the Qwen3 series and GPT-OSS, and the vision encoder using InternViT-300M and InternViT-6B.136The Dynamic High Resolution strategy introduced in InternVL1.5 is also retained in our design.137 138 139`InternVL3.5-Flash`:140Compared to InternVL3.5, InternVL3.5-Flash further integrates the *Visual Resolution Router (ViR)*, thus yielding a series of  efficient variants friendly  suitable for  resource-constrained scenarios. 141Specifically, in InternVL3.5, each image patch is initially represented as 1024 visual tokens for the vision encoder, which are then compressed into 256 tokens via a pixel shuffle module before being passed to the Large Language Model (LLM).142In InternVL3.5-Flash, as shown in the Figure below, an additional pixel shuffle module with a higher compression rate is included, enabling the compression of visual tokens down to 64 tokens.143For each patch, the patch router determines the appropriate compression rate by assessing its semantic richness, and routes it to the corresponding pixel shuffle module accordingly.144Benefiting from this patch-aware compression mechanism, InternVL3.5-Flash is able to reduce the number of visual tokens by 50\% while maintaining nearly 100\% of the performance of InternVL3.5.145 146 147![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/architecture.jpg)148 149## Training and Deployment Strategy150 151### Pre-Training152 153During the pre-training stage, we update all model parameters jointly using the combination of large-scale text and multimodal corpora. Specifically, given an arbitrary training sample consisting of a multimodal token sequence \\(\mathbf{x}=\left(x_1, x_2, \ldots, x_L\right)\\), the next token prediction (NTP) loss is calculated on each text token as follows:154 155$$156    \mathcal{L}_{i}=-\log p_\theta\left(x_i \mid x_1, \ldots, x_{i-1}\right),157$$158 159where \\(x_i\\) is the predicted token and  prefix tokens in \\(\{x_1, x_2, \ldots, x_{i-1}\}\\) can be either  text tokens or  image tokens. Notably, for conversation samples, only response tokens  are included for the calculation of the loss.160Additionally, to mitigate bias toward either longer or shorter responses during training, we adopt the square averaging to re-weight the NTP loss  as follows:161 162$$163\mathcal{L}_{i}^{'} = \frac{w_i}{\sum_j w_j} \cdot \mathcal{L}_i, \quad w_i = \frac{1}{N^{0.5}},164$$165 166where \\(N\\) denotes the number of tokens in the training sample on which the loss needs to be calculated. The random JPEG compression is also included to enhance the model's real-world performance.167 168### Supervised Fine-Tuning169 170During the SFT phase, we adopt the same objective as in the pre-training stage and use the  square-root averaging strategy to calculate the final loss.  In this stage, the context window is set to 32K tokens to adapt long-context information.171Compared to InternVL3, the SFT stage of InternVL3.5 contains  more high-quality and  diverse training data derived from three sources: 172 173(1) Instruction-following data from InternVL3, which are reused to preserve broad coverage of vision–language tasks. 174 175(2) Multimodal reasoning data in the "Thinking" mode, which are included to instill long-thinking capabilities in the model. To construct such data, we first use InternVL3-78B to describe the image and then input the description into DeepSeek-R1 to sample rollouts with detailed reasoning processes. Rollouts with an incorrect final answer are filtered out. The questions in these datasets cover various expert domains, such as mathematics and scientific disciplines, thereby strengthening performance on different reasoning tasks. 176 177(3) Capability-expansion datasets, which endow InternVL3.5 with new skills, including GUI-based interaction, embodied interaction, and scalable vect178 179### Cascade Reinforcement Learning180 181Cascade RL aims to combine the benefits of offline RL and online RL to progressively facilitate the post-training of MLLMs in an efficient manner.182Specifically, we first fine-tune the model using an offline RL algorithm as an efficient warm-up stage to reach a satisfied results, which can guarantee the high-quality rollouts for the latter stage. 183Subsequently, we employ an online RL algorithm to further refine the output distribution based on rollouts generated by the model itself.  Compared to the single offline or online RL stage, our cascaded RL achieves significant performance improvements at a fraction of the GPU time cost.184 185 186 187During the offline RL stage, we employ mixed preference optimization (MPO) to fine-tune the model. Specifically, the training objective of MPO is a combination of preference loss \\(\mathcal{L}_{p}\\), quality loss \\(\mathcal{L}_{q}\\), and generation loss \\(\mathcal{L}_{g}\\), which can be formulated as follows:188 189$$190    \mathcal{L}_{\text{MPO}}=191    w_{p} \mathcal{L}_{p}192    +193    w_{q} \mathcal{L}_{q}194    +195    w_{g} \mathcal{L}_{g}196    ,197$$198 199where \\(w_{*}\\) represents the weight assigned to each loss component.200The DPO loss, BCO loss, and LM loss serve as the preference loss, quality loss, and generation loss, respectively.201 202 203During the online RL stage, we employ GSPO, without reference model constraints, as our online RL algorithm, which we find more effective in training both dense and mixture-of-experts (MoE) models. Similar to GRPO, the advantage is defined as the normalized reward across responses sampled from the same query.204The training objective of GSPO is given by:205 206$$207    \mathcal{L}_{\mathrm{GSPO}}(\theta)=\mathbb{E}_{x \sim \mathcal{D},\left\{y_i\right\}_{i=1}^G \sim \pi_{\theta \text { old }}(\cdot \mid x)}\left[\frac{1}{G} \sum_{i=1}^G \min \left(s_i(\theta) \widehat{A}_i, \operatorname{clip}\left(s_i(\theta), 1-\varepsilon, 1+\varepsilon\right) \widehat{A}_i\right)\right],208$$209 210where the importance sampling ratio is defined as the geometric mean of the per-token ratios.211 212> Please see [our paper](https://huggingface.co/papers/2508.18265) for more technical and experimental details.213 214 215### Visual Consistency Learning216 217 218We further include ViCO as an additional training stage to integrate the *visual resolution router (ViR)* into InternVL3.5, thereby reducing the inference cost of InternVL3.5. The obtained efficient version of InternVL3.5 are termed as *InternVL3.5-Flash*. In particular, ViCO comprises two stages:219 220`Consistency training`:221In this stage, the entire model is trained to minimize the divergence between response distributions conditioned on visual tokens with different compression rates.222In practice, we introduce an extra reference model, which is frozen and initialized with InternVL3.5.223Given a sample, each image patch is represented as either 256 or 64 tokens, and the training objective is defined as follows:224 225 226$$227\mathcal{L}_\text{ViCO} =228\mathbb{E}_{\xi \sim \mathcal{R}} \Bigg[229\frac{1}{N} \sum_{i=1}^{N} \mathrm{KL} \Big(230\pi_{\theta_{ref}}\left(y_i \mid y_{<i}, I\right) \;\Big\|\;231\pi_{\theta_{policy}}\left(y_i \mid y_{<i}, I_\xi\right)232\Big)233\Bigg],234$$235 236where \\(\mathrm{KL}\) denotes the KL divergence and \(\xi\) denotes the compression rate, which is uniformly sampled from \(\{\frac{1}{4},\frac{1}{16}\}\). The image \(I_\xi\) is represented as 256 tokens when \(\xi=\frac{1}{4}\) and 64 tokens when \(\xi=\frac{1}{16}\). Notably, the reference model always performs inference with \(\xi=\frac{1}{4}\).237 238 239`Router training`:240This stage aims to train the ViR to select an appropriate trade-off resolution for different inputs.241ViR is formulated as a binary classifier and trained using standard cross-entropy loss.242To construct the route targets, we first compute the KL divergence between the model outputs conditioned on uncompressed visual tokens (i.e., 256 tokens per patch) and those conditioned on compressed visual tokens (i.e., 64 tokens per patch).243During this stage, the main MLLM (ViT, MLP and LLM) is kept frozen, and only the ViR is trained.244Specifically, we first compute the loss ratio for each patch:245 246$$247r_i = \frac{\mathcal{L}_\text{ViCO}\big(y_i \mid I_{\frac{1}{16}}\big)}{\mathcal{L}_\text{ViCO}\big(y_i \mid I_{\frac{1}{4}}\big)},248$$249 250which quantifies the relative increase in loss caused by compressing the visual tokens. Based on this ratio, the binary ground-truth label for the patch router is defined as:251 252$$253y_i^\text{router} =254\begin{cases}2550, & r_i < \tau \; \text{(compression has negligible impact)} \\2561, & r_i \ge \tau \; \text{(compression has significant impact)},257\end{cases}258$$259 260where \(y_i^{\text{router}}=0\) and \(y_i^{\text{router}}=1\)  indicate that the compression rate \(\xi\) is set to \(\tfrac{1}{16}\) and \(\tfrac{1}{4}\), respectively.261 262> Please see [our paper](https://huggingface.co/papers/2508.18265) for more technical and experimental details.263 264 265### Test-Time Scaling266 267 268Test-time scaling (TTS) has been empirically demonstrated as an effective approach to enhance the reasoning capabilities of LLMs and MLLMs, particularly for complex tasks necessitating multi-step inference.269In this work, we implement a comprehensive test-time scaling approach that simultaneously improves reasoning depth (i.e., deep thinking) and breadth (i.e., parallel thinking).270 271`Deep Thinking`: By activating the Thinking mode, we guide the model to deliberately engage in step-by-step reasoning (i.e., decomposing complex problems into logical steps and validating intermediate conclusions) prior to generating the final answer. This approach systematically improves the logical structure of solutions for complex problems, particularly those requiring multi-step inference, and enhances reasoning depth.272 273`Parallel Thinking`: Following InternVL3, for reasoning tasks, we adopt the Best-of-N (BoN) strategy by employing [VisualPRM-v1.1](https://huggingface.co/OpenGVLab/VisualPRM-8B-v1_1) as the critic model to select the optimal response from multiple reasoning candidates.274This approach improves reasoning breadth.275 276> Notably, unless otherwise specified, the experimental results reported in our paper are obtained without applying TTS. Thus far, we have only applied TTS to reasoning benchmarks, since we found that the model already exhibits strong perception and understanding capabilities, and initiating TTS yields no significant improvement.277 278 279### Decoupled Vision-Language Deployment280 281In multimodal inference, the vision encoder and language model have distinct computational characteristics. The vision encoder that transforms images into semantic features is highly parallelizable and does not rely on long-term history state.  In contrast,  the language model adopts the inference in an autoregressive manner, which requires previous states to compute the next one. This sequential property makes the language part more sensitive to memory bandwidth and latency. 282When MLLMs are deployed online at scale, the vision and language models often block each other, thus incurring additional inference cost. This effect becomes more pronounced with larger vision models or higher-resolution images.283 284![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/DvD.jpg)285 286As shown in the Figure above, we propose decoupled vision-language deployment (DvD) to address this issue by separating vision and language processing, with a particular focus on optimizing the prefilling stage. The vision subsystem batches and processes images to produce compact feature embeddings, which are then transmitted to the language subsystem for fusion with the text context prior to decoding. This separation alleviates blocking and brings multimodal prefilling performance closer to that of pure language models.287In our system implementation, the ViT and MLP (and ViR for InternVL3.5-Flash) are deployed on the vision server, while the language server executes only the LLM. The communication is unidirectional, transmitting BF16 visual features over TCP, with RDMA optionally employed to achieve higher transmission speed. Vision processing, feature transmission, and language processing are organized into an asynchronous three-stage pipeline, enabling overlapped execution and minimizing pipeline stalls.288 289 290DvD increases GPU utilization and processing efficiency on the vision side, while enabling the language server to focus exclusively on the LLM’s prefilling and decoding without being blocked by vision computation. This design leads to improved throughput and responsiveness. Moreover, the architecture supports independent hardware cost optimization for the vision and language modules, and facilitates the seamless integration of new modules without requiring modifications to the language server deployment.291 292 293## Evaluation on Multimodal Capability294 295### Multimodal Reasoning and Mathematics296 297![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_reasoning.jpg)298 299### OCR, Chart, and Document Understanding300 301![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_ocr.jpg)302 303### Multi-Image Understanding & Real-World Comprehension304 305![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_multi_images.jpg)306 307### Comprehensive Multimodal Understanding & Multimodal Hallucination Evaluation308 309![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_comprehensive.jpg)310 311### Visual Grounding312 313![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_grounding.jpg)314 315### Multimodal Multilingual Understanding316 317![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_multilingual.jpg)318 319### Video Understanding320 321![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_video.jpg)322 323### GUI Tasks324 325![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_gui.jpg)326 327### Embodied Tasks328 329![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_embody.jpg)330 331### SVG Tasks332 333![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_svg.jpg)334 335![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_svg_gen.jpg)336 337## Evaluation on Language Capability338 339![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/performance_text.jpg)340 341## Ablation Study342 343### Cascade Reinforcement Learning344 345![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/ablation_cascade_rl.jpg)346 347![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/ablation_cascade_rl_table.jpg)348 349### Decoupled Vision-Language Deployment350 351 352![image/jpg](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B/resolve/main/images/ablation_dvd.jpg)353 354## Quick Start355 356We provide an example code to run `InternVL3.5-8B` using `transformers`. Please note that our models with up to 30B parameters can be deployed on a single A100 GPU, while the 38B model requires two A100 GPUs and the 235B model requires eight A100 GPUs.357 358> In most cases, both [LMDeploy](https://github.com/InternLM/lmdeploy) and [vLLM](https://github.com/vllm-project/vllm) can be used for model deployment. However, for InternVL3.5-20B-A4B, we recommend using vLLM since lmdeploy has not yet supported GPT-OSS.359 360> Please use transformers>=4.52.1 to ensure the model works normally. For the 20B version of our model, transformers>=4.55.0 is required.361 362### Model Loading363 364#### 16-bit (bf16 / fp16)365 366```python367import torch368from transformers import AutoTokenizer, AutoModel369path = "OpenGVLab/InternVL3_5-8B"370model = AutoModel.from_pretrained(371    path,372    torch_dtype=torch.bfloat16,373    low_cpu_mem_usage=True,374    use_flash_attn=True,375    trust_remote_code=True).eval().cuda()376```377 378#### BNB 8-bit Quantization379 380```python381import torch382from transformers import AutoTokenizer, AutoModel383path = "OpenGVLab/InternVL3_5-8B"384model = AutoModel.from_pretrained(385    path,386    torch_dtype=torch.bfloat16,387    load_in_8bit=True,388    low_cpu_mem_usage=True,389    use_flash_attn=True,390    trust_remote_code=True).eval()391```392 393#### Multiple GPUs394 395```python396import math397import torch398from transformers import AutoTokenizer, AutoModel399 400path = "OpenGVLab/InternVL3_5-8B"401model = AutoModel.from_pretrained(402    path,403    torch_dtype=torch.bfloat16,404    low_cpu_mem_usage=True,405    use_flash_attn=True,406    trust_remote_code=True,407    device_map="auto").eval()408```409 410### Thinking Mode411 412To enable thinking mode, please set the system prompt to our Thinking System Prompt. When enabling Thinking mode, we recommend setting `do_sample=True` and `temperature=0.6` to mitigate undesired repetition.413 414```python415R1_SYSTEM_PROMPT = """416You are an AI assistant that rigorously follows this response protocol:417 4181. First, conduct a detailed analysis of the question. Consider different angles, potential solutions, and reason through the problem step-by-step. Enclose this entire thinking process within <think> and </think> tags.419 4202. After the thinking section, provide a clear, concise, and direct answer to the user's question. Separate the answer from the think section with a newline.421 422Ensure that the thinking process is thorough but remains focused on the query. The final answer should be standalone and not reference the thinking section.423""".strip()424 425model.system_message = R1_SYSTEMP_PROMPT426```427 428### Inference with Transformers429 430```python431import math432import numpy as np433import torch434import torchvision.transforms as T435from decord import VideoReader, cpu436from PIL import Image437from torchvision.transforms.functional import InterpolationMode438from transformers import AutoModel, AutoTokenizer439 440IMAGENET_MEAN = (0.485, 0.456, 0.406)441IMAGENET_STD = (0.229, 0.224, 0.225)442 443def build_transform(input_size):444    MEAN, STD = IMAGENET_MEAN, IMAGENET_STD445    transform = T.Compose([446        T.Lambda(lambda img: img.convert('RGB') if img.mode != 'RGB' else img),447        T.Resize((input_size, input_size), interpolation=InterpolationMode.BICUBIC),448        T.ToTensor(),449        T.Normalize(mean=MEAN, std=STD)450    ])451    return transform452 453def find_closest_aspect_ratio(aspect_ratio, target_ratios, width, height, image_size):454    best_ratio_diff = float('inf')455    best_ratio = (1, 1)456    area = width * height457    for ratio in target_ratios:458        target_aspect_ratio = ratio[0] / ratio[1]459        ratio_diff = abs(aspect_ratio - target_aspect_ratio)460        if ratio_diff < best_ratio_diff:461            best_ratio_diff = ratio_diff462            best_ratio = ratio463        elif ratio_diff == best_ratio_diff:464            if area > 0.5 * image_size * image_size * ratio[0] * ratio[1]:465                best_ratio = ratio466    return best_ratio467 468def dynamic_preprocess(image, min_num=1, max_num=12, image_size=448, use_thumbnail=False):469    orig_width, orig_height = image.size470    aspect_ratio = orig_width / orig_height471 472    # calculate the existing image aspect ratio473    target_ratios = set(474        (i, j) for n in range(min_num, max_num + 1) for i in range(1, n + 1) for j in range(1, n + 1) if475        i * j <= max_num and i * j >= min_num)476    target_ratios = sorted(target_ratios, key=lambda x: x[0] * x[1])477 478    # find the closest aspect ratio to the target479    target_aspect_ratio = find_closest_aspect_ratio(480        aspect_ratio, target_ratios, orig_width, orig_height, image_size)481 482    # calculate the target width and height483    target_width = image_size * target_aspect_ratio[0]484    target_height = image_size * target_aspect_ratio[1]485    blocks = target_aspect_ratio[0] * target_aspect_ratio[1]486 487    # resize the image488    resized_img = image.resize((target_width, target_height))489    processed_images = []490    for i in range(blocks):491        box = (492            (i % (target_width // image_size)) * image_size,493            (i // (target_width // image_size)) * image_size,494            ((i % (target_width // image_size)) + 1) * image_size,495            ((i // (target_width // image_size)) + 1) * image_size496        )497        # split the image498        split_img = resized_img.crop(box)499        processed_images.append(split_img)500    assert len(processed_images) == blocks501    if use_thumbnail and len(processed_images) != 1:502        thumbnail_img = image.resize((image_size, image_size))503        processed_images.append(thumbnail_img)504    return processed_images505 506def load_image(image_file, input_size=448, max_num=12):507    image = Image.open(image_file).convert('RGB')508    transform = build_transform(input_size=input_size)509    images = dynamic_preprocess(image, image_size=input_size, use_thumbnail=True, max_num=max_num)510    pixel_values = [transform(image) for image in images]511    pixel_values = torch.stack(pixel_values)512    return pixel_values513 514path = 'OpenGVLab/InternVL3_5-8B'515model = AutoModel.from_pretrained(516    path,517    torch_dtype=torch.bfloat16,518    load_in_8bit=False,519    low_cpu_mem_usage=True,520    use_flash_attn=True,521    trust_remote_code=True,522    device_map="auto").eval()523tokenizer = AutoTokenizer.from_pretrained(path, trust_remote_code=True, use_fast=False)524 525# set the max number of tiles in `max_num`526pixel_values = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()527generation_config = dict(max_new_tokens=1024, do_sample=True)528 529# pure-text conversation (纯文本对话)530question = 'Hello, who are you?'531response, history = model.chat(tokenizer, None, question, generation_config, history=None, return_history=True)532print(f'User: {question}\nAssistant: {response}')533 534question = 'Can you tell me a story?'535response, history = model.chat(tokenizer, None, question, generation_config, history=history, return_history=True)536print(f'User: {question}\nAssistant: {response}')537 538# single-image single-round conversation (单图单轮对话)539question = '<image>\nPlease describe the image shortly.'540response = model.chat(tokenizer, pixel_values, question, generation_config)541print(f'User: {question}\nAssistant: {response}')542 543# single-image multi-round conversation (单图多轮对话)544question = '<image>\nPlease describe the image in detail.'545response, history = model.chat(tokenizer, pixel_values, question, generation_config, history=None, return_history=True)546print(f'User: {question}\nAssistant: {response}')547 548question = 'Please write a poem according to the image.'549response, history = model.chat(tokenizer, pixel_values, question, generation_config, history=history, return_history=True)550print(f'User: {question}\nAssistant: {response}')551 552# multi-image multi-round conversation, combined images (多图多轮对话,拼接图像)553pixel_values1 = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()554pixel_values2 = load_image('./examples/image2.jpg', max_num=12).to(torch.bfloat16).cuda()555pixel_values = torch.cat((pixel_values1, pixel_values2), dim=0)556 557question = '<image>\nDescribe the two images in detail.'558response, history = model.chat(tokenizer, pixel_values, question, generation_config,559                               history=None, return_history=True)560print(f'User: {question}\nAssistant: {response}')561 562question = 'What are the similarities and differences between these two images.'563response, history = model.chat(tokenizer, pixel_values, question, generation_config,564                               history=history, return_history=True)565print(f'User: {question}\nAssistant: {response}')566 567# multi-image multi-round conversation, separate images (多图多轮对话,独立图像)568pixel_values1 = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()569pixel_values2 = load_image('./examples/image2.jpg', max_num=12).to(torch.bfloat16).cuda()570pixel_values = torch.cat((pixel_values1, pixel_values2), dim=0)571num_patches_list = [pixel_values1.size(0), pixel_values2.size(0)]572 573question = 'Image-1: <image>\nImage-2: <image>\nDescribe the two images in detail.'574response, history = model.chat(tokenizer, pixel_values, question, generation_config,575                               num_patches_list=num_patches_list,576                               history=None, return_history=True)577print(f'User: {question}\nAssistant: {response}')578 579question = 'What are the similarities and differences between these two images.'580response, history = model.chat(tokenizer, pixel_values, question, generation_config,581                               num_patches_list=num_patches_list,582                               history=history, return_history=True)583print(f'User: {question}\nAssistant: {response}')584 585# batch inference, single image per sample (单图批处理)586pixel_values1 = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()587pixel_values2 = load_image('./examples/image2.jpg', max_num=12).to(torch.bfloat16).cuda()588num_patches_list = [pixel_values1.size(0), pixel_values2.size(0)]589pixel_values = torch.cat((pixel_values1, pixel_values2), dim=0)590 591questions = ['<image>\nDescribe the image in detail.'] * len(num_patches_list)592responses = model.batch_chat(tokenizer, pixel_values,593                             num_patches_list=num_patches_list,594                             questions=questions,595                             generation_config=generation_config)596for question, response in zip(questions, responses):597    print(f'User: {question}\nAssistant: {response}')598 599# video multi-round conversation (视频多轮对话)600def get_index(bound, fps, max_frame, first_idx=0, num_segments=32):601    if bound:602        start, end = bound[0], bound[1]603    else:604        start, end = -100000, 100000605    start_idx = max(first_idx, round(start * fps))606    end_idx = min(round(end * fps), max_frame)607    seg_size = float(end_idx - start_idx) / num_segments608    frame_indices = np.array([609        int(start_idx + (seg_size / 2) + np.round(seg_size * idx))610        for idx in range(num_segments)611    ])612    return frame_indices613 614def load_video(video_path, bound=None, input_size=448, max_num=1, num_segments=32):615    vr = VideoReader(video_path, ctx=cpu(0), num_threads=1)616    max_frame = len(vr) - 1617    fps = float(vr.get_avg_fps())618 619    pixel_values_list, num_patches_list = [], []620    transform = build_transform(input_size=input_size)621    frame_indices = get_index(bound, fps, max_frame, first_idx=0, num_segments=num_segments)622    for frame_index in frame_indices:623        img = Image.fromarray(vr[frame_index].asnumpy()).convert('RGB')624        img = dynamic_preprocess(img, image_size=input_size, use_thumbnail=True, max_num=max_num)625        pixel_values = [transform(tile) for tile in img]626        pixel_values = torch.stack(pixel_values)627        num_patches_list.append(pixel_values.shape[0])628        pixel_values_list.append(pixel_values)629    pixel_values = torch.cat(pixel_values_list)630    return pixel_values, num_patches_list631 632video_path = './examples/red-panda.mp4'633pixel_values, num_patches_list = load_video(video_path, num_segments=8, max_num=1)634pixel_values = pixel_values.to(torch.bfloat16).cuda()635video_prefix = ''.join([f'Frame{i+1}: <image>\n' for i in range(len(num_patches_list))])636question = video_prefix + 'What is the red panda doing?'637# Frame1: <image>\nFrame2: <image>\n...\nFrame8: <image>\n{question}638response, history = model.chat(tokenizer, pixel_values, question, generation_config,639                               num_patches_list=num_patches_list, history=None, return_history=True)640print(f'User: {question}\nAssistant: {response}')641 642question = 'Describe this video in detail.'643response, history = model.chat(tokenizer, pixel_values, question, generation_config,644                               num_patches_list=num_patches_list, history=history, return_history=True)645print(f'User: {question}\nAssistant: {response}')646```647 648#### Streaming Output649 650Besides this method, you can also use the following code to get streamed output.651 652```python653from transformers import TextIteratorStreamer654from threading import Thread655 656# Initialize the streamer657streamer = TextIteratorStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True, timeout=10)658# Define the generation configuration659generation_config = dict(max_new_tokens=1024, do_sample=False, streamer=streamer)660# Start the model chat in a separate thread661thread = Thread(target=model.chat, kwargs=dict(662    tokenizer=tokenizer, pixel_values=pixel_values, question=question,663    history=None, return_history=False, generation_config=generation_config,664))665thread.start()666 667# Initialize an empty string to store the generated text668generated_text = ''669# Loop through the streamer to get the new text as it is generated670for new_text in streamer:671    if new_text == model.conv_template.sep:672        break673    generated_text += new_text674    print(new_text, end='', flush=True)  # Print each new chunk of generated text on the same line675```676 677## Finetune678 679Many repositories now support fine-tuning of the InternVL series models, including [InternVL](https://github.com/OpenGVLab/InternVL), [SWIFT](https://github.com/modelscope/ms-swift), [XTuner](https://github.com/InternLM/xtuner), and others. Please refer to their documentation for more details on fine-tuning.680 681## Deployment682 683### LMDeploy684 685LMDeploy is a toolkit for compressing, deploying, and serving LLMs & VLMs.686 687```sh688pip install lmdeploy>=0.9.1689```690 691LMDeploy abstracts the complex inference process of multi-modal Vision-Language Models (VLM) into an easy-to-use pipeline, similar to the Large Language Model (LLM) inference pipeline.692 693#### A 'Hello, world' Example694 695```python696from lmdeploy import pipeline, PytorchEngineConfig697from lmdeploy.vl import load_image698 699image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/tests/data/tiger.jpeg')700 701# Please set tp=2 for the 38B version and tp=8 for the 241B-A28B version.702model = 'OpenGVLab/InternVL3_5-8B'703pipe = pipeline(model, backend_config=PytorchEngineConfig(session_len=32768, tp=1))704 705response = pipe(('describe this image', image))706print(response.text)707```708 709#### Multi-images Inference710 711When dealing with multiple images, you can put them all in one list. Keep in mind that multiple images will lead to a higher number of input tokens, and as a result, the size of the context window typically needs to be increased.712 713```python714from lmdeploy import pipeline, PytorchEngineConfig715from lmdeploy.vl import load_image716from lmdeploy.vl.constants import IMAGE_TOKEN717 718# Please set tp=2 for the 38B version and tp=8 for the 241B-A28B version.719model = 'OpenGVLab/InternVL3_5-8B'720pipe = pipeline(model, backend_config=PytorchEngineConfig(session_len=32768, tp=1))721 722image_urls=[723    'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg',724    'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/det.jpg'725]726 727images = [load_image(img_url) for img_url in image_urls]728# Numbering images improves multi-image conversations729response = pipe((f'Image-1: {IMAGE_TOKEN}\nImage-2: {IMAGE_TOKEN}\ndescribe these two images', images))730print(response.text)731```732 733#### Batch Prompts Inference734 735Conducting inference with batch prompts is quite straightforward; just place them within a list structure:736 737```python738from lmdeploy import pipeline, PytorchEngineConfig739from lmdeploy.vl import load_image740 741# Please set tp=2 for the 38B version and tp=8 for the 241B-A28B version.742model = 'OpenGVLab/InternVL3_5-8B'743pipe = pipeline(model, backend_config=PytorchEngineConfig(session_len=32768, tp=1))744 745image_urls=[746    "https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg",747    "https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/det.jpg"748]749prompts = [('describe this image', load_image(img_url)) for img_url in image_urls]750response = pipe(prompts)751print(response)752```753 754#### Multi-turn Conversation755 756There are two ways to do the multi-turn conversations with the pipeline. One is to construct messages according to the format of OpenAI and use above introduced method, the other is to use the `pipeline.chat` interface.757 758```python759from lmdeploy import pipeline, PytorchEngineConfig, GenerationConfig760from lmdeploy.vl import load_image761 762# Please set tp=2 for the 38B version and tp=8 for the 241B-A28B version.763model = 'OpenGVLab/InternVL3_5-8B'764pipe = pipeline(model, backend_config=PytorchEngineConfig(session_len=32768, tp=1))765 766image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg')767gen_config = GenerationConfig(top_k=50, top_p=0.95, temperature=0.6, max_new_tokens=8192)768sess = pipe.chat(('describe this image', image), gen_config=gen_config)769print(sess.response.text)770sess = pipe.chat('What is the woman doing?', session=sess, gen_config=gen_config)771print(sess.response.text)772```773 774#### Service775 776LMDeploy's `api_server` enables models to be easily packed into services with a single command. The provided RESTful APIs are compatible with OpenAI's interfaces. Below are an example of service startup:777 778```shell779lmdeploy serve api_server OpenGVLab/InternVL3_5-8B --server-port 23333 --tp 1 --backend pytorch780```781 782To use the OpenAI-style interface, you need to install OpenAI:783 784```shell785pip install openai786```787 788Then, use the code below to make the API call:789 790```python791from openai import OpenAI792 793client = OpenAI(api_key='YOUR_API_KEY', base_url='http://0.0.0.0:23333/v1')794model_name = client.models.list().data[0].id795response = client.chat.completions.create(796    model=model_name,797    messages=[{798        'role':799        'user',800        'content': [{801            'type': 'text',802            'text': 'describe this image',803        }, {804            'type': 'image_url',805            'image_url': {806                'url':807                'https://modelscope.oss-cn-beijing.aliyuncs.com/resource/tiger.jpeg',808            },809        }],810    }],811    temperature=0.8,812    top_p=0.8)813print(response)814```815 816## License817 818This project is released under the apache-2.0 License. This project uses the pre-trained Qwen3 as a component, which is licensed under the apache-2.0 License.819 820## Citation821 822If you find this project useful in your research, please consider citing:823 824```BibTeX825@article{wang2025internvl3_5,826  title={InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency},827  author={Wang, Weiyun and Gao, Zhangwei and Gu, Lixin and Pu, Hengjun and Cui, Long and Wei, Xingguang and Liu, Zhaoyang and Jing, Linglin and Ye, Shenglong and Shao, Jie and others},828  journal={arXiv preprint arXiv:2508.18265},829  year={2025}830}831```832