OpenGVLab/InternVL3_5-1B-Flash
11906
1---2license: apache-2.03pipeline_tag: image-text-to-text4library_name: transformers5base_model:6 - OpenGVLab/InternVL3_5-1B7base_model_relation: finetune8datasets:9 - OpenGVLab/MMPR-v1.210 - OpenGVLab/MMPR-Tiny11language:12 - multilingual13tags:14 - internvl15 - custom_code16---17 18# InternVL3_5-1B-Flash19 20[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggingface.co/papers/2411.10442) [\[📜 InternVL3\]](https://huggingface.co/papers/2504.10479) [\[📜 InternVL3.5\]](https://huggingface.co/papers/2508.18265)21 22[\[🆕 Blog\]](https://internvl.github.io/blog/) [\[🗨️ Chat Demo\]](https://chat.intern-ai.org.cn/) [\[🚀 Quick Start\]](#quick-start) [\[📖 Documents\]](https://internvl.readthedocs.io/en/latest/)23 24<div align="center">25 <img width="500" alt="image" src="https://cdn-uploads.huggingface.co/production/uploads/64006c09330a45b03605bba3/zJsd2hqd3EevgXo6fNgC-.png">26</div>27 28## Introduction29 30We introduce *InternVL3.5*, a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL series. A key innovation is the *Cascade Reinforcement Learning (Cascade RL)* framework, which enhances reasoning through a two-stage process: offline RL for stable convergence and online RL for refined alignment. This coarse-to-fine training strategy leads to substantial improvements on downstream reasoning tasks, e.g., MMMU and MathVista. To optimize efficiency, we propose a *Visual Resolution Router (ViR)* that dynamically adjusts the resolution of visual tokens without compromising performance. Coupled with ViR, our Decoupled *Vision-Language Deployment (DvD)* strategy separates the vision encoder and language model across different GPUs, effectively balancing computational load. These contributions collectively enable InternVL3.5 to achieve up to a +16.0\% gain in overall reasoning performance and a 4.05 \\(\times\\) inference speedup compared to its predecessor, i.e., InternVL3. In addition, InternVL3.5 supports novel capabilities such as GUI interaction and embodied agency. Notably, our largest model, i.e., InternVL3.5-241B-A28B, attains state-of-the-art results among open-source MLLMs across general multimodal, reasoning, text, and agentic tasks—narrowing the performance gap with leading commercial models like GPT-5. All models and code are publicly released.31 3233 34> Hatched bars represent closed-source commercial models. We report average scores on a set of multimodal general, reasoning, text, and agentic benchmarks: MMBench v1.1 (en), MMStar,BLINK, HallusionBench, AI2D, OCRBench, MMVet, MME-RealWorld (en), MVBench, VideoMME, MMMU, MathVista, MathVision, MathVerse, DynaMath, WeMath, LogicVista, MATH500, AIME24, AIME25, GPQA, MMLU-Pro, GAOKAO, IFEval, SGP-Bench, VSI-Bench, ERQA, SpaCE-10, and OmniSpatial.35 36See [quick start](#quick-start) for how to use our model.37 38## InternVL3.5 Family39 40In the following table, we provide an overview of the InternVL3.5 series.41To maintain consistency with earlier generations, we provide two model formats: [the GitHub format](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B), consistent with prior releases, and [the HF format](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B-HF), aligned with the official Transformers standard.42 43> If you want to convert the checkpoint between these two formats, please refer to the scripts about [custom2hf](https://github.com/OpenGVLab/InternVL/blob/main/internvl_chat/tools/internvl_custom2hf.py) and [hf2custom](https://github.com/OpenGVLab/InternVL/blob/main/internvl_chat/tools/internvl_hf2custom.py).44 45 46### Github Format47 48 49| Model | #Vision Param | #Language Param | #Total Param | HF Link | ModelScope Link |50| --------------------- | ------------- | --------------- | ------------ | ------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------- |51| InternVL3.5-1B | 0.3B | 0.8B | 1.1B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-1B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-1B) |52| InternVL3.5-2B | 0.3B | 2.0B | 2.3B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-2B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-2B) |53| InternVL3.5-4B | 0.3B | 4.4B | 4.7B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-4B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-4B) |54| InternVL3.5-8B | 0.3B | 8.2B | 8.5B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-8B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-8B) |55| InternVL3.5-14B | 0.3B | 14.8B | 15.1B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-14B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-14B) |56| InternVL3.5-38B | 5.5B | 32.8B | 38.4B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-38B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-38B) |57| InternVL3.5-20B-A4B | 0.3B | 20.9B | 21.2B-A4B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-GPT-OSS-20B-A4B-Preview) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-GPT-OSS-20B-A4B-Preview) |58| InternVL3.5-30B-A3B | 0.3B | 30.5B | 30.8B-A3B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-30B-A3B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-30B-A3B) |59| InternVL3.5-241B-A28B | 5.5B | 235.1B | 240.7B-A28B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-241B-A28B) |60 61 62### HuggingFace Format63 64 65| Model | #Vision Param | #Language Param | #Total Param | HF Link | ModelScope Link |66| ------------------------ | ------------- | --------------- | ------------ | --------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- |67| InternVL3.5-1B-HF | 0.3B | 0.8B | 1.1B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-1B-HF) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-1B-HF) |68| InternVL3.5-2B-HF | 0.3B | 2.0B | 2.3B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-2B-HF) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-2B-HF) |69| InternVL3.5-4B-HF | 0.3B | 4.4B | 4.7B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-4B-HF) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-4B-HF) |70| InternVL3.5-8B-HF | 0.3B | 8.2B | 8.5B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-8B-HF) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-8B-HF) |71| InternVL3.5-14B-HF | 0.3B | 14.8B | 15.1B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-14B-HF) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-14B-HF) |72| InternVL3.5-38B-HF | 5.5B | 32.8B | 38.4B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-38B-HF) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-38B-HF) |73| InternVL3.5-20B-A4B-HF | 0.3B | 20.9B | 21.2B-A4B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-GPT-OSS-20B-A4B-Preview-HF) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-GPT-OSS-20B-A4B-Preview-HF) |74| InternVL3.5-30B-A3B-HF | 0.3B | 30.5B | 30.8B-A3B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-30B-A3B-HF) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-30B-A3B-HF) |75| InternVL3.5-241B-A28B-HF | 5.5B | 235.1B | 240.7B-A28B | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B-HF) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-241B-A28B-HF) |76 77 7879 80> We conduct the evaluation with [VLMEvalkit](https://github.com/open-compass/VLMEvalKit). ***To enable the Thinking mode of our model, please set the system prompt to [R1_SYSTEM_PROMPT](https://github.com/open-compass/VLMEvalKit/blob/main/vlmeval/vlm/internvl/internvl_chat.py#L38).*** When enabling Thinking mode, we recommend setting `do_sample=True` and `temperature=0.6` to mitigate undesired repetition.81 82Our training pipeline comprises four stages: Multimodal Continual Pre-Training (**CPT**), Supervised Fine-Tuning (**SFT**), and Cascade Reinforcement Learning (**CascadeRL**). In CascadeRL, we first fine-tune the model using Mixed Preference Optimization (**MPO**) under an offline RL setting, followed by **GSPO** under an oneline RL setting.83For the Flash version of InternVL3.5, we additionally introduce a lightweight training stage, termed Visual Consistency Learning (**ViCO**), which reduces the token cost required to represent an image patch.84 8586 87Here, we also open-source the model weights after different training stages for potential research usage.88***If you're unsure which version to use, please select the one without any suffix, as it has completed the full training pipeline.***89 90 91| Model | Training Pipeline | HF Link | ModelScope Link |92| -------------------------------- | --------------------- | --------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- |93| InternVL3.5-1B-Pretrained | CPT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-1B-Pretrained) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-1B-Pretrained) |94| InternVL3.5-1B-Instruct | CPT + SFT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-1B-Instruct) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-1B-Instruct) |95| InternVL3.5-1B-MPO | CPT + SFT + MPO | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-1B-MPO) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-1B-MPO) |96| InternVL3.5-1B | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-1B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-1B) |97| InternVL3.5-2B-Pretrained | CPT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-2B-Pretrained) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-2B-Pretrained) |98| InternVL3.5-2B-Instruct | CPT + SFT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-2B-Instruct) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-2B-Instruct) |99| InternVL3.5-2B-MPO | CPT + SFT + MPO | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-2B-MPO) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-2B-MPO) |100| InternVL3.5-2B | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-2B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-2B) |101| InternVL3.5-4B-Pretrained | CPT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-4B-Pretrained) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-4B-Pretrained) |102| InternVL3.5-4B-Instruct | CPT + SFT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-4B-Instruct) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-4B-Instruct) |103| InternVL3.5-4B-MPO | CPT + SFT + MPO | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-4B-MPO) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-4B-MPO) |104| InternVL3.5-4B | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-4B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-4B) |105| InternVL3.5-8B-Pretrained | CPT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-8B-Pretrained) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-8B-Pretrained) |106| InternVL3.5-8B-Instruct | CPT + SFT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-8B-Instruct) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-8B-Instruct) |107| InternVL3.5-8B-MPO | CPT + SFT + MPO | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-8B-MPO) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-8B-MPO) |108| InternVL3.5-8B | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-8B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-8B) |109| InternVL3.5-14B-Pretrained | CPT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-14B-Pretrained) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-14B-Pretrained) |110| InternVL3.5-14B-Instruct | CPT + SFT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-14B-Instruct) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-14B-Instruct) |111| InternVL3.5-14B-MPO | CPT + SFT + MPO | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-14B-MPO) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-14B-MPO) |112| InternVL3.5-14B | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-14B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-14B) |113| InternVL3.5-30B-A3B-Pretrained | CPT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-30B-A3B-Pretrained) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-30B-A3B-Pretrained) |114| InternVL3.5-30B-A3B-Instruct | CPT + SFT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-30B-A3B-Instruct) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-30B-A3B-Instruct) |115| InternVL3.5-30B-A3B-MPO | CPT + SFT + MPO | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-30B-A3B-MPO) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-30B-A3B-MPO) |116| InternVL3.5-30B-A3B | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-30B-A3B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-30B-A3B) |117| InternVL3.5-38B-Pretrained | CPT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-38B-Pretrained) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-38B-Pretrained) |118| InternVL3.5-38B-Instruct | CPT + SFT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-38B-Instruct) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-38B-Instruct) |119| InternVL3.5-38B-MPO | CPT + SFT + MPO | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-38B-MPO) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-38B-MPO) |120| InternVL3.5-38B | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-38B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-38B) |121| InternVL3.5-241B-A28B-Pretrained | CPT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B-Pretrained) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-241B-A28B-Pretrained) |122| InternVL3.5-241B-A28B-Instruct | CPT + SFT | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B-Instruct) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-241B-A28B-Instruct) |123| InternVL3.5-241B-A28B-MPO | CPT + SFT + MPO | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B-MPO) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-241B-A28B-MPO) |124| InternVL3.5-241B-A28B | CPT + SFT + CascadeRL | [🤗 link](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B) | [🤖 link](https://www.modelscope.cn/models/OpenGVLab/InternVL3_5-241B-A28B) |125 126 127The Flash version of our model will be released as soon as possible.128 129 130 131## Model Architecture132 133`InternVL3.5`:134This series of models follow the "ViT–MLP–LLM" paradigm adopted in previous versions of InternVL.135We initialize the language model using the Qwen3 series and GPT-OSS, and the vision encoder using InternViT-300M and InternViT-6B.136The Dynamic High Resolution strategy introduced in InternVL1.5 is also retained in our design.137 138 139`InternVL3.5-Flash`:140Compared to InternVL3.5, InternVL3.5-Flash further integrates the *Visual Resolution Router (ViR)*, thus yielding a series of efficient variants friendly suitable for resource-constrained scenarios. 141Specifically, in InternVL3.5, each image patch is initially represented as 1024 visual tokens for the vision encoder, which are then compressed into 256 tokens via a pixel shuffle module before being passed to the Large Language Model (LLM).142In InternVL3.5-Flash, as shown in the Figure below, an additional pixel shuffle module with a higher compression rate is included, enabling the compression of visual tokens down to 64 tokens.143For each patch, the patch router determines the appropriate compression rate by assessing its semantic richness, and routes it to the corresponding pixel shuffle module accordingly.144Benefiting from this patch-aware compression mechanism, InternVL3.5-Flash is able to reduce the number of visual tokens by 50\% while maintaining nearly 100\% of the performance of InternVL3.5.145 146 147148 149## Training and Deployment Strategy150 151### Pre-Training152 153During the pre-training stage, we update all model parameters jointly using the combination of large-scale text and multimodal corpora. Specifically, given an arbitrary training sample consisting of a multimodal token sequence \\(\mathbf{x}=\left(x_1, x_2, \ldots, x_L\right)\\), the next token prediction (NTP) loss is calculated on each text token as follows:154 155$$156 \mathcal{L}_{i}=-\log p_\theta\left(x_i \mid x_1, \ldots, x_{i-1}\right),157$$158 159where \\(x_i\\) is the predicted token and prefix tokens in \\(\{x_1, x_2, \ldots, x_{i-1}\}\\) can be either text tokens or image tokens. Notably, for conversation samples, only response tokens are included for the calculation of the loss.160Additionally, to mitigate bias toward either longer or shorter responses during training, we adopt the square averaging to re-weight the NTP loss as follows:161 162$$163\mathcal{L}_{i}^{'} = \frac{w_i}{\sum_j w_j} \cdot \mathcal{L}_i, \quad w_i = \frac{1}{N^{0.5}},164$$165 166where \\(N\\) denotes the number of tokens in the training sample on which the loss needs to be calculated. The random JPEG compression is also included to enhance the model's real-world performance.167 168### Supervised Fine-Tuning169 170During the SFT phase, we adopt the same objective as in the pre-training stage and use the square-root averaging strategy to calculate the final loss. In this stage, the context window is set to 32K tokens to adapt long-context information.171Compared to InternVL3, the SFT stage of InternVL3.5 contains more high-quality and diverse training data derived from three sources: 172 173(1) Instruction-following data from InternVL3, which are reused to preserve broad coverage of vision–language tasks. 174 175(2) Multimodal reasoning data in the "Thinking" mode, which are included to instill long-thinking capabilities in the model. To construct such data, we first use InternVL3-78B to describe the image and then input the description into DeepSeek-R1 to sample rollouts with detailed reasoning processes. Rollouts with an incorrect final answer are filtered out. The questions in these datasets cover various expert domains, such as mathematics and scientific disciplines, thereby strengthening performance on different reasoning tasks. 176 177(3) Capability-expansion datasets, which endow InternVL3.5 with new skills, including GUI-based interaction, embodied interaction, and scalable vect178 179### Cascade Reinforcement Learning180 181Cascade RL aims to combine the benefits of offline RL and online RL to progressively facilitate the post-training of MLLMs in an efficient manner.182Specifically, we first fine-tune the model using an offline RL algorithm as an efficient warm-up stage to reach a satisfied results, which can guarantee the high-quality rollouts for the latter stage. 183Subsequently, we employ an online RL algorithm to further refine the output distribution based on rollouts generated by the model itself. Compared to the single offline or online RL stage, our cascaded RL achieves significant performance improvements at a fraction of the GPU time cost.184 185 186 187During the offline RL stage, we employ mixed preference optimization (MPO) to fine-tune the model. Specifically, the training objective of MPO is a combination of preference loss \\(\mathcal{L}_{p}\\), quality loss \\(\mathcal{L}_{q}\\), and generation loss \\(\mathcal{L}_{g}\\), which can be formulated as follows:188 189$$190 \mathcal{L}_{\text{MPO}}=191 w_{p} \mathcal{L}_{p}192 +193 w_{q} \mathcal{L}_{q}194 +195 w_{g} \mathcal{L}_{g}196 ,197$$198 199where \\(w_{*}\\) represents the weight assigned to each loss component.200The DPO loss, BCO loss, and LM loss serve as the preference loss, quality loss, and generation loss, respectively.201 202 203During the online RL stage, we employ GSPO, without reference model constraints, as our online RL algorithm, which we find more effective in training both dense and mixture-of-experts (MoE) models. Similar to GRPO, the advantage is defined as the normalized reward across responses sampled from the same query.204The training objective of GSPO is given by:205 206$$207 \mathcal{L}_{\mathrm{GSPO}}(\theta)=\mathbb{E}_{x \sim \mathcal{D},\left\{y_i\right\}_{i=1}^G \sim \pi_{\theta \text { old }}(\cdot \mid x)}\left[\frac{1}{G} \sum_{i=1}^G \min \left(s_i(\theta) \widehat{A}_i, \operatorname{clip}\left(s_i(\theta), 1-\varepsilon, 1+\varepsilon\right) \widehat{A}_i\right)\right],208$$209 210where the importance sampling ratio is defined as the geometric mean of the per-token ratios.211 212> Please see [our paper](https://huggingface.co/papers/2508.18265) for more technical and experimental details.213 214 215### Visual Consistency Learning216 217 218We further include ViCO as an additional training stage to integrate the *visual resolution router (ViR)* into InternVL3.5, thereby reducing the inference cost of InternVL3.5. The obtained efficient version of InternVL3.5 are termed as *InternVL3.5-Flash*. In particular, ViCO comprises two stages:219 220`Consistency training`:221In this stage, the entire model is trained to minimize the divergence between response distributions conditioned on visual tokens with different compression rates.222In practice, we introduce an extra reference model, which is frozen and initialized with InternVL3.5.223Given a sample, each image patch is represented as either 256 or 64 tokens, and the training objective is defined as follows:224 225 226$$227\mathcal{L}_\text{ViCO} =228\mathbb{E}_{\xi \sim \mathcal{R}} \Bigg[229\frac{1}{N} \sum_{i=1}^{N} \mathrm{KL} \Big(230\pi_{\theta_{ref}}\left(y_i \mid y_{<i}, I\right) \;\Big\|\;231\pi_{\theta_{policy}}\left(y_i \mid y_{<i}, I_\xi\right)232\Big)233\Bigg],234$$235 236where \\(\mathrm{KL}\) denotes the KL divergence and \(\xi\) denotes the compression rate, which is uniformly sampled from \(\{\frac{1}{4},\frac{1}{16}\}\). The image \(I_\xi\) is represented as 256 tokens when \(\xi=\frac{1}{4}\) and 64 tokens when \(\xi=\frac{1}{16}\). Notably, the reference model always performs inference with \(\xi=\frac{1}{4}\).237 238 239`Router training`:240This stage aims to train the ViR to select an appropriate trade-off resolution for different inputs.241ViR is formulated as a binary classifier and trained using standard cross-entropy loss.242To construct the route targets, we first compute the KL divergence between the model outputs conditioned on uncompressed visual tokens (i.e., 256 tokens per patch) and those conditioned on compressed visual tokens (i.e., 64 tokens per patch).243During this stage, the main MLLM (ViT, MLP and LLM) is kept frozen, and only the ViR is trained.244Specifically, we first compute the loss ratio for each patch:245 246$$247r_i = \frac{\mathcal{L}_\text{ViCO}\big(y_i \mid I_{\frac{1}{16}}\big)}{\mathcal{L}_\text{ViCO}\big(y_i \mid I_{\frac{1}{4}}\big)},248$$249 250which quantifies the relative increase in loss caused by compressing the visual tokens. Based on this ratio, the binary ground-truth label for the patch router is defined as:251 252$$253y_i^\text{router} =254\begin{cases}2550, & r_i < \tau \; \text{(compression has negligible impact)} \\2561, & r_i \ge \tau \; \text{(compression has significant impact)},257\end{cases}258$$259 260where \(y_i^{\text{router}}=0\) and \(y_i^{\text{router}}=1\) indicate that the compression rate \(\xi\) is set to \(\tfrac{1}{16}\) and \(\tfrac{1}{4}\), respectively.261 262> Please see [our paper](https://huggingface.co/papers/2508.18265) for more technical and experimental details.263 264 265### Test-Time Scaling266 267 268Test-time scaling (TTS) has been empirically demonstrated as an effective approach to enhance the reasoning capabilities of LLMs and MLLMs, particularly for complex tasks necessitating multi-step inference.269In this work, we implement a comprehensive test-time scaling approach that simultaneously improves reasoning depth (i.e., deep thinking) and breadth (i.e., parallel thinking).270 271`Deep Thinking`: By activating the Thinking mode, we guide the model to deliberately engage in step-by-step reasoning (i.e., decomposing complex problems into logical steps and validating intermediate conclusions) prior to generating the final answer. This approach systematically improves the logical structure of solutions for complex problems, particularly those requiring multi-step inference, and enhances reasoning depth.272 273`Parallel Thinking`: Following InternVL3, for reasoning tasks, we adopt the Best-of-N (BoN) strategy by employing [VisualPRM-v1.1](https://huggingface.co/OpenGVLab/VisualPRM-8B-v1_1) as the critic model to select the optimal response from multiple reasoning candidates.274This approach improves reasoning breadth.275 276> Notably, unless otherwise specified, the experimental results reported in our paper are obtained without applying TTS. Thus far, we have only applied TTS to reasoning benchmarks, since we found that the model already exhibits strong perception and understanding capabilities, and initiating TTS yields no significant improvement.277 278 279### Decoupled Vision-Language Deployment280 281In multimodal inference, the vision encoder and language model have distinct computational characteristics. The vision encoder that transforms images into semantic features is highly parallelizable and does not rely on long-term history state. In contrast, the language model adopts the inference in an autoregressive manner, which requires previous states to compute the next one. This sequential property makes the language part more sensitive to memory bandwidth and latency. 282When MLLMs are deployed online at scale, the vision and language models often block each other, thus incurring additional inference cost. This effect becomes more pronounced with larger vision models or higher-resolution images.283 284285 286As shown in the Figure above, we propose decoupled vision-language deployment (DvD) to address this issue by separating vision and language processing, with a particular focus on optimizing the prefilling stage. The vision subsystem batches and processes images to produce compact feature embeddings, which are then transmitted to the language subsystem for fusion with the text context prior to decoding. This separation alleviates blocking and brings multimodal prefilling performance closer to that of pure language models.287In our system implementation, the ViT and MLP (and ViR for InternVL3.5-Flash) are deployed on the vision server, while the language server executes only the LLM. The communication is unidirectional, transmitting BF16 visual features over TCP, with RDMA optionally employed to achieve higher transmission speed. Vision processing, feature transmission, and language processing are organized into an asynchronous three-stage pipeline, enabling overlapped execution and minimizing pipeline stalls.288 289 290DvD increases GPU utilization and processing efficiency on the vision side, while enabling the language server to focus exclusively on the LLM’s prefilling and decoding without being blocked by vision computation. This design leads to improved throughput and responsiveness. Moreover, the architecture supports independent hardware cost optimization for the vision and language modules, and facilitates the seamless integration of new modules without requiring modifications to the language server deployment.291 292 293## Evaluation on Multimodal Capability294 295### Multimodal Reasoning and Mathematics296 297298 299### OCR, Chart, and Document Understanding300 301302 303### Multi-Image Understanding & Real-World Comprehension304 305306 307### Comprehensive Multimodal Understanding & Multimodal Hallucination Evaluation308 309310 311### Visual Grounding312 313314 315### Multimodal Multilingual Understanding316 317318 319### Video Understanding320 321322 323### GUI Tasks324 325326 327### Embodied Tasks328 329330 331### SVG Tasks332 333334 335336 337## Evaluation on Language Capability338 339340 341## Ablation Study342 343### Cascade Reinforcement Learning344 345346 347348 349### Decoupled Vision-Language Deployment350 351 352353 354## Quick Start355 356We provide an example code to run `InternVL3.5-8B` using `transformers`. Please note that our models with up to 30B parameters can be deployed on a single A100 GPU, while the 38B model requires two A100 GPUs and the 235B model requires eight A100 GPUs.357 358> In most cases, both [LMDeploy](https://github.com/InternLM/lmdeploy) and [vLLM](https://github.com/vllm-project/vllm) can be used for model deployment. However, for InternVL3.5-20B-A4B, we recommend using vLLM since lmdeploy has not yet supported GPT-OSS.359 360> Please use transformers>=4.52.1 to ensure the model works normally. For the 20B version of our model, transformers>=4.55.0 is required.361 362### Model Loading363 364#### 16-bit (bf16 / fp16)365 366```python367import torch368from transformers import AutoTokenizer, AutoModel369path = "OpenGVLab/InternVL3_5-8B"370model = AutoModel.from_pretrained(371 path,372 torch_dtype=torch.bfloat16,373 low_cpu_mem_usage=True,374 use_flash_attn=True,375 trust_remote_code=True).eval().cuda()376```377 378#### BNB 8-bit Quantization379 380```python381import torch382from transformers import AutoTokenizer, AutoModel383path = "OpenGVLab/InternVL3_5-8B"384model = AutoModel.from_pretrained(385 path,386 torch_dtype=torch.bfloat16,387 load_in_8bit=True,388 low_cpu_mem_usage=True,389 use_flash_attn=True,390 trust_remote_code=True).eval()391```392 393#### Multiple GPUs394 395```python396import math397import torch398from transformers import AutoTokenizer, AutoModel399 400path = "OpenGVLab/InternVL3_5-8B"401model = AutoModel.from_pretrained(402 path,403 torch_dtype=torch.bfloat16,404 low_cpu_mem_usage=True,405 use_flash_attn=True,406 trust_remote_code=True,407 device_map="auto").eval()408```409 410### Thinking Mode411 412To enable thinking mode, please set the system prompt to our Thinking System Prompt. When enabling Thinking mode, we recommend setting `do_sample=True` and `temperature=0.6` to mitigate undesired repetition.413 414```python415R1_SYSTEM_PROMPT = """416You are an AI assistant that rigorously follows this response protocol:417 4181. First, conduct a detailed analysis of the question. Consider different angles, potential solutions, and reason through the problem step-by-step. Enclose this entire thinking process within <think> and </think> tags.419 4202. After the thinking section, provide a clear, concise, and direct answer to the user's question. Separate the answer from the think section with a newline.421 422Ensure that the thinking process is thorough but remains focused on the query. The final answer should be standalone and not reference the thinking section.423""".strip()424 425model.system_message = R1_SYSTEMP_PROMPT426```427 428### Inference with Transformers429 430```python431import math432import numpy as np433import torch434import torchvision.transforms as T435from decord import VideoReader, cpu436from PIL import Image437from torchvision.transforms.functional import InterpolationMode438from transformers import AutoModel, AutoTokenizer439 440IMAGENET_MEAN = (0.485, 0.456, 0.406)441IMAGENET_STD = (0.229, 0.224, 0.225)442 443def build_transform(input_size):444 MEAN, STD = IMAGENET_MEAN, IMAGENET_STD445 transform = T.Compose([446 T.Lambda(lambda img: img.convert('RGB') if img.mode != 'RGB' else img),447 T.Resize((input_size, input_size), interpolation=InterpolationMode.BICUBIC),448 T.ToTensor(),449 T.Normalize(mean=MEAN, std=STD)450 ])451 return transform452 453def find_closest_aspect_ratio(aspect_ratio, target_ratios, width, height, image_size):454 best_ratio_diff = float('inf')455 best_ratio = (1, 1)456 area = width * height457 for ratio in target_ratios:458 target_aspect_ratio = ratio[0] / ratio[1]459 ratio_diff = abs(aspect_ratio - target_aspect_ratio)460 if ratio_diff < best_ratio_diff:461 best_ratio_diff = ratio_diff462 best_ratio = ratio463 elif ratio_diff == best_ratio_diff:464 if area > 0.5 * image_size * image_size * ratio[0] * ratio[1]:465 best_ratio = ratio466 return best_ratio467 468def dynamic_preprocess(image, min_num=1, max_num=12, image_size=448, use_thumbnail=False):469 orig_width, orig_height = image.size470 aspect_ratio = orig_width / orig_height471 472 # calculate the existing image aspect ratio473 target_ratios = set(474 (i, j) for n in range(min_num, max_num + 1) for i in range(1, n + 1) for j in range(1, n + 1) if475 i * j <= max_num and i * j >= min_num)476 target_ratios = sorted(target_ratios, key=lambda x: x[0] * x[1])477 478 # find the closest aspect ratio to the target479 target_aspect_ratio = find_closest_aspect_ratio(480 aspect_ratio, target_ratios, orig_width, orig_height, image_size)481 482 # calculate the target width and height483 target_width = image_size * target_aspect_ratio[0]484 target_height = image_size * target_aspect_ratio[1]485 blocks = target_aspect_ratio[0] * target_aspect_ratio[1]486 487 # resize the image488 resized_img = image.resize((target_width, target_height))489 processed_images = []490 for i in range(blocks):491 box = (492 (i % (target_width // image_size)) * image_size,493 (i // (target_width // image_size)) * image_size,494 ((i % (target_width // image_size)) + 1) * image_size,495 ((i // (target_width // image_size)) + 1) * image_size496 )497 # split the image498 split_img = resized_img.crop(box)499 processed_images.append(split_img)500 assert len(processed_images) == blocks501 if use_thumbnail and len(processed_images) != 1:502 thumbnail_img = image.resize((image_size, image_size))503 processed_images.append(thumbnail_img)504 return processed_images505 506def load_image(image_file, input_size=448, max_num=12):507 image = Image.open(image_file).convert('RGB')508 transform = build_transform(input_size=input_size)509 images = dynamic_preprocess(image, image_size=input_size, use_thumbnail=True, max_num=max_num)510 pixel_values = [transform(image) for image in images]511 pixel_values = torch.stack(pixel_values)512 return pixel_values513 514path = 'OpenGVLab/InternVL3_5-8B'515model = AutoModel.from_pretrained(516 path,517 torch_dtype=torch.bfloat16,518 load_in_8bit=False,519 low_cpu_mem_usage=True,520 use_flash_attn=True,521 trust_remote_code=True,522 device_map="auto").eval()523tokenizer = AutoTokenizer.from_pretrained(path, trust_remote_code=True, use_fast=False)524 525# set the max number of tiles in `max_num`526pixel_values = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()527generation_config = dict(max_new_tokens=1024, do_sample=True)528 529# pure-text conversation (纯文本对话)530question = 'Hello, who are you?'531response, history = model.chat(tokenizer, None, question, generation_config, history=None, return_history=True)532print(f'User: {question}\nAssistant: {response}')533 534question = 'Can you tell me a story?'535response, history = model.chat(tokenizer, None, question, generation_config, history=history, return_history=True)536print(f'User: {question}\nAssistant: {response}')537 538# single-image single-round conversation (单图单轮对话)539question = '<image>\nPlease describe the image shortly.'540response = model.chat(tokenizer, pixel_values, question, generation_config)541print(f'User: {question}\nAssistant: {response}')542 543# single-image multi-round conversation (单图多轮对话)544question = '<image>\nPlease describe the image in detail.'545response, history = model.chat(tokenizer, pixel_values, question, generation_config, history=None, return_history=True)546print(f'User: {question}\nAssistant: {response}')547 548question = 'Please write a poem according to the image.'549response, history = model.chat(tokenizer, pixel_values, question, generation_config, history=history, return_history=True)550print(f'User: {question}\nAssistant: {response}')551 552# multi-image multi-round conversation, combined images (多图多轮对话,拼接图像)553pixel_values1 = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()554pixel_values2 = load_image('./examples/image2.jpg', max_num=12).to(torch.bfloat16).cuda()555pixel_values = torch.cat((pixel_values1, pixel_values2), dim=0)556 557question = '<image>\nDescribe the two images in detail.'558response, history = model.chat(tokenizer, pixel_values, question, generation_config,559 history=None, return_history=True)560print(f'User: {question}\nAssistant: {response}')561 562question = 'What are the similarities and differences between these two images.'563response, history = model.chat(tokenizer, pixel_values, question, generation_config,564 history=history, return_history=True)565print(f'User: {question}\nAssistant: {response}')566 567# multi-image multi-round conversation, separate images (多图多轮对话,独立图像)568pixel_values1 = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()569pixel_values2 = load_image('./examples/image2.jpg', max_num=12).to(torch.bfloat16).cuda()570pixel_values = torch.cat((pixel_values1, pixel_values2), dim=0)571num_patches_list = [pixel_values1.size(0), pixel_values2.size(0)]572 573question = 'Image-1: <image>\nImage-2: <image>\nDescribe the two images in detail.'574response, history = model.chat(tokenizer, pixel_values, question, generation_config,575 num_patches_list=num_patches_list,576 history=None, return_history=True)577print(f'User: {question}\nAssistant: {response}')578 579question = 'What are the similarities and differences between these two images.'580response, history = model.chat(tokenizer, pixel_values, question, generation_config,581 num_patches_list=num_patches_list,582 history=history, return_history=True)583print(f'User: {question}\nAssistant: {response}')584 585# batch inference, single image per sample (单图批处理)586pixel_values1 = load_image('./examples/image1.jpg', max_num=12).to(torch.bfloat16).cuda()587pixel_values2 = load_image('./examples/image2.jpg', max_num=12).to(torch.bfloat16).cuda()588num_patches_list = [pixel_values1.size(0), pixel_values2.size(0)]589pixel_values = torch.cat((pixel_values1, pixel_values2), dim=0)590 591questions = ['<image>\nDescribe the image in detail.'] * len(num_patches_list)592responses = model.batch_chat(tokenizer, pixel_values,593 num_patches_list=num_patches_list,594 questions=questions,595 generation_config=generation_config)596for question, response in zip(questions, responses):597 print(f'User: {question}\nAssistant: {response}')598 599# video multi-round conversation (视频多轮对话)600def get_index(bound, fps, max_frame, first_idx=0, num_segments=32):601 if bound:602 start, end = bound[0], bound[1]603 else:604 start, end = -100000, 100000605 start_idx = max(first_idx, round(start * fps))606 end_idx = min(round(end * fps), max_frame)607 seg_size = float(end_idx - start_idx) / num_segments608 frame_indices = np.array([609 int(start_idx + (seg_size / 2) + np.round(seg_size * idx))610 for idx in range(num_segments)611 ])612 return frame_indices613 614def load_video(video_path, bound=None, input_size=448, max_num=1, num_segments=32):615 vr = VideoReader(video_path, ctx=cpu(0), num_threads=1)616 max_frame = len(vr) - 1617 fps = float(vr.get_avg_fps())618 619 pixel_values_list, num_patches_list = [], []620 transform = build_transform(input_size=input_size)621 frame_indices = get_index(bound, fps, max_frame, first_idx=0, num_segments=num_segments)622 for frame_index in frame_indices:623 img = Image.fromarray(vr[frame_index].asnumpy()).convert('RGB')624 img = dynamic_preprocess(img, image_size=input_size, use_thumbnail=True, max_num=max_num)625 pixel_values = [transform(tile) for tile in img]626 pixel_values = torch.stack(pixel_values)627 num_patches_list.append(pixel_values.shape[0])628 pixel_values_list.append(pixel_values)629 pixel_values = torch.cat(pixel_values_list)630 return pixel_values, num_patches_list631 632video_path = './examples/red-panda.mp4'633pixel_values, num_patches_list = load_video(video_path, num_segments=8, max_num=1)634pixel_values = pixel_values.to(torch.bfloat16).cuda()635video_prefix = ''.join([f'Frame{i+1}: <image>\n' for i in range(len(num_patches_list))])636question = video_prefix + 'What is the red panda doing?'637# Frame1: <image>\nFrame2: <image>\n...\nFrame8: <image>\n{question}638response, history = model.chat(tokenizer, pixel_values, question, generation_config,639 num_patches_list=num_patches_list, history=None, return_history=True)640print(f'User: {question}\nAssistant: {response}')641 642question = 'Describe this video in detail.'643response, history = model.chat(tokenizer, pixel_values, question, generation_config,644 num_patches_list=num_patches_list, history=history, return_history=True)645print(f'User: {question}\nAssistant: {response}')646```647 648#### Streaming Output649 650Besides this method, you can also use the following code to get streamed output.651 652```python653from transformers import TextIteratorStreamer654from threading import Thread655 656# Initialize the streamer657streamer = TextIteratorStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True, timeout=10)658# Define the generation configuration659generation_config = dict(max_new_tokens=1024, do_sample=False, streamer=streamer)660# Start the model chat in a separate thread661thread = Thread(target=model.chat, kwargs=dict(662 tokenizer=tokenizer, pixel_values=pixel_values, question=question,663 history=None, return_history=False, generation_config=generation_config,664))665thread.start()666 667# Initialize an empty string to store the generated text668generated_text = ''669# Loop through the streamer to get the new text as it is generated670for new_text in streamer:671 if new_text == model.conv_template.sep:672 break673 generated_text += new_text674 print(new_text, end='', flush=True) # Print each new chunk of generated text on the same line675```676 677## Finetune678 679Many repositories now support fine-tuning of the InternVL series models, including [InternVL](https://github.com/OpenGVLab/InternVL), [SWIFT](https://github.com/modelscope/ms-swift), [XTuner](https://github.com/InternLM/xtuner), and others. Please refer to their documentation for more details on fine-tuning.680 681## Deployment682 683### LMDeploy684 685LMDeploy is a toolkit for compressing, deploying, and serving LLMs & VLMs.686 687```sh688pip install lmdeploy>=0.9.1689```690 691LMDeploy abstracts the complex inference process of multi-modal Vision-Language Models (VLM) into an easy-to-use pipeline, similar to the Large Language Model (LLM) inference pipeline.692 693#### A 'Hello, world' Example694 695```python696from lmdeploy import pipeline, PytorchEngineConfig697from lmdeploy.vl import load_image698 699image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/tests/data/tiger.jpeg')700 701# Please set tp=2 for the 38B version and tp=8 for the 241B-A28B version.702model = 'OpenGVLab/InternVL3_5-8B'703pipe = pipeline(model, backend_config=PytorchEngineConfig(session_len=32768, tp=1))704 705response = pipe(('describe this image', image))706print(response.text)707```708 709#### Multi-images Inference710 711When dealing with multiple images, you can put them all in one list. Keep in mind that multiple images will lead to a higher number of input tokens, and as a result, the size of the context window typically needs to be increased.712 713```python714from lmdeploy import pipeline, PytorchEngineConfig715from lmdeploy.vl import load_image716from lmdeploy.vl.constants import IMAGE_TOKEN717 718# Please set tp=2 for the 38B version and tp=8 for the 241B-A28B version.719model = 'OpenGVLab/InternVL3_5-8B'720pipe = pipeline(model, backend_config=PytorchEngineConfig(session_len=32768, tp=1))721 722image_urls=[723 'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg',724 'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/det.jpg'725]726 727images = [load_image(img_url) for img_url in image_urls]728# Numbering images improves multi-image conversations729response = pipe((f'Image-1: {IMAGE_TOKEN}\nImage-2: {IMAGE_TOKEN}\ndescribe these two images', images))730print(response.text)731```732 733#### Batch Prompts Inference734 735Conducting inference with batch prompts is quite straightforward; just place them within a list structure:736 737```python738from lmdeploy import pipeline, PytorchEngineConfig739from lmdeploy.vl import load_image740 741# Please set tp=2 for the 38B version and tp=8 for the 241B-A28B version.742model = 'OpenGVLab/InternVL3_5-8B'743pipe = pipeline(model, backend_config=PytorchEngineConfig(session_len=32768, tp=1))744 745image_urls=[746 "https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg",747 "https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/det.jpg"748]749prompts = [('describe this image', load_image(img_url)) for img_url in image_urls]750response = pipe(prompts)751print(response)752```753 754#### Multi-turn Conversation755 756There are two ways to do the multi-turn conversations with the pipeline. One is to construct messages according to the format of OpenAI and use above introduced method, the other is to use the `pipeline.chat` interface.757 758```python759from lmdeploy import pipeline, PytorchEngineConfig, GenerationConfig760from lmdeploy.vl import load_image761 762# Please set tp=2 for the 38B version and tp=8 for the 241B-A28B version.763model = 'OpenGVLab/InternVL3_5-8B'764pipe = pipeline(model, backend_config=PytorchEngineConfig(session_len=32768, tp=1))765 766image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg')767gen_config = GenerationConfig(top_k=50, top_p=0.95, temperature=0.6, max_new_tokens=8192)768sess = pipe.chat(('describe this image', image), gen_config=gen_config)769print(sess.response.text)770sess = pipe.chat('What is the woman doing?', session=sess, gen_config=gen_config)771print(sess.response.text)772```773 774#### Service775 776LMDeploy's `api_server` enables models to be easily packed into services with a single command. The provided RESTful APIs are compatible with OpenAI's interfaces. Below are an example of service startup:777 778```shell779lmdeploy serve api_server OpenGVLab/InternVL3_5-8B --server-port 23333 --tp 1 --backend pytorch780```781 782To use the OpenAI-style interface, you need to install OpenAI:783 784```shell785pip install openai786```787 788Then, use the code below to make the API call:789 790```python791from openai import OpenAI792 793client = OpenAI(api_key='YOUR_API_KEY', base_url='http://0.0.0.0:23333/v1')794model_name = client.models.list().data[0].id795response = client.chat.completions.create(796 model=model_name,797 messages=[{798 'role':799 'user',800 'content': [{801 'type': 'text',802 'text': 'describe this image',803 }, {804 'type': 'image_url',805 'image_url': {806 'url':807 'https://modelscope.oss-cn-beijing.aliyuncs.com/resource/tiger.jpeg',808 },809 }],810 }],811 temperature=0.8,812 top_p=0.8)813print(response)814```815 816## License817 818This project is released under the apache-2.0 License. This project uses the pre-trained Qwen3 as a component, which is licensed under the apache-2.0 License.819 820## Citation821 822If you find this project useful in your research, please consider citing:823 824```BibTeX825@article{wang2025internvl3_5,826 title={InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency},827 author={Wang, Weiyun and Gao, Zhangwei and Gu, Lixin and Pu, Hengjun and Cui, Long and Wei, Xingguang and Liu, Zhaoyang and Jing, Linglin and Ye, Shenglong and Shao, Jie and others},828 journal={arXiv preprint arXiv:2508.18265},829 year={2025}830}831```832 