CoolFace
Modelpublic

nvidia/EGM-8B-SFT

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
5likes186downloads
Model Card

EGM-Qwen3-VL-8B-SFT

<p align="center"> <a href="https://nvlabs.github.io/EGM">[Project Page]</a> &nbsp; <a href="https://github.com/NVlabs/EGM">[Code]</a> &nbsp; </p>

<div align="center"> <img src="https://nvlabs.github.io/EGM/figure4.jpeg" width="90%"/> </div>

Model Summary

EGM-Qwen3-VL-8B-SFT is the supervised fine-tuning (SFT) checkpoint from the first stage of the EGM (Efficient Visual Grounding Language Models) training pipeline. It is built on top of Qwen3-VL-8B-Thinking.

This is an intermediate checkpoint intended for further reinforcement learning training. For the final model with best performance, see nvidia/EGM-8B.

Training Details

SFT Stage

In the SFT stage, a proprietary VLM generates detailed chain-of-thought reasoning steps for visual grounding training data. The base Qwen3-VL-8B-Thinking model is then fine-tuned on this reasoning-augmented data to learn structured visual grounding with explicit reasoning.

This SFT checkpoint serves as the initialization for the subsequent RL stage (GRPO), which yields the final EGM-8B model.

How to Use for RL Training

bash
pip install -U huggingface_hub
huggingface-cli download nvidia/EGM-8B-SFT --local-dir ./models/EGM-8B-SFT

Then follow the installation instructions in the EGM repository, prepare the RL data and start training:

bash
export BASE_DIR=$(pwd)
export MODEL_PATH="${BASE_DIR}/models/EGM-8B-SFT"
export OUTPUT_DIR="${BASE_DIR}/checkpoint/"
export DATA_DIR="${BASE_DIR}/data/EGM_Datasets/processed_rl_data/"

cd verl
bash scripts/grounding_qwen.sh

See the EGM repository for full RL training instructions.

Model Architecture

ComponentDetails
ArchitectureQwen3VLForConditionalGeneration
Precisionbfloat16
Text Hidden Size4096
Text Layers36
Attention Heads32 (8 KV heads)
Text Intermediate Size12,288
Vision Hidden Size1152
Vision Layers27
Patch Size16 x 16
Max Position Embeddings262,144
Vocabulary Size151,936

Related Models

ModelDescription
nvidia/EGM-8BFinal RL-trained model (best performance)
nvidia/EGM-4B-SFTSFT checkpoint for the 4B variant
nvidia/EGM-4BFinal RL-trained 4B model

Citation

bibtex
@article{zhan2026EGM,
    author = {Zhan, Guanqi and Li, Changye and Liu, Zhijian and Lu, Yao and Wu, Yi and Han, Song and Zhu, Ligeng},
    title = {EGM: Efficient Visual Grounding Language Models},
    booktitle = {arXiv},
    year = {2026}
}

Acknowledgment

This repository benefits from Qwen3-VL, InternVL, verl and verl-internvl.