CoolFace
Modelpublic

JamesK2W/viewagent-ai2thor-qwen25vl7b-ivp

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes6downloads
Model Card

ViewAgent — Qwen2.5-VL-7B trained on AI2-THOR Interactive View Planning (IVP)

Qwen2.5-VL-7B-Instruct trained with the GraphRL pipeline (verl RL ↔ view-graph SFT distillation) on the AI2-THOR port of the ViewSuite proxy-task suite.

Test results (AI2-THOR test split, n=252)

ModelP2VV2P**IVP**Overall
Qwen2.5-VL-7B (base)40.927.44.024.1
This model (trained)25.041.361.142.5
GPT-5.4 (zero-shot)73.869.436.159.8

IVP (multi-turn active exploration): 61.1% — +57 over base, +25 over the best frontier zero-shot model (GPT-5.4). IVP-only RL trades some single-turn MCQ accuracy (P2V 40.9→25.0). Checkpoint = iter3 RL globalstep460.

  • —Code: https://github.com/JamesKrW/ViewAgent (branch feat/ai2thor-env)
  • —Dataset: https://huggingface.co/datasets/JamesK2W/viewagent-ai2thor