JamesK2W/viewagent-ai2thor-qwen25vl7b-ivp
06
ViewAgent — Qwen2.5-VL-7B trained on AI2-THOR Interactive View Planning (IVP)
Qwen2.5-VL-7B-Instruct trained with the GraphRL pipeline (verl RL ↔ view-graph SFT distillation) on the AI2-THOR port of the ViewSuite proxy-task suite.
Test results (AI2-THOR test split, n=252)
IVP (multi-turn active exploration): 61.1% — +57 over base, +25 over the best frontier zero-shot model (GPT-5.4). IVP-only RL trades some single-turn MCQ accuracy (P2V 40.9→25.0). Checkpoint = iter3 RL globalstep460.
- Code: https://github.com/JamesKrW/ViewAgent (branch
feat/ai2thor-env) - Dataset: https://huggingface.co/datasets/JamesK2W/viewagent-ai2thor
