Myungkyu/qwen3_vl_8b_rmbench_taco_lora_3d5k
045
qwen3vl8brmbenchtacolora3d5k
High-level planner for RMBench (9 simulated tabletop tasks): Qwen/Qwen3-VL-8B-Instruct fine-tuned with LoRA on the per-tick planning records of `Myungkyu/RMBench-taco-gemini` — RMBench demonstrations with dense subtask labels from the task-specific context (offline annotation).
- Qwen3-VL-8B-Instruct, LoRA r=32 / alpha 64 / dropout 0 (Unsloth FastVisionModel), merged weights of checkpoint 3500
- Contract: 4 recent head-camera frames + task goal + carried memory text → JSON with the current subtask, the updated memory, a keyframe flag with caption and a retrieval query
- Optimizer batch 16; this repo holds the merged checkpoint at step 3500 only
