CoolFace
Modelpublic

jucamohedano/qwen3-vl-4b-oven-grpo-traversal-wikidata-exact

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes8downloads
Model Card

Qwen3-VL-4B OVEN GRPO — traversal wikidata exact (LoRA)

LoRA adapter for `Qwen/Qwen3-VL-4B-Instruct`, trained with GRPO on the OVEN open-domain visual entity recognition task using verl.

Traversal in a Wikidata-style format under an exact-match reward (isolates the prompt-format axis).

Exploratory run from the training sweep, not reported individually in the thesis. Released to document the search over reward designs, prompt formats, and data curations.

Recipe

  • —Base model: Qwen/Qwen3-VL-4B-Instruct
  • —Adapter: LoRA, rank 64, alpha 32 (task_type: CAUSAL_LM)
  • —Algorithm: GRPO (verl), group size 16, seed 42, ~300 steps
  • —Prompt / reward: traversal, 'wikidata' format, exact-match reward. Shaped OVEN reward (format 0.05, exact/fuzzy 0.70, specificity-weighted hF 0.15, path-match 0.05, aggregation 0.05); a missing \boxed{} answer returns 0.00.

Training dynamics (from the wandb run)

metricstartendnotes
policy entropy0.1400.106min 0.075
KL to reference0.0860.834max 0.895
training reward0.0570.087—
response length31.20333.484tokens
val exact-match@10.1380.137held-out

Training reward moved 0.057 → 0.087 while held-out exact-match went 0.138 → 0.137 (essentially flat). Consistent with the thesis: outcome-only GRPO does not raise validation accuracy above the prompt-only elicitation floor. As a small-KL LoRA update the policy stays close to the reference model.

Usage

python
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoProcessor

base = AutoModelForImageTextToText.from_pretrained("Qwen/Qwen3-VL-4B-Instruct", trust_remote_code=True, device_map="auto")
model = PeftModel.from_pretrained(base, "jucamohedano/qwen3-vl-4b-oven-grpo-traversal-wikidata-exact")
processor = AutoProcessor.from_pretrained("Qwen/Qwen3-VL-4B-Instruct", trust_remote_code=True)

Intended use

Research artifact for reproducing the thesis's (largely negative) reinforcement-learning results on OVEN and for studying GRPO training dynamics on a multimodal open-world task. Not tuned or recommended for production classification.

Provenance

MSc thesis, University of Trento (Juan Camacho Mohedano). Training: verl (branch grpo-oven-v080); reward in verl/utils/reward_score/oven_boxed.py; data by oven-mllm-eval/scripts/build_verl_oven_parquet.py. wandb run: offline-run-20260704_140446-qwen3-vl-4b-oven-grpo-traversal-wikidata-exact-clean-seed42.