CoolFace
Modelpublic

namhokaist/appgen-qwen3-model-e-uedgrpo-awlegacy-randdense-v38-b200x4-s100-20260728-step100

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes12downloads
Model Card

AppGen Qwen3-VL Model E — Random-GRPO dense v38, step 100

This is the Hugging Face export of global_step_100/actor from the completed AppGen Random-GRPO dense-reward v38 run on four NVIDIA B200 GPUs.

Important training outcome

The run completed all 100 requested steps and saved checkpoints every 10 steps, but no actor optimization update was applied. All 100 update attempts were skipped because every training batch was marked score_only. A subsequent bitwise audit compared all 750 tensors (8,767,123,696 elements) and confirmed that this step-100 export is exactly equal to the pinned Model E base checkpoint. The held-out validation success rate was also unchanged at every evaluation: 11/193 (5.6995%).

This repository preserves the requested completed-run artifact and its full provenance. It should not be interpreted as an improved GRPO model.

Artifact identity

  • —Run: uedgpo_q3_model_e_awlegacy_random_dense_v38_b200x4_s100_20260728
  • —Step: 100
  • —Base model: namhokaist/appgen-qwen3-vl-8b-sft-ngc-amex-avariant-E-ngc-lr2p5e7-1ep
  • —Base revision: 6cdf0aa413850771f9a6f4c4da38f53d9d060f1c
  • —Training code: grpo_ued_train at 750fce45b6dfb8550c352be950582a1bbccd654d
  • —W&B run: https://wandb.ai/namhokoh-korea-advanced-institute-of-science-and-technology/ued-train-grpo/runs/q3e-awlegacy-v38-b200x4-randdense-s100-20260728

Random-GRPO contract

  • —Reward shaping: potential-based dense reward, delta emission
  • —Actor learning rate: 2e-7
  • —Group size: 8; train rows: 8; PPO minibatch: 64
  • —Maximum environment steps: 25
  • —Replay probability: 0.0; adaptive replay disabled
  • —UCB: disabled
  • —Category sampling: target prior, hierarchical scope
  • —Total steps: 100; checkpoint/evaluation interval: 10
  • —Prompt profile: AndroidWorld legacy, text-first, 10-turn history
  • —Coordinate space: normalized with denominator 999

Prompt and data provenance

  • —System prompt: appgen_system_prompt.txt
  • —Prompt bytes: 4,041
  • —Prompt SHA-256: 8795391af87a58a1670a33ae3f3e568e2af46899bc9b93fdc795c5f1fbbedfaf
  • —Environment data: luca0621/appgen-training-data@e96ee5f9f53b394a23ca1ce1912cb9ad40720842
  • —Requested random pool: 331 environments; 327 passed runtime validation
  • —Goal-text data: luca0621/appgen-sft-data@c195ae15abd3d6aaa07f971b8d732e7a28fa8dbf
  • —Goal-text artifact: sft_qwen3_UNIFIED.json, 2,768 rows
  • —Goal-text SHA-256: 0f6817d5c522d629437389aeec02d4d4f1fa88e7b10dd9770ad9d9569557f8bf

Final held-out validation

CategorySuccess / tasksRate
communication0 / 280.00%
finance1 / 137.69%
food and drink1 / 175.88%
health and fitness0 / 180.00%
maps and navigation2 / 219.52%
productivity1 / 224.55%
shopping1 / 119.09%
social1 / 156.67%
tools3 / 2015.00%
travel and local1 / 283.57%
pooled11 / 1935.70%

The final mean partial-credit test score was 0.3516881456426913.

Verification files

  • —appgen_grpo_manifest.json: exact run, prompt, base-model, and training contract
  • —grpo_publication_identity.json: identity of this GRPO publication
  • —step100_vs_base_weight_audit.json: zero-update and exact-weight-equality audit
  • —run_manifest.json: inherited Model E SFT provenance
  • —publication_identity.json and training_verification.json: inherited Model E publication records

Limitations

This artifact did not change the base model weights, so it should not be used to claim a Random-GRPO improvement. It is retained for reproducibility and diagnosis of the score-only training path. This model is for GUI-agent research and is not intended for safety-critical autonomous deployment.