namhokaist/appgen-qwen3-model-e-uedgrpo-awlegacy-randdense-v38-b200x4-s100-20260728-step100
AppGen Qwen3-VL Model E — Random-GRPO dense v38, step 100
This is the Hugging Face export of global_step_100/actor from the completed AppGen Random-GRPO dense-reward v38 run on four NVIDIA B200 GPUs.
Important training outcome
The run completed all 100 requested steps and saved checkpoints every 10 steps, but no actor optimization update was applied. All 100 update attempts were skipped because every training batch was marked score_only. A subsequent bitwise audit compared all 750 tensors (8,767,123,696 elements) and confirmed that this step-100 export is exactly equal to the pinned Model E base checkpoint. The held-out validation success rate was also unchanged at every evaluation: 11/193 (5.6995%).
This repository preserves the requested completed-run artifact and its full provenance. It should not be interpreted as an improved GRPO model.
Artifact identity
- Run:
uedgpo_q3_model_e_awlegacy_random_dense_v38_b200x4_s100_20260728 - Step:
100 - Base model:
namhokaist/appgen-qwen3-vl-8b-sft-ngc-amex-avariant-E-ngc-lr2p5e7-1ep - Base revision:
6cdf0aa413850771f9a6f4c4da38f53d9d060f1c - Training code:
grpo_ued_trainat750fce45b6dfb8550c352be950582a1bbccd654d - W&B run: https://wandb.ai/namhokoh-korea-advanced-institute-of-science-and-technology/ued-train-grpo/runs/q3e-awlegacy-v38-b200x4-randdense-s100-20260728
Random-GRPO contract
- Reward shaping: potential-based dense reward, delta emission
- Actor learning rate:
2e-7 - Group size:
8; train rows:8; PPO minibatch:64 - Maximum environment steps:
25 - Replay probability:
0.0; adaptive replay disabled - UCB: disabled
- Category sampling: target prior, hierarchical scope
- Total steps:
100; checkpoint/evaluation interval:10 - Prompt profile: AndroidWorld legacy, text-first, 10-turn history
- Coordinate space: normalized with denominator 999
Prompt and data provenance
- System prompt:
appgen_system_prompt.txt - Prompt bytes:
4,041 - Prompt SHA-256:
8795391af87a58a1670a33ae3f3e568e2af46899bc9b93fdc795c5f1fbbedfaf - Environment data:
luca0621/appgen-training-data@e96ee5f9f53b394a23ca1ce1912cb9ad40720842 - Requested random pool: 331 environments; 327 passed runtime validation
- Goal-text data:
luca0621/appgen-sft-data@c195ae15abd3d6aaa07f971b8d732e7a28fa8dbf - Goal-text artifact:
sft_qwen3_UNIFIED.json, 2,768 rows - Goal-text SHA-256:
0f6817d5c522d629437389aeec02d4d4f1fa88e7b10dd9770ad9d9569557f8bf
Final held-out validation
The final mean partial-credit test score was 0.3516881456426913.
Verification files
appgen_grpo_manifest.json: exact run, prompt, base-model, and training contractgrpo_publication_identity.json: identity of this GRPO publicationstep100_vs_base_weight_audit.json: zero-update and exact-weight-equality auditrun_manifest.json: inherited Model E SFT provenancepublication_identity.jsonandtraining_verification.json: inherited Model E publication records
Limitations
This artifact did not change the base model weights, so it should not be used to claim a Random-GRPO improvement. It is retained for reproducibility and diagnosis of the score-only training path. This model is for GUI-agent research and is not intended for safety-critical autonomous deployment.
