ZhidongWang/seerdrive-pdm-lite
Upload post_train_round1/metadata.json with huggingface_hub
Upload post_train_round1/policy_accepted12_ep9s1100.pt with huggingface_hub
Stage A full run stagea_full_0813 ep30 (v/q=0.10/0.25, value v4, R-19 stuck-frame gate)
full_f2_sft30_0712 SFT ep30 (D-37 F2 consequence-conditioned scoring, pure SFT, no RL stage)
CR-2 full-set two-stage ckpt: SFT30 + GRPO-RL5 (ep35, 0708 run full_cr2_sft30rl5_0707)
full_r1d29_0706 RL ep35: R1(R1-3b)+R2(p3)+D-29 statics; val geo_ade 0.144 fde 0.154 acc 0.928 IoU_veh 0.336
full-set GRPO ms3 run full_strat955_ms3_sft30rl5_0704: SFT30+RL5, 8xGPU bs8, MULTI-STEP WM sim_speed_steps=3 (0.5/1.0/1.5s), stratified 95:5 split (train 561088 / val 33451), 24.5M params. val ep35: geo_ade=0.144 geo_fde=0.142 speed_acc=0.936 speed_l1=0.606 mIoU=0.505 IoU_veh=0.357. deploy model-only ckpt.
full-set GRPO run full_strat955_sft30rl5_0704: SFT30+RL5, 8xGPU bs8, stratified 95:5 route split (train 561088 / val 33451). val ep35: geo_ade=0.145 geo_fde=0.182 speed_acc=0.919 speed_l1=0.694 mIoU=0.491 IoU_veh=0.357. deploy model-only ckpt (21.1M params).
Upload latest checkpoint (D-24/D-25 run mkutw4j7, ~ep33, full state)
Add README
D-23 GT-nearest retrain checkpoint (epoch 50)
initial commit
