JamesK2W/gamevla
0
gamevla — CS2 VLA / behavioral-cloning checkpoints
Checkpoints for real-environment testing. One folder per experiment; each folder has config.yaml, dataset_statistics.json, the selected checkpoint(s), and an EVAL.md with the exact run command. For experiments with more than one saved checkpoint we keep the final step and one intermediate step; single-checkpoint experiments keep the only one.
▶ To run these in the CS2 environment, see [EVALUATION.md](EVALUATION.md) (score = mean total episode reward) and each folder's EVAL.md.
aim_cos = eval mouse-aim cosine similarity (the aim quality metric; ~0 means aim did not learn). key_f1 = eval h0 key macro-F1. attack_f1 = eval firing F1.
Notes:
- ⚠️ `aimflow` (v1) aim_cos ≈ 0.475 is a metric artifact, not real aim: v1 used a different mouse normalization, so its aim_cos is NOT comparable to v2/v3 (which read ~0). Across all experiments here, aim is not convincingly learned — treat it as an open question to verify in the real environment.
csbcpeaks early (step 5000) then overfits on keys;steps_15000is more trained but slightly worse.nitrogen_starvla/final_model.ptis the only fully-completed (60k-step) run.
