zhiqilicg/flowmap_grpo_sCM
Flow-Map GRPO checkpoints (FLUX.1-lite MeanFlow / sCM students)
LoRA adapters (r=64, alpha=128, 12 attention/FF projections) trained with Flow-Map GRPO on the T2I-Distill MeanFlow (flux1lite_meanflow.pt) and sCM (flux1lite_scm.pt) students of Freepik/flux.1-lite-8B, plus the online-SFT / Diffusion-NFT baselines. Code: ZhiqiLi-CG/flowmap_grpo (branch and commit in each run folder's README).
Layout: checkpoints/<family>/<task>/<method>[-<wandb run id>]/checkpoint-<step>/{lora, lora_ema, training_state}. lora/ is the raw learner, lora_ema/ the EMA adapter used for all reported numbers, training_state/ the full resume point. Run folders carry a README.md and the training-time full_eval_metrics.jsonl where available.
Evaluated Flow-Map GRPO runs (Genesis cluster; offline step sweeps in experiment_results.md, sections in the last column)
Other runs in this repo (longer Flow-Map GRPO runs and SFT / NFT baselines)
Evaluation protocol for the GRPO runs: K = sampling steps - 1 (K=4 is the 5-step training sampler), EMA adapter, OCR test set (1018 prompts), PickScore test set (2048), GenEval (553 x 4), DrawBench (200 prompts x 5) with PickScore / Aesthetic / DeQA / ImageReward / UnifiedReward; details and all numbers in the run READMEs and experiment_results.md.
