HoneyLane/vlmphysics-human-eval
0
physiq: re-encode GT to h264 720x480 (fixes black screen)
physicsiq: 198 pairs (real GT vs CogX-5B-I2V generated)
swap likephys -> physicsiq option
remove likephys_human (not a model-aware benchmark)
likephys: 96 pairs (valid vs invalid)
add likephys human-eval option
fix: prompt off-by-one (cogx5b)
fix: prompt off-by-one (wan14b)
cogx5b: re-shuffle A/B with seed=17 (different from wan14b seed=42)
fix: chooser links use explicit index.html
Initial deploy: human-eval pairwise UI
initial commit
