GameWorldData/HUD-UI-Validation-Pilot
HUD/UI annotation validation pilot This public artifact compares 10 gameplay screenshots across four columns: raw frame; GPT-5.6 Sol X-High final reference; GPT-5.6 Terra High refined annotation; confidence-routed final annotation. Files: analysis.md: aggregate metrics and per-game error table; per_sample_metrics.csv: machine-readable sample metrics; overlay_comparison_contact_sheet.jpg: full comparison sheet. The 0--100 confidence value is a conservative pipeline routing… See the full description on the dataset page: https://huggingface.co/datasets/GameWorldData/HUD-UI-Validation-Pilot.
HUD/UI annotation validation pilot
This public artifact compares 10 gameplay screenshots across four columns:
- raw frame;
- GPT-5.6 Sol X-High final reference;
- GPT-5.6 Terra High refined annotation;
- confidence-routed final annotation.
Files:
- `analysis.md`: aggregate metrics and per-game error table;
- `per_sample_metrics.csv`: machine-readable sample metrics;
- `overlay_comparison_contact_sheet.jpg`: full comparison sheet.
The 0--100 confidence value is a conservative pipeline routing score, not a calibrated probability of correctness. On this 10-sample pilot, pure Terra High element F1 is 0.7755 and mean slot Jaccard is 0.6875. The sample is too small for a final model-selection conclusion.
