GameWorldData/HUD-UI-Validation-Pilot
HUD/UI annotation validation pilot This public artifact compares 10 gameplay screenshots across four columns: raw frame; GPT-5.6 Sol X-High final reference; GPT-5.6 Terra High refined annotation; confidence-routed final annotation. Files: analysis.md: aggregate metrics and per-game error table; per_sample_metrics.csv: machine-readable sample metrics; overlay_comparison_contact_sheet.jpg: full comparison sheet. The 0--100 confidence value is a conservative pipeline routing… See the full description on the dataset page: https://huggingface.co/datasets/GameWorldData/HUD-UI-Validation-Pilot.
Upload formal 100-sample exploratory overlay_comparison_contact_sheet.jpg
Upload formal 100-sample exploratory per_sample_metrics.csv
Upload formal 100-sample exploratory analysis.md
Add four formal Ground Truth human-review bundles
Add human review artifact human_review/31bea7df6706952cf58b1ff7/confidence.json
Add human review artifact human_review/31bea7df6706952cf58b1ff7/round3_adjudicated.json
Add human review artifact human_review/31bea7df6706952cf58b1ff7/round2_refined.json
Add human review artifact human_review/31bea7df6706952cf58b1ff7/round3_overlay.jpg
Add human review artifact human_review/31bea7df6706952cf58b1ff7/round2_overlay.jpg
Add human review artifact human_review/31bea7df6706952cf58b1ff7/raw.jpg
Add human review artifact human_review/31bea7df6706952cf58b1ff7/README.md
Upload overlay_comparison_contact_sheet.jpg
Upload per_sample_metrics.csv
Upload analysis.md
Upload README.md
initial commit
