zcahjl3/figmirror-paper-derived-pilot
Paper-derived Oracle Figures — Pilot v7 (agent-judged, 2026-07-14) Every pick vetted by an independent Haiku 4.5 vision-agent on two axes: benchmark_worthy — is the SOURCE non-trivial (challenge a 7B code-gen model)? fidelity_pass — does REPRODUCED faithfully echo SOURCE? Only picks where BOTH = YES ship. Stats 50 picks (28 ML + 22 SCI) 20 distinct subtypes; 12/12 majors 13 distinct venues Per-pick files <domain>/<safe-id>/: source.png… See the full description on the dataset page: https://huggingface.co/datasets/zcahjl3/figmirror-paper-derived-pilot.
v7: 50 picks (28 ML + 22 SCI), all agent-judged by Haiku 4.5 vision on benchmark_worthy + fidelity_pass
v5: 47 picks (31 ML + 16 SCI), +7 from incremental D4 round
Add paper-derived oracle figures pilot v4 (40 picks, 33/17 ML/SCI)
initial commit
