zcahjl3/figmirror-paper-derived-pilot
Paper-derived Oracle Figures — Pilot v7 (agent-judged, 2026-07-14) Every pick vetted by an independent Haiku 4.5 vision-agent on two axes: benchmark_worthy — is the SOURCE non-trivial (challenge a 7B code-gen model)? fidelity_pass — does REPRODUCED faithfully echo SOURCE? Only picks where BOTH = YES ship. Stats 50 picks (28 ML + 22 SCI) 20 distinct subtypes; 12/12 majors 13 distinct venues Per-pick files <domain>/<safe-id>/: source.png… See the full description on the dataset page: https://huggingface.co/datasets/zcahjl3/figmirror-paper-derived-pilot.
Paper-derived Oracle Figures — Pilot v7 (agent-judged, 2026-07-14)
Every pick vetted by an independent Haiku 4.5 vision-agent on two axes:
- benchmark_worthy — is the SOURCE non-trivial (challenge a 7B code-gen model)?
- fidelity_pass — does REPRODUCED faithfully echo SOURCE? Only picks where BOTH = YES ship.
Stats
- 50 picks (28 ML + 22 SCI)
- 20 distinct subtypes; 12/12 majors
- 13 distinct venues
Per-pick files
<domain>/<safe-id>/: source.png, reproduced.png, code.py, data.{csv|npz}, dataecho.md, caption.txt, metadata.json (with `agentjudge`).
Related docs
- Spec:
docs/superpowers/specs/2026-05-24-paper-derived-oracle-figures-design.md - Gate audit:
docs/superpowers/specs/2026-07-14-evaluation-gates-audit.md
