t2ance/atlas-29-supergpqa-records-as-a-candidate-pool
29. Is SuperGPQA-Records a candidate pool an orchestrator can learn from? 1. Question and links Does a candidate pool built from SuperGPQA-Records (128 hard questions, eight candidates each from the eight strongest single-shot models of the records) leave room for an untrained orchestrator, Qwen3.6-27B or Qwen3.5-9B, to learn: how does each model's greedy accuracy move from no candidate to eight against the pool's coverage? Report source:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-29-supergpqa-records-as-a-candidate-pool.
29. Is SuperGPQA-Records a candidate pool an orchestrator can learn from?
1. Question and links
Does a candidate pool built from SuperGPQA-Records (128 hard questions, eight candidates each from the eight strongest single-shot models of the records) leave room for an untrained orchestrator, Qwen3.6-27B or Qwen3.5-9B, to learn: how does each model's greedy accuracy move from no candidate to eight against the pool's coverage? Report source: .claude/skills/atlas-experimenting/experiments/rl-training/29-supergpqa-records-as-a-candidate-pool/main.tex in t2ance/ATLAS. No W&B run: the probe is inference only. Status: closed 2026-09-14. The probe ran 2026-09-13 19:02 to 20:11 UTC on server GPU3 (640 states a model); the rerun of the truncated no-candidate states at a 65,536-token context ran 22:09 to 23:58 UTC the same day (83 states for the 27B, 47 for the 9B). Judged with a truncated state counted wrong and the rerun standing in for the truncated solo states: the candidates lift the 27B from 0.555 to 0.641 and the 9B from 0.523 to 0.594 on the 128 questions; both read the candidates as a vote; the 9B is the model to train.
Read this first. The reading copy of this directory is t2ance/atlas-experiments under 29-supergpqa-records-as-a-candidate-pool/; the Hub repository holds the same files. No saved step and no per-token array exists here.
2. Directory tree
README.md: this front page.data/probe_questions.jsonl: the 128 questions with their eight candidates in order, written byprobe.py prepare(fields: uuid, question, options, answer_letter, discipline, field, subfield, candidates[model, answer, reasoning, status, correct, answered]).probe/<served model>.jsonl: one row per (question, k) state written byprobe.py run: the answer, its correctness, the finish reason, token counts, seconds, the JSON content and the think text.probe/<served model>.k0-65536.shard0of1.jsonl: the rerun of every k = 0 state that hit the 16,384-token budget, written byprobe.py rerunat a 65,536-token context (same fields, pluscontext); 83 rows for the 27B, 47 for the 9B.probe/summary.md: the tables written byprobe.py summarize, with a rerun section per model (states ended, states still at the budget, correct, B_0 over all 128 with the rerun in place).logs/: the probe's two servers and two launchers (serve_*.log,launch_*.log), the rerun's (rerun_serve_*.log,rerun_*.log), andgpu_watch/(the machine layer's samples and verdicts across both runs).
3. Inputs
- SuperGPQA (
m-a-p/SuperGPQA,SuperGPQA-all.jsonl) and SuperGPQA-Records (m-a-p/SuperGPQA-Records,zero-shot/<model>_SuperGPQA-all_zero-shot.jsonlfor the eight models named inprobe.py), downloaded 2026-09-13 toExperiment/datasets/hub/on the home host; not copied here. - Qwen3.6-27B: report 15's checkpoint copy (
15-selector-capacity-vs-training/checkpoints/Qwen3.6-27B, Hub snapshot 6a9e13b). Qwen3.5-9B: Hub snapshot c202236.
4. How to read one file
probe/qwen3_6_27b_base.jsonl: each line is a JSON object; k is the number of candidates shown, correct compares answer with the question's answer_letter in data/probe_questions.jsonl by uuid. python probe.py summarize in the report folder rebuilds probe/summary.md from these files.
5. Shared weights
None.
