CoolFace
Datasetpublic

pantomiman/reason-over-search-eval-m5

Reason-over-Search M5: unified held-out evaluation (9 runs) Held-out 7-benchmark QA evaluation of the M5 reward-shape x seed ablation: Qwen3.5-0.8B GRPO-trained on MuSiQue with three reward shapes (F1-only / F1+format / EM-only) at three seeds (42 / 43 / 44). This repo is the uniform, light index of every checkpoint's eval scores; it carries the per-cell metric_score.txt + config.yaml and the per-run aggregated CSVs, in ONE consistent layout. (The heavy intermediate_data.json… See the full description on the dataset page: https://huggingface.co/datasets/pantomiman/reason-over-search-eval-m5.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes1.2kdownloads
2 commits on main
47f34134mo ago

M5 9-run unified held-out eval: metric_score + config + per-run CSVs (uniform layout)

pantomiman
4330d6b4mo ago

initial commit

pantomiman