pantomiman/reason-over-search-eval-m5
Reason-over-Search M5: unified held-out evaluation (9 runs) Held-out 7-benchmark QA evaluation of the M5 reward-shape x seed ablation: Qwen3.5-0.8B GRPO-trained on MuSiQue with three reward shapes (F1-only / F1+format / EM-only) at three seeds (42 / 43 / 44). This repo is the uniform, light index of every checkpoint's eval scores; it carries the per-cell metric_score.txt + config.yaml and the per-run aggregated CSVs, in ONE consistent layout. (The heavy intermediate_data.json… See the full description on the dataset page: https://huggingface.co/datasets/pantomiman/reason-over-search-eval-m5.
This repository belongs to pantomiman on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
