t2ance/atlas-30-openmathreasoning-genselect-training
30. Does Qwen3.5-9B learn to use candidates on OpenMathReasoning? 1. Question and links Trained by reinforcement learning on rows of eight OpenMathReasoning GenSelect candidates with one to seven of them correct, does Qwen3.5-9B's accuracy with eight candidates on held-out rows of the same distribution rise above its untrained accuracy and above the majority vote of the eight? Report source:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-30-openmathreasoning-genselect-training.
04.3k
This repository belongs to t2ance on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
atlas-30-openmathreasoning-genselect-training
public
not set
no
t2ance
