CoolFace
Datasetpublic

eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b

Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct) Raw rollout-level reward observations from the empirical Section of "Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis". This is the data that produced the headline Theorem 3 (correlation-dependent MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on GSM8K test prompts at temperature 0.7.… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes97downloads
settings

This repository belongs to eagle0504 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemultireward-grpo-gsm8k-rewards-qwen2.5-7b
visibilitypublic
licencecc-by-4.0
gatedno
ownereagle0504
Account settings
eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b · CoolFace