CoolFace
Datasetpublic

dvyomkesh/nemo-grpo-weak3-from084-prompts

Nemo Weak-3 GRPO Prompt Dataset This dataset is a clean GRPO/RLVR prompt set for the three weak Nemotron challenge types identified after the 0.84 SDPO adapter diagnostics: bit_manipulation, unit_conversion, and gravity. The training rows are intentionally modeled as: prompt x + gold answer r + verifier/reward spec There are no source CoT traces, teacher completions, SDPO samples, RLSD privileged traces, or eval predictions in the training split. GRPO should sample completions… See the full description on the dataset page: https://huggingface.co/datasets/dvyomkesh/nemo-grpo-weak3-from084-prompts.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes20downloads
settings

This repository belongs to dvyomkesh on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namenemo-grpo-weak3-from084-prompts
visibilitypublic
licenceother
gatedno
ownerdvyomkesh
Account settings
dvyomkesh/nemo-grpo-weak3-from084-prompts · CoolFace