ZoeyZou/grm-reproduce-align30k-beavertails-v
Fused GRM reproduction data This is the self-contained Stage-1 preference dataset used by GRM-Reproduce to reproduce the generative reward model (GRM) stage of Generative RLHF-V. It is an independent reproduction artifact, not an official dataset release by the paper authors. The release fuses exactly 30,000 Align-Anything pairs and 9,369 BeaverTails-V pairs. Every row already follows the final VERL Stage-1 schema produced by grlhfv_repro.data.to_verl_preference_example, and… See the full description on the dataset page: https://huggingface.co/datasets/ZoeyZou/grm-reproduce-align30k-beavertails-v.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face