CoolFace
Datasetpublic

ZoeyZou/grm-reproduce-align30k-beavertails-v

Fused GRM reproduction data This is the self-contained Stage-1 preference dataset used by GRM-Reproduce to reproduce the generative reward model (GRM) stage of Generative RLHF-V. It is an independent reproduction artifact, not an official dataset release by the paper authors. The release fuses exactly 30,000 Align-Anything pairs and 9,369 BeaverTails-V pairs. Every row already follows the final VERL Stage-1 schema produced by grlhfv_repro.data.to_verl_preference_example, and… See the full description on the dataset page: https://huggingface.co/datasets/ZoeyZou/grm-reproduce-align30k-beavertails-v.

sourceHugging Facecc-by-nc-4.0updated 18d agoView on Hugging Face
0likes276downloads
settings

This repository belongs to ZoeyZou on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namegrm-reproduce-align30k-beavertails-v
visibilitypublic
licencecc-by-nc-4.0
gatedno
ownerZoeyZou
Account settings
ZoeyZou/grm-reproduce-align30k-beavertails-v · CoolFace