sheng22213/pku-safe-rlhf-125-eval-bundle
PKU SafeRLHF 125-Subset Evaluation Bundle This bundle contains the code and prepared data used for the 125-sample evaluation subset under: figure6_eval728_without_train_overlap_125_repro_seed0 It is intended as a portable reproduction package for: generation on the fixed 125-prompt subset pairwise helpfulness / harmlessness judging single-model humanness judging the newer mark-error humanness judging Layout code/: generation, judging, merge, and cluster… See the full description on the dataset page: https://huggingface.co/datasets/sheng22213/pku-safe-rlhf-125-eval-bundle.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face