CoolFace
Datasetpublic

MWilinski/hh-rlhf-harmless-base-rollouts-gpt-5.1-policy

Gemma Reward-Scored Rollouts Dataset Generation Parameters { "input": { "source_type": "hf", "repo_id": "MWilinski/hh-rlhf-harmless-base-rollouts-gpt-5.1-policy", "split": "train", "path": "" }, "base_dataset": { "id": "MWilinski/hh-rlhf-harmless-base", "split": "train", "prompt_field": "prompt" }, "selection": { "prompt_indices": [] }, "scoring": { "backend": "openrouter", "model": "google/gemma-3-27b-it"… See the full description on the dataset page: https://huggingface.co/datasets/MWilinski/hh-rlhf-harmless-base-rollouts-gpt-5.1-policy.

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes20downloads

MWilinski/hh-rlhf-harmless-base-rollouts-gpt-5.1-policy · main · files are served by the source, never re-hosted here