datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-RLHF-GenRM-v1
Dataset Description:
This dataset is designed to train Generative Reward Models (GenRMs). It leverages reinforcement learning at scale to train accurate and robust GenRMs that generalize better than traditional Bradley-Terry models and reduce the risk of reward hacking.
The dataset is composed of:
Preference data focused on diverse domains
A synthetic safety blend
The data follows a "meta-prompt" structure where the model is instructed to act as an expert evaluation judge. For… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RLHF-GenRM-v1.hh-rlhf-strength-cleaned
Dataset Card for hh-rlhf-strength-cleaned
Other Language Versions: English, 中文.
Dataset Description
In the paper titled "Secrets of RLHF in Large Language Models Part II: Reward Modeling" we measured the preference strength of each preference pair in the hh-rlhf dataset through model ensemble and annotated the valid set with GPT-4. In this repository, we provide:
Metadata of preference strength for both the training and valid sets.
GPT-4 annotations on the valid set.
We… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/hh-rlhf-strength-cleaned.hh-rlhf
Dataset Card for HH-RLHF
Dataset Summary
This repository provides access to two different kinds of data:
Human preference data about helpfulness and harmlessness from Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. These data are meant to train preference (or reward) models for subsequent RLHF training. These data are not meant for supervised training of dialogue agents. Training dialogue agents on these data is likely to lead… See the full description on the dataset page: https://huggingface.co/datasets/polinaeterna/hh-rlhf.korean_rlhf_content_filtered
Korean RLHF Content Filtered
Dataset Summary
This dataset is a cleaned, content-only derivative of:
Source dataset: jojo0217/korean_rlhf_dataset
Source URL: https://huggingface.co/datasets/jojo0217/korean_rlhf_dataset
Each row has a single content field suitable for LM pretraining/SFT-style text modeling.
Construction
Content construction rule
For each source row:
If input is empty: content = instruction + "\n" + output
If input is not empty:… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/korean_rlhf_content_filtered.WiredBrain-RLHF
WiredBrain-RLHF: Production-Grade Data for High-Integrity AI Alignment
Overview
WiredBrain-RLHF transforms the foundational Anthropic HH-RLHF dataset (148K conversations) into a production-grade training resource through systematic data enrichment. Engineered by Shubham Dev (Department of Computer Science, Jaypee University of Information Technology, India), this dataset powers high-stakes AI alignment research requiring factual integrity and entity preservation.… See the full description on the dataset page: https://huggingface.co/datasets/pheonix-delta/WiredBrain-RLHF.static-rlhf-interface-dataRLHFlow__ArmoRM-Llama3-8B-v0.1-details
Dataset Card for Evaluation run of RLHFlow/ArmoRM-Llama3-8B-v0.1
Dataset automatically created during the evaluation run of model RLHFlow/ArmoRM-Llama3-8B-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/RLHFlow__ArmoRM-Llama3-8B-v0.1-details.pku-safe-rlhf-125-eval-bundle
PKU SafeRLHF 125-Subset Evaluation Bundle
This bundle contains the code and prepared data used for the 125-sample
evaluation subset under:
figure6_eval728_without_train_overlap_125_repro_seed0
It is intended as a portable reproduction package for:
generation on the fixed 125-prompt subset
pairwise helpfulness / harmlessness judging
single-model humanness judging
the newer mark-error humanness judging
Layout
code/: generation, judging, merge, and cluster… See the full description on the dataset page: https://huggingface.co/datasets/sheng22213/pku-safe-rlhf-125-eval-bundle.rlhf_dialogue_relevance_startrek_capitan_kirk
Citation
@misc{luo_tinyxgpt,
title = {TinyXGPT: A Minimal End-to-End State-Space-Attention Large Language Model},
author = {R\'ois\'in Luo},
url = {https://github.com/roisincrtai/tinyxgpt}
}
jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-details
Dataset Card for Evaluation run of jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model
Dataset automatically created during the evaluation run of model jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-details.tag_system_RLHFRLHFlow__LLaMA3-iterative-DPO-final-details
Dataset Card for Evaluation run of RLHFlow/LLaMA3-iterative-DPO-final
Dataset automatically created during the evaluation run of model RLHFlow/LLaMA3-iterative-DPO-final
The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/RLHFlow__LLaMA3-iterative-DPO-final-details.
