CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-RLHF-GenRM-v1 Dataset Description: This dataset is designed to train Generative Reward Models (GenRMs). It leverages reinforcement learning at scale to train accurate and robust GenRMs that generalize better than traditional Bradley-Terry models and reduce the risk of reward hacking. The dataset is composed of: Preference data focused on diverse domains A synthetic safety blend The data follows a "meta-prompt" structure where the model is instructed to act as an expert evaluation judge. For… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RLHF-GenRM-v1.tabularreinforcement-learning100K<n<1M5 likes186 downloads7mo agoHugging Face02OpenMOSS-Team /hh-rlhf-strength-cleaned Dataset Card for hh-rlhf-strength-cleaned Other Language Versions: English, 中文. Dataset Description In the paper titled "Secrets of RLHF in Large Language Models Part II: Reward Modeling" we measured the preference strength of each preference pair in the hh-rlhf dataset through model ensemble and annotated the valid set with GPT-4. In this repository, we provide: Metadata of preference strength for both the training and valid sets. GPT-4 annotations on the valid set. We… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/hh-rlhf-strength-cleaned.tabular100K<n<1M26 likes130 downloads3y agoHugging Face03polinaeterna /hh-rlhf Dataset Card for HH-RLHF Dataset Summary This repository provides access to two different kinds of data: Human preference data about helpfulness and harmlessness from Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. These data are meant to train preference (or reward) models for subsequent RLHF training. These data are not meant for supervised training of dialogue agents. Training dialogue agents on these data is likely to lead… See the full description on the dataset page: https://huggingface.co/datasets/polinaeterna/hh-rlhf.tabular100K<n<1M1 likes114 downloads3y agoHugging Face04alwaysgood /korean_rlhf_content_filtered Korean RLHF Content Filtered Dataset Summary This dataset is a cleaned, content-only derivative of: Source dataset: jojo0217/korean_rlhf_dataset Source URL: https://huggingface.co/datasets/jojo0217/korean_rlhf_dataset Each row has a single content field suitable for LM pretraining/SFT-style text modeling. Construction Content construction rule For each source row: If input is empty: content = instruction + "\n" + output If input is not empty:… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/korean_rlhf_content_filtered.tabulartext-generation100K<n<1M0 likes50 downloads6mo agoHugging Face05pheonix-delta /WiredBrain-RLHF WiredBrain-RLHF: Production-Grade Data for High-Integrity AI Alignment Overview WiredBrain-RLHF transforms the foundational Anthropic HH-RLHF dataset (148K conversations) into a production-grade training resource through systematic data enrichment. Engineered by Shubham Dev (Department of Computer Science, Jaypee University of Information Technology, India), this dataset powers high-stakes AI alignment research requiring factual integrity and entity preservation.… See the full description on the dataset page: https://huggingface.co/datasets/pheonix-delta/WiredBrain-RLHF.tabularreinforcement-learning100K<n<1M1 likes28 downloads7mo agoHugging Face06Tristan /static-rlhf-interface-datatabularn<1K0 likes16 downloads3y agoHugging Face07open-llm-leaderboard /RLHFlow__ArmoRM-Llama3-8B-v0.1-detailsgated Dataset Card for Evaluation run of RLHFlow/ArmoRM-Llama3-8B-v0.1 Dataset automatically created during the evaluation run of model RLHFlow/ArmoRM-Llama3-8B-v0.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/RLHFlow__ArmoRM-Llama3-8B-v0.1-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face08sheng22213 /pku-safe-rlhf-125-eval-bundle PKU SafeRLHF 125-Subset Evaluation Bundle This bundle contains the code and prepared data used for the 125-sample evaluation subset under: figure6_eval728_without_train_overlap_125_repro_seed0 It is intended as a portable reproduction package for: generation on the fixed 125-prompt subset pairwise helpfulness / harmlessness judging single-model humanness judging the newer mark-error humanness judging Layout code/: generation, judging, merge, and cluster… See the full description on the dataset page: https://huggingface.co/datasets/sheng22213/pku-safe-rlhf-125-eval-bundle.tabulartext-generationn<1K0 likes11 downloads2mo agoHugging Face09roisincrtai /rlhf_dialogue_relevance_startrek_capitan_kirk Citation @misc{luo_tinyxgpt, title = {TinyXGPT: A Minimal End-to-End State-Space-Attention Large Language Model}, author = {R\'ois\'in Luo}, url = {https://github.com/roisincrtai/tinyxgpt} } tabularn<1K0 likes11 downloads2mo agoHugging Face10open-llm-leaderboard /jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-detailsgated Dataset Card for Evaluation run of jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model Dataset automatically created during the evaluation run of model jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-details.tabular10K<n<100K0 likes9 downloads2y agoHugging Face11hsy0202 /tag_system_RLHFtabularn<1K0 likes6 downloads2y agoHugging Face12open-llm-leaderboard /RLHFlow__LLaMA3-iterative-DPO-final-detailsgated Dataset Card for Evaluation run of RLHFlow/LLaMA3-iterative-DPO-final Dataset automatically created during the evaluation run of model RLHFlow/LLaMA3-iterative-DPO-final The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/RLHFlow__LLaMA3-iterative-DPO-final-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.