helpsteer
Weyaxi_-_HelpSteer-filtered-Solar-Instruct-ggufLLaMA-3.2-3B-DPO-HelpSteer3-Nemotron-EN-GGUFWeyaxi-HelpSteer-filtered-Solar-Instruct-GGUFLLaMA-3.2-3B-DPO-HelpSteer3-RMR1-14B-GGUFWeyaxi_-_HelpSteer-filtered-neural-chat-7b-v3-1-7B-ggufLLaMA-3.2-3B-DPO-HelpSteer3-Nemotron-Qwen3-GGUFWeyaxi_-_HelpSteer-filtered-7B-ggufWeyaxi-HelpSteer-filtered-neural-chat-7b-v3-1-7B-GGUF
HelpSteer2
HelpSteer2: Open-source dataset for training top-performing reward models
HelpSteer2 is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses.
This dataset has been created in partnership with Scale AI.
When used to tune a Llama 3.1 70B Instruct Model, we achieve 94.1% on RewardBench, which makes it the best Reward Model as… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer2.HelpSteer3
HelpSteer3
HelpSteer3 is an open-source dataset (CC-BY-4.0) that supports aligning models to become more helpful in responding to user prompts.
HelpSteer3-Preference can be used to train Llama 3.3 Nemotron Super 49B v1 (for Generative RMs) and Llama 3.3 70B Instruct Models (for Bradley-Terry RMs) to produce Reward Models that score as high as 85.5% on RM-Bench and 78.6% on JudgeBench, which substantially surpass existing Reward Models on these benchmarks.
HelpSteer3-Feedback and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer3.HelpSteer
HelpSteer: Helpfulness SteerLM Dataset
HelpSteer is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses.
Leveraging this dataset and SteerLM, we train a Llama 2 70B to reach 7.54 on MT Bench, the highest among models trained on open-source datasets based on MT Bench Leaderboard as of 15 Nov 2023.
This model is available on… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer.helpsteer2-pref-subsampleshelpsteer2_dpo_nonverboseHelperSteer 2, formatted in DPO format (prompt, chosen, rejected).
in main branch there is a custom scoring correct > helpful > -verbosity
in each branch we have preference pairs for only correct, helpful, verbosity, coherence, complexity
Please note that only correct and helpful has strong inter-rater agreement in the HelpSteer2 paper
This is the notebook used to produce the dataset… See the full description on the dataset page: https://huggingface.co/datasets/wassname/helpsteer2_dpo_nonverbose.Helpsteer-preference-standard
