datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HelpSteer2
HelpSteer2: Open-source dataset for training top-performing reward models
HelpSteer2 is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses.
This dataset has been created in partnership with Scale AI.
When used to tune a Llama 3.1 70B Instruct Model, we achieve 94.1% on RewardBench, which makes it the best Reward Model as… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer2.HelpSteer3
HelpSteer3
HelpSteer3 is an open-source dataset (CC-BY-4.0) that supports aligning models to become more helpful in responding to user prompts.
HelpSteer3-Preference can be used to train Llama 3.3 Nemotron Super 49B v1 (for Generative RMs) and Llama 3.3 70B Instruct Models (for Bradley-Terry RMs) to produce Reward Models that score as high as 85.5% on RM-Bench and 78.6% on JudgeBench, which substantially surpass existing Reward Models on these benchmarks.
HelpSteer3-Feedback and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer3.HelpSteer
HelpSteer: Helpfulness SteerLM Dataset
HelpSteer is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses.
Leveraging this dataset and SteerLM, we train a Llama 2 70B to reach 7.54 on MT Bench, the highest among models trained on open-source datasets based on MT Bench Leaderboard as of 15 Nov 2023.
This model is available on… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer.rlhf_helpful_evalouroboros-trace-help
Trace Help — does an execution trace help a model answer questions about a run?
In one minute. Twelve small programs in six languages (Python, JavaScript, C,
C++, Go, Elixir). Each was run once with a fixed command. Five questions per
program ask what actually happened on that one run: how many times a function
was called, what a particular call returned, what it was called with, whether a
function ran at all, which function raised. Sixty questions in total.
Every record carries… See the full description on the dataset page: https://huggingface.co/datasets/digitable-lol/ouroboros-trace-help.review_helpfulness_prediction
Dataset Card for Review Helpfulness Prediction (RHP) Dataset
Dataset Summary
The success of e-commerce services is largely dependent on helpful reviews that aid customers in making informed purchasing decisions. However, some reviews may be spammy or biased, making it challenging to identify which ones are helpful. Current methods for identifying helpful reviews only focus on the review text, ignoring the importance of who posted the review and when it was posted.… See the full description on the dataset page: https://huggingface.co/datasets/tafseer-nayeem/review_helpfulness_prediction.HelpingAI__Dhanishtha-Large-details
Dataset Card for Evaluation run of HelpingAI/Dhanishtha-Large
Dataset automatically created during the evaluation run of model HelpingAI/Dhanishtha-Large
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HelpingAI__Dhanishtha-Large-details.HelpSteer2-Preference-WarmStartOEvortex__HelpingAI2.5-10B-details
Dataset Card for Evaluation run of OEvortex/HelpingAI2.5-10B
Dataset automatically created during the evaluation run of model OEvortex/HelpingAI2.5-10B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OEvortex__HelpingAI2.5-10B-details.Weyaxi_HelpSteer-filtered-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Weyaxi_HelpSteer-filtered-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Weyaxi/HelpSteer-filtered with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/Weyaxi_HelpSteer-filtered-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.Dhanishtha-2.0-SUPERTHINKER📦 Dhanishtha-2.0-SUPERTHINKER
A distilled corpus of 11.7K high-quality samples showcasing multi-phase reasoning and structured emotional cognition. Sourced directly from the internal training data of Dhanishtha-2.0 — the world’s first Large Language Model (LLM) to implement Intermediate Thinking, featuring multiple <think> and <ser> blocks per response
📊 Overview
11.7K multilingual samples (languages listed below)
Instruction-Output format, ideal for supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/HelpingAI/Dhanishtha-2.0-SUPERTHINKER.ro_dpo_helpsteer2
Dataset Description
HelpSteer2 dataset contains 10k human-annotated preferences entries.
Here we provide the Romanian translation of the HelpSteer2 dataset, translated with GPT-4o mini.
This dataset is a next step of the alignment protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions (Masala et al., 2024).
Citation
@misc{wang2024helpsteer2preferencecomplementingratingspreferences… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_dpo_helpsteer2.LLM-Preferences-HelpSteer2
LLM-Preferences-HelpSteer2
Author: Min Li
Blog: https://rlhflow.github.io/posts/2025-01-22-decision-tree-reward-model/
Dataset Description
This dataset contains pairwise preference judgments from 34 modern LLMs on response pairs from the HelpSteer2 dataset.
Key Features
Contains 9,125 response pairs from HelpSteer2-Preference
Includes preferences from 9 closed-source and 25 open-source LLMs
Documents position bias analysis and preference consistency metrics… See the full description on the dataset page: https://huggingface.co/datasets/RLHFlow/LLM-Preferences-HelpSteer2.ro_dpo_helpsteer
Dataset Description
HelpSteer dataset contains 37k samples each containing a prompt, a response, as well as five human-annotated attributes of the response.
Here we provide the Romanian translation of the HelpSteer dataset, translated with Systran.
This dataset is part of the alignment protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions (Masala et al., 2024).
Citation… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_dpo_helpsteer.HelpSteer2-20k-jaNVIDIA が公開している SteerLM 向けのトライアルデータセット HelpSteer2を日本語に自動翻訳したデータセットになります。HelpSteer2 は Nemotron-4-430B-Reward でも利用されています。SteerLM でのアライメントや報酬モデルの作成にご活用下さい。
NVIDIA Releases Open Synthetic Data Generation Pipeline for Training Large Language Models
SteerLM での LLM トレーニング方法については以下の URL を参考にして下さい。
Announcing NVIDIA SteerLM : https://developer.nvidia.com/blog/announcing-steerlm-a-simple-and-practical-technique-to-customize-llms-during-inference
NeMo Aligner : https://github.com/NVIDIA/NeMo-Aligner… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/HelpSteer2-20k-ja.helpfulness-safety-calibration-dpo-100k
Helpfulness-Safety Calibration DPO (100K)
100,000 DPO preference pairs for calibrating the helpfulness-safety tradeoff in language models. Each example contains a prompt, a chosen response (correct handling), and a rejected response (incorrect handling) — covering both over-refusal and under-refusal failure modes.
Motivation
Safety-trained models often swing between two failure modes:
Over-refusal: Refusing legitimate requests because they superficially resemble… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/helpfulness-safety-calibration-dpo-100k.Ko.HelpSteer원본 데이터셋: nvidia/HelpSteer
help-steer-alpacaHelpSteer-35k-jaNVIDIA が公開している SteerLM 向けのトライアルデータセット HelpSteerを日本語に自動翻訳したデータセットになります。SteerLM でのアライメントをお試ししたい際にご活用下さい。
SteerLM での LLM トレーニング方法については以下の URL を参考にして下さい。
Announcing NVIDIA SteerLM : https://developer.nvidia.com/blog/announcing-steerlm-a-simple-and-practical-technique-to-customize-llms-during-inference
NeMo Aligner : https://github.com/NVIDIA/NeMo-Aligner
SteerLM training user guide : https://docs.nvidia.com/nemo-framework/user-guide/latest/modelalignment/steerlm.html
[参考] SteerLM :… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/HelpSteer-35k-ja.Intermediate-Thinking-130k
Intermediate-Thinking-130k
A comprehensive dataset of 135,000 high-quality samples designed to advance language model reasoning capabilities through structured intermediate thinking processes. This dataset enables training and evaluation of models with sophisticated self-correction and iterative reasoning abilities across 42 languages.
Overview
Intermediate-Thinking-130k addresses a fundamental limitation in current language models: their inability to pause, reflect, and… See the full description on the dataset page: https://huggingface.co/datasets/HelpingAI/Intermediate-Thinking-130k.HelpSteer-hindiHelpSteer3
HelpSteer3
HelpSteer3 is an open-source dataset (CC-BY-4.0) that supports aligning models to become more helpful in responding to user prompts.
HelpSteer3-Preference can be used to train Llama 3.3 Nemotron Super 49B v1 (for Generative RMs) and Llama 3.3 70B Instruct Models (for Bradley-Terry RMs) to produce Reward Models that score as high as 85.5% on RM-Bench and 78.6% on JudgeBench, which substantially surpass existing Reward Models on these benchmarks.
HelpSteer3-Feedback and… See the full description on the dataset page: https://huggingface.co/datasets/chilomax/HelpSteer3.hh-rlhf-helpful-base-jahttps://github.com/anthropics/hh-rlhf の内容のうち、helpful-base内のchosenに記載されている英文をfuguMTで翻訳、うまく翻訳できていないものを除外、修正したものです。
dpo-general-helpfulness-15k
General Helpfulness DPO Pairs (15K)
DPO preference pairs for training LLMs to give specific, actionable, genuinely useful responses instead of generic, hedged, or platitudinous ones.
Motivation
The most common failure mode in production LLMs isn't hallucination — it's unhelpfulness: vague answers, excessive caveats, refusals where none are needed, and generic advice that could apply to anyone. This dataset trains models to be genuinely helpful by rewarding… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/dpo-general-helpfulness-15k.HelpSteer_alpaca_reformattedM-Help
M-HELP: Mental Health Help-Seeking Behavior Dataset
📖 Description
M-HELP is the first dataset focused on identifying help-seeking behavior on social media.The dataset also includes multi-label mental health disorder annotations, enabling research not only in detecting whether a person is seeking help, but also in classifying possible underlying mental health concerns.
This dataset was introduced in our EMNLP Findings 2025 paper:"M-HELP: Using Social Media Data to… See the full description on the dataset page: https://huggingface.co/datasets/zuhashaik/M-Help.help-request-messages-v2
📊 Help Classifier Dataset (v2)
🧠 Overview
The Help Classifier Dataset (v2) is a curated NLP dataset designed to classify student help requests into meaningful categories within a collaborative learning environment.
This dataset was developed as part of a larger AI system for the Coding in Color (CIC) ecosystem, where students work across domains such as AI development, game development, 2D/3D art, and robotics.
The goal of this dataset is to enable models to:… See the full description on the dataset page: https://huggingface.co/datasets/King-8/help-request-messages-v2.Tauri-Helpsteer-3-Preference-KTOHelpSteer3_dpo_format
Dataset: nvidia/HelpSteer3
HelpSteer3 is an open-source dataset (CC-BY-4.0) that supports aligning models to become more helpful in responding to user prompts.
Preference Score Integer from -3 to 3, corresponding to:
-3: Response 1 is much better than Response 2
-2: Response 1 is better than Response 2
-1: Response 1 is slightly better than Response 2
0: Response 1 is about the same as Response 2
1: Response 2 is slightly better than Response 1
2: Response 2 is better than Response… See the full description on the dataset page: https://huggingface.co/datasets/CarrotAI/HelpSteer3_dpo_format.HelpSteer3
HelpSteer3
HelpSteer3 is an open-source dataset (CC-BY-4.0) that supports aligning models to become more helpful in responding to user prompts.
HelpSteer3-Preference can be used to train Llama 3.3 Nemotron Super 49B v1 (for Generative RMs) and Llama 3.3 70B Instruct Models (for Bradley-Terry RMs) to produce Reward Models that score as high as 85.5% on RM-Bench and 78.6% on JudgeBench, which substantially surpass existing Reward Models on these benchmarks.
HelpSteer3-Feedback and… See the full description on the dataset page: https://huggingface.co/datasets/sunorme/HelpSteer3.
