datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
review_helpfulness_prediction
Dataset Card for Review Helpfulness Prediction (RHP) Dataset
Dataset Summary
The success of e-commerce services is largely dependent on helpful reviews that aid customers in making informed purchasing decisions. However, some reviews may be spammy or biased, making it challenging to identify which ones are helpful. Current methods for identifying helpful reviews only focus on the review text, ignoring the importance of who posted the review and when it was posted.… See the full description on the dataset page: https://huggingface.co/datasets/tafseer-nayeem/review_helpfulness_prediction.RiC_harmless_helpfulThe hhrlhf dataset for RiC (https://huggingface.co/papers/2402.10207) training with harmless (R1) and helpful (R2) rewards.
The 'input_ids' are obtained from Llama2 tokenizer. If you want to use other base models, replace it using other tokenizers.
Note: the rewards are already normalized accroding to their corresponding mean and std. The mean and std data for R1 and R2 are saved into all_reward_stat_harmhelp_Rlarge.npy.
The mean and std for R1 and R2 is (-0.94732502, 1.92034349)… See the full description on the dataset page: https://huggingface.co/datasets/Ray2333/RiC_harmless_helpful.train_data_Helpful_drdpo_preferencehelpsteer2-helpfulness-preference
Citation
@misc{wang2024helpsteer2preferencecomplementingratingspreferences,
title={HelpSteer2-Preference: Complementing Ratings with Preferences},
author={Zhilin Wang and Alexander Bukharin and Olivier Delalleau and Daniel Egert and Gerald Shen and Jiaqi Zeng and Oleksii Kuchaiev and Yi Dong},
year={2024},
eprint={2410.01257},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2410.01257},
}
@misc{wang2024helpsteer2… See the full description on the dataset page: https://huggingface.co/datasets/Jennny/helpsteer2-helpfulness-preference.train_data_Helpful_drdpo_preference_7b_sft_1eqwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.45-beta-0p3-margin-log
jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.45-beta-0p3-margin-log
Per-step margin summary statistics exported from a New-DPO training run.
Source Run
Model repo id: jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.45-beta-0p3
Base model: jackf857/qwen3-8b-base-sft-hh-helpful-4xh200-batch-64-20260417-214452
Training run name:… See the full description on the dataset page: https://huggingface.co/datasets/jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.45-beta-0p3-margin-log.hhrlhf_helpful_filteredultrabin_clean_max_chosen_min_rejected_rationalized_helpfulnessultrafeedback_binarized_helpfulness_prefsultra-50k-samples-dataset-helpfulnesshh_helpfulness_mc_rewards_IS_cliptrain_data_Helpful_drdpo_preference_7b_075qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-s_star-0.4-eta-0.01-margin-log
jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-s_star-0.4-eta-0.01-margin-log
Per-step margin summary statistics exported from a New-DPO training run.
Source Run
Model repo id: jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-s_star-0.4-eta-0.01
Base model: jackf857/qwen3-8b-base-sft-hh-helpful-4xh200-batch-64-20260417-214452
Training run name: qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-s_star-0.4-eta-0.01
W&B project:… See the full description on the dataset page: https://huggingface.co/datasets/jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-s_star-0.4-eta-0.01-margin-log.qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.35-margin-log
jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.35-margin-log
Per-step margin summary statistics exported from a New-DPO training run.
Source Run
Model repo id: jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.35
Base model: jackf857/qwen3-8b-base-sft-hh-helpful-4xh200-batch-64-20260417-214452
Training run name: qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.35
W&B project:… See the full description on the dataset page: https://huggingface.co/datasets/jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.35-margin-log.qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48-margin-log
jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48-margin-log
Per-step margin summary statistics exported from a New-DPO training run.
Source Run
Model repo id: jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48
Base model: jackf857/qwen3-8b-base-sft-hh-helpful-4xh200-batch-64-20260417-214452
Training run name: qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48
W&B project:… See the full description on the dataset page: https://huggingface.co/datasets/jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.48-margin-log.ontocord__Llama_3.2_1b-autoredteam_helpfulness-train-details
Dataset Card for Evaluation run of ontocord/Llama_3.2_1b-autoredteam_helpfulness-train
Dataset automatically created during the evaluation run of model ontocord/Llama_3.2_1b-autoredteam_helpfulness-train
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__Llama_3.2_1b-autoredteam_helpfulness-train-details.Llama-3.2-3B-Instruct-UF-REBEL_multi_armo_renamed_tokenized_helpfulness_logprobqwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.43-margin-log
W-61/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.43-margin-log
Per-step margin summary statistics exported from a New-DPO training run.
Source Run
Model repo id: W-61/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.43
Base model: jackf857/qwen3-8b-base-sft-hh-helpful-4xh200-batch-64-20260417-214452
Training run name: qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.43
W&B project:… See the full description on the dataset page: https://huggingface.co/datasets/W-61/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.43-margin-log.qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.5-margin-log
jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.5-margin-log
Per-step margin summary statistics exported from a New-DPO training run.
Source Run
Model repo id: jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.5
Base model: jackf857/qwen3-8b-base-sft-hh-helpful-4xh200-batch-64-20260417-214452
Training run name: qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.5
W&B project:… See the full description on the dataset page: https://huggingface.co/datasets/jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-s_star-0.4-eta-0.1-q_t-0.5-margin-log.helpful_harmless_data_10kdeceptive_helpfulness_boundary_v01Clarus Deceptive Helpfulness Boundary v0.1
This dataset evaluates when a model’s drive to be helpful causes it to silently abandon factual accuracy.
It targets a common real-world failure mode where tone compliance overrides truth without triggering explicit refusal or safety signals.
This is not training data.
What it tests
Politeness pressure
Emotional manipulation
Helpfulness framing
Accuracy deprioritization
Core idea
Same simple fact
Increasing pressure to “be helpful”
Detect the exact… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/deceptive_helpfulness_boundary_v01.llama-3-8b-base-margin-dpo-hh-helpful-margin-log
W-61/llama-3-8b-base-margin-dpo-hh-helpful-margin-log
Per-step margin summary statistics exported from a margin-DPO training run.
Source Run
Model repo id: W-61/llama-3-8b-base-margin-dpo-hh-helpful-8xh200
Base model: W-61/llama-3-8b-base-sft-hh-helpful-8xh200
Run name: llama-3-8b-base-margin-dpo-hh-helpful-8xh200-20260410-172009
Margin log path: /scratch/feng.yulu/dynamic-dpo-v4/outputs/llama-3-8b-base-margin-dpo-hh-helpful-8xh200-20260410-172009/margin_logs
Published… See the full description on the dataset page: https://huggingface.co/datasets/W-61/llama-3-8b-base-margin-dpo-hh-helpful-margin-log.qwen3-8b-base-epsilon-dpo-hh-helpful-4xh200-margin-loggenerative_rm_prompting_datahh_helpful_modelgpt-4o-mini_num200Llama-3.2-3B-Instruct-UF-REBEL_multi_armo_renamed_tokenized_helpfulnessqwen3-8b-base-margin-dpo-hh-helpful-4xh200-margin-logqwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.6-margin-log
jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.6-margin-log
Per-step margin summary statistics exported from a New-DPO training run.
Source Run
Model repo id: jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.6
Base model: jackf857/qwen3-8b-base-sft-hh-helpful-4xh200-batch-64-20260417-214452
Training run name: qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.6
W&B project:… See the full description on the dataset page: https://huggingface.co/datasets/jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.6-margin-log.qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-s_star-0.4-eta-0.05-margin-log
jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-s_star-0.4-eta-0.05-margin-log
Per-step margin summary statistics exported from a New-DPO training run.
Source Run
Model repo id: jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-s_star-0.4-eta-0.05
Base model: jackf857/qwen3-8b-base-sft-hh-helpful-4xh200-batch-64-20260417-214452
Training run name: qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-s_star-0.4-eta-0.05
W&B project:… See the full description on the dataset page: https://huggingface.co/datasets/jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-s_star-0.4-eta-0.05-margin-log.ontocord__RedPajama3b_v1-autoredteam_helpfulness-train-details
Dataset Card for Evaluation run of ontocord/RedPajama3b_v1-autoredteam_helpfulness-train
Dataset automatically created during the evaluation run of model ontocord/RedPajama3b_v1-autoredteam_helpfulness-train
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__RedPajama3b_v1-autoredteam_helpfulness-train-details.hh_helpfulness_qwen2.5_1.5b_generation
