datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
movie_rationalesThe movie rationale dataset contains human annotated rationales for movie
reviews.RationaleRM
English | 中文
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
[📄 Paper] •
[🤗 Dataset] •
[📜 Citation]
Outcome Accuracy vs Rationale Consistency: Rationale Consistency effectively distinguishes frontier models and detects deceptive alignment
📖 Overview
RationaleRM is a research project that investigates how to align not just the outcomes but also the reasoning processes of reward models with human judgments.… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/RationaleRM.ultrabin_clean_max_chosen_min_rejected_rationalized_truthfulnessultrabin_clean_max_chosen_min_rejected_rationalized_honestyesnli_with_rationaleRationalRewards-SFTDataTLDR: this is the SFT trajectories for training reasoning reward model for text-to-image generation and image editing, from the following paper.
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
Haozhe Wang1
Cong Wei2
Weiming Ren2
Jiaming Liu3
Fangzhen Lin1
Wenhu Chen2
1 HKUST
2 University of Waterloo
3 Alibaba… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/RationalRewards-SFTData.LitBench-RationalesIf you are the author of any comment in this dataset and would like it removed, please contact us and we will comply promptly.
RationalRewards-EvalData-GenAIBench-MMRB2-ERBenchTLDR: this is the RewardModel Evaluation dataset for text-to-image generation and image editing, from the following paper.
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
Haozhe Wang1
Cong Wei2
Weiming Ren2
Jiaming Liu3
Fangzhen Lin1
Wenhu Chen2
1 HKUST
2 University of Waterloo
3 Alibaba
RationalRewards is a… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/RationalRewards-EvalData-GenAIBench-MMRB2-ERBench.fair-rationalesExplainability methods are used to benchmark
the extent to which model predictions align
with human rationales i.e., are 'right for the
right reasons'. Previous work has failed to acknowledge, however,
that what counts as a rationale is sometimes subjective. This paper
presents what we think is a first of its kind, a
collection of human rationale annotations augmented with the annotators demographic information.magpie-reasoning-v1-20k-math-verifiable-step-by-step-rationale-alpaca-formatmagpie-reasoning-v1-20k-math-verifiable-step-by-step-rationaleMind2Web-HTML-cleaned-lite-with-desc_w_tao_value_rationalemagpie-reasoning-v1-10k-step-by-step-rationale-alpaca-formatmagpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-llama3.1RationaleVQASynthetic_Rationale
Synthetic Rationale Dataset: Enabling LLMs to Perform Explainable Assessment via Preference Optimization on MCTS
The Synthetic Rationale dataset is composed of intermediate assessment rationales generated by large language models (LLMs). Described as "noisy", these rationales may include errors or approximations, designed specifically for response-level explainable assessment of student answers in science and biology subjects. The rationales are derived from the thought tree data… See the full description on the dataset page: https://huggingface.co/datasets/jiazhengli/Synthetic_Rationale.ultrafeedback_rationale_Qwen2.5-3B-Instruct_cotbest_n_no_rationale_poc_agent_withjava_vulnllmSFT_PN_Rationales
SFT Dataset
generated from Qwen/Qwen3-VL-32B-Instruct
verified from OpenGVLab/InternVL3-78B
Domain Distribution of Positive/Negative Rationales
Per-Dataset Positive/Negative Rationale Counts by Domain
RationalRewards_DiffusionNFT_TrainDataTLDR: this is the diffusion RL training dataset for text-to-image generation and image editing, from the following paper.
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
Haozhe Wang1
Cong Wei2
Weiming Ren2
Jiaming Liu3
Fangzhen Lin1
Wenhu Chen2
1 HKUST
2 University of Waterloo
3 Alibaba
RationalRewards is a… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/RationalRewards_DiffusionNFT_TrainData.litbench-rationales-gpt4
LitBench Rationales - GPT-4 Rubric Evaluations
This dataset contains new rationales for story pair evaluations from the LitBench dataset, generated using GPT-4 with a structured rubric-based evaluation approach.
Evaluation Rubric
The rationales were generated using a 5-criterion rubric:
Creativity & Originality (25 points): Uniqueness of concept, innovative elements, fresh perspective
Writing Quality & Style (25 points): Prose quality, voice consistency, grammar and… See the full description on the dataset page: https://huggingface.co/datasets/SAA-Lab/litbench-rationales-gpt4.echr_rational
Dataset Card for echr_rational
Dataset Summary
Deconfounding Legal Judgment Prediction for European Court of Human
Rights Cases Towards Better Alignment with Experts
This work demonstrates that Legal Judgement Prediction systems without expert-informed adjustments can be vulnerable to shallow, distracting surface signals that arise from corpus construction, case distribution, and confounding factors. To mitigate this, we use domain expertise to strategically identify… See the full description on the dataset page: https://huggingface.co/datasets/TUMLegalTech/echr_rational.ultrafeedback_rationale_Llama-3.2-3B-Instruct_cotultrafeedback_rationale_Qwen2.5-3B-Instruct_ultra_sft_2e-5_thre-0.7_packing_42_cotbest_n_no_rationale_poc_onlyagent_vulnllmRationale_MCTS
Rationale MCTS Dataset: Enabling LLMs to Assess Through Rationale Thought Trees
The Rationale MCTS dataset consists of intermediate assessment rationales generated by large language models (LLMs). These rationales are "noisy," meaning they might contain errors or approximate reasoning, tailored for step-by-step explainable assessment of student answers in science and biology. The dataset targets questions from the The Hewlett Foundation: Short Answer Scoring competition, available… See the full description on the dataset page: https://huggingface.co/datasets/jiazhengli/Rationale_MCTS.LitBench-new-rationales
Dataset Card for "LitBench-Rationales-GPT4-Complete"
More Information needed
ultrabin_clean_max_chosen_min_rejected_rationalized_instruction_followingclimate_fever_rationalesThe Climate-Fever dataset was first collected and published by Diggelmann et al, 2020.For our study, we are interested in token-level rationales which are not available from the initial publication of Climate-Fever. Therefore, we manually selected a subset of 102 claims (510 claim-evidence pairs) based on clarity of the claim formulation and balanced claim labels. Each sample was annotated on token-level by 3 annotators as either supporting the claim (label=1), contradicting the claim… See the full description on the dataset page: https://huggingface.co/datasets/stephaniebrandl/climate_fever_rationales.movie_rationales_truncated
