rational
Heejindo_-_rationale_model_e10_save5000-ggufHeejindo_-_rationale_model_e10_save5000_eos-ggufHeejindo_-_rationale_model_e3_save5000_f2-ggufHeejindo_-_rationale_model_e3_save5000_f3-ggufHeejindo_-_rationale_model_e3_save5000_f4-ggufHeejindo_-_rationale_model_e10-ggufHeejindo_-_rationale_model_e3_save5000_rp_f1-ggufHeejindo_-_rationale_model_e3_save5000_rp-gguf
Datasets
All datasets matching “rational”movie_rationalesThe movie rationale dataset contains human annotated rationales for movie
reviews.RationaleRM
English | 中文
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
[📄 Paper] •
[🤗 Dataset] •
[📜 Citation]
Outcome Accuracy vs Rationale Consistency: Rationale Consistency effectively distinguishes frontier models and detects deceptive alignment
📖 Overview
RationaleRM is a research project that investigates how to align not just the outcomes but also the reasoning processes of reward models with human judgments.… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/RationaleRM.ultrabin_clean_max_chosen_min_rejected_rationalized_truthfulnessultrabin_clean_max_chosen_min_rejected_rationalized_honestyesnli_with_rationaleRationalRewards-SFTDataTLDR: this is the SFT trajectories for training reasoning reward model for text-to-image generation and image editing, from the following paper.
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
Haozhe Wang1
Cong Wei2
Weiming Ren2
Jiaming Liu3
Fangzhen Lin1
Wenhu Chen2
1 HKUST
2 University of Waterloo
3 Alibaba… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/RationalRewards-SFTData.
