CoolFace
20 results

rational

eraser-benchmark /movie_rationalesThe movie rationale dataset contains human annotated rationales for movie reviews.text-classification1K<n<10K5 likes620 downloads3y agoHugging FaceQwen /RationaleRM English | 中文 Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models [📄 Paper] • [🤗 Dataset] • [📜 Citation] Outcome Accuracy vs Rationale Consistency: Rationale Consistency effectively distinguishes frontier models and detects deceptive alignment 📖 Overview RationaleRM is a research project that investigates how to align not just the outcomes but also the reasoning processes of reward models with human judgments.… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/RationaleRM.text-classification10K<n<100K29 likes618 downloads8mo agoHugging FaceContextualAI /ultrabin_clean_max_chosen_min_rejected_rationalized_truthfulnesstabular10K<n<100K0 likes538 downloads2y agoHugging FaceContextualAI /ultrabin_clean_max_chosen_min_rejected_rationalized_honestytabular10K<n<100K0 likes360 downloads2y agoHugging Facebatalovme /esnli_with_rationaletext100K<n<1M0 likes286 downloads2y agoHugging FaceTIGER-Lab /RationalRewards-SFTDataTLDR: this is the SFT trajectories for training reasoning reward model for text-to-image generation and image editing, from the following paper. RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time Haozhe Wang1   Cong Wei2   Weiming Ren2   Jiaming Liu3   Fangzhen Lin1   Wenhu Chen2 1 HKUST   2 University of Waterloo   3 Alibaba… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/RationalRewards-SFTData.texttext-to-image100K<n<1M1 likes128 downloads5mo agoHugging Face