pairwise-preference
pairwise_preferencesv2orm-pairwise-preference-pairs
Pairwise Outcome Reward Model (ORM)
A Robust Preference Learning Model for Agentic Reasoning Systems
📋 Model Description
This is a Pairwise Outcome Reward Model (ORM) designed for agentic reasoning systems. The model learns to rank reasoning traces through relative preference judgments rather than absolute quality scores, achieving superior stability and reproducibility compared to traditional pointwise approaches.
Key Achievements:
✅ 96.3% pairwise accuracy with… See the full description on the dataset page: https://huggingface.co/datasets/LossFunctionLover/orm-pairwise-preference-pairs.pairwise_preferenceslfqa_expert_pairwise_human_preference_no_reasoningtasksource_oasst2_pairwise_rlhf_reward-PreferenceShareGPTpairwise_preferencesv3
