rl-reward
multimodal-rewardbench-2Paper: https://arxiv.org/abs/2512.16899
Multimodal RewardBench 2 (MMRB2). Processed from https://github.com/facebookresearch/MMRB2 .
If you find this useful, please cite with following bibtex:
@article{hu2025multimodalrewardbench2,
title={Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image},
author={Hu, Yushi and Askari-Hemmat, Reyhane and Hall, Melissa and Dinan, Emily and Zettlemoyer, Luke and Ghazvininejad, Marjan},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/rl-research/multimodal-rewardbench-2.recipe_RL_data_bert-base-uncased_BASE_REWARD_32_STEPDIFF_rewardrecipe_RL_data_roberta-base_BASE_REWARD_32_STEPDIFF_rewardR1-Reward-RL
[📖 arXiv Paper]
[📊 R1-Reward Code]
[📝 R1-Reward Model]
Training Multimodal Reward Model Through Stable Reinforcement Learning
🔥 We are proud to open-source R1-Reward, a comprehensive project for improve reward modeling through reinforcement learning. This release includes:
R1-Reward Model: A state-of-the-art (SOTA) multimodal reward model demonstrating substantial gains (Voting@15):
13.5% improvement on VL Reward-Bench.3.5% improvement on MM-RLHF Reward-Bench.… See the full description on the dataset page: https://huggingface.co/datasets/yifanzhang114/R1-Reward-RL.qwen35-4b-dci-rl-rewardv2-step10-bcp100-eval
Qwen3.5-4B DCI RL reward-v2 step-10: BCP100 evaluation
This private evaluation artifact records the 100-query BrowseComp-Plus run of
eigentom/qwen35-4b-dci-rl-rewardv2-step10
in the DR-DCI BM25 local-retrieval environment.
Headline result
Model stage
Semantic accuracy
Vanilla Qwen3.5-4B (archived reference)
33/100
Best 4B SFT checkpoint, epoch 4 / step 850 (archived audit)
52/100
RL reward-v2, step 10 (DeepSeek-V4-Flash judge)
75/100
RL reward-v2… See the full description on the dataset page: https://huggingface.co/datasets/eigentom/qwen35-4b-dci-rl-rewardv2-step10-bcp100-eval.sotopia-rl-reward-annotation
Sotopia-RL: Reward Design for Social Intelligence Dataset
This repository contains the dataset and related resources for the paper Sotopia-RL: Reward Design for Social Intelligence.
Sotopia-RL proposes a novel framework that refines coarse episode-level feedback into utterance-level, multi-dimensional rewards. This enables more effective training of socially intelligent agents through reinforcement learning, particularly addressing challenges like partial observability and… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/sotopia-rl-reward-annotation.
