CoolFace
18 results

rl-reward

rl-research /multimodal-rewardbench-2Paper: https://arxiv.org/abs/2512.16899 Multimodal RewardBench 2 (MMRB2). Processed from https://github.com/facebookresearch/MMRB2 . If you find this useful, please cite with following bibtex: @article{hu2025multimodalrewardbench2, title={Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image}, author={Hu, Yushi and Askari-Hemmat, Reyhane and Hall, Melissa and Dinan, Emily and Zettlemoyer, Luke and Ghazvininejad, Marjan}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/rl-research/multimodal-rewardbench-2.image1K<n<10K5 likes182 downloads9mo agoHugging FaceAnonymousSub /recipe_RL_data_bert-base-uncased_BASE_REWARD_32_STEPDIFF_reward1M<n<10M0 likes168 downloads4y agoHugging FaceAnonymousSub /recipe_RL_data_roberta-base_BASE_REWARD_32_STEPDIFF_reward1M<n<10M0 likes93 downloads4y agoHugging Faceyifanzhang114 /R1-Reward-RL [📖 arXiv Paper] [📊 R1-Reward Code] [📝 R1-Reward Model] Training Multimodal Reward Model Through Stable Reinforcement Learning 🔥 We are proud to open-source R1-Reward, a comprehensive project for improve reward modeling through reinforcement learning. This release includes: R1-Reward Model: A state-of-the-art (SOTA) multimodal reward model demonstrating substantial gains (Voting@15): 13.5% improvement on VL Reward-Bench.3.5% improvement on MM-RLHF Reward-Bench.… See the full description on the dataset page: https://huggingface.co/datasets/yifanzhang114/R1-Reward-RL.textimage-text-to-text10K<n<100K6 likes47 downloads1y agoHugging Faceeigentom /qwen35-4b-dci-rl-rewardv2-step10-bcp100-eval Qwen3.5-4B DCI RL reward-v2 step-10: BCP100 evaluation This private evaluation artifact records the 100-query BrowseComp-Plus run of eigentom/qwen35-4b-dci-rl-rewardv2-step10 in the DR-DCI BM25 local-retrieval environment. Headline result Model stage Semantic accuracy Vanilla Qwen3.5-4B (archived reference) 33/100 Best 4B SFT checkpoint, epoch 4 / step 850 (archived audit) 52/100 RL reward-v2, step 10 (DeepSeek-V4-Flash judge) 75/100 RL reward-v2… See the full description on the dataset page: https://huggingface.co/datasets/eigentom/qwen35-4b-dci-rl-rewardv2-step10-bcp100-eval.0 likes39 downloads2mo agoHugging Faceulab-ai /sotopia-rl-reward-annotation Sotopia-RL: Reward Design for Social Intelligence Dataset This repository contains the dataset and related resources for the paper Sotopia-RL: Reward Design for Social Intelligence. Sotopia-RL proposes a novel framework that refines coarse episode-level feedback into utterance-level, multi-dimensional rewards. This enables more effective training of socially intelligent agents through reinforcement learning, particularly addressing challenges like partial observability and… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/sotopia-rl-reward-annotation.texttext-generation1K<n<10K2 likes36 downloads1y agoHugging Face