datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VC-RewardBench
Visual-ERM
Visual-ERM is a multimodal generative reward model for vision-to-code tasks.It evaluates outputs directly in the rendered visual space and produces fine-grained, interpretable, and task-agnostic discrepancy feedback for structured visual reconstruction.
📄 Paper |
💻 GitHub |
📊 VC-RewardBench
Model Overview
Existing rewards for vision-to-code usually fall into two categories:
Text-based rewards such as edit distance or TEDS, which ignore… See the full description on the dataset page: https://huggingface.co/datasets/internlm/VC-RewardBench.VL-RewardBench
Dataset Card for VLRewardBench
Project Page:
https://vl-rewardbench.github.io
Dataset Summary
VLRewardBench is a comprehensive benchmark designed to evaluate vision-language generative reward models (VL-GenRMs) across visual perception, hallucination detection, and reasoning tasks. The benchmark contains 1,250 high-quality examples specifically curated to probe model limitations.
Dataset Structure
Each instance consists of multimodal queries spanning three key… See the full description on the dataset page: https://huggingface.co/datasets/MMInstruction/VL-RewardBench.multimodal-rewardbench-2Paper: https://arxiv.org/abs/2512.16899
Multimodal RewardBench 2 (MMRB2). Processed from https://github.com/facebookresearch/MMRB2 .
If you find this useful, please cite with following bibtex:
@article{hu2025multimodalrewardbench2,
title={Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image},
author={Hu, Yushi and Askari-Hemmat, Reyhane and Hall, Melissa and Dinan, Emily and Zettlemoyer, Luke and Ghazvininejad, Marjan},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/rl-research/multimodal-rewardbench-2.MM-RLHF-RewardBench
[📖 arXiv Paper]
[📊 MM-RLHF Data]
[📝 Homepage]
[🏆 Reward Model]
[🔮 MM-RewardBench]
[🔮 MM-SafetyBench]
[📈 Evaluation Suite]
The Next Step Forward in Multimodal LLM Alignment
[2025/02/10] 🔥 We are proud to open-source MM-RLHF, a comprehensive project for aligning Multimodal Large Language Models (MLLMs) with human preferences. This release includes:
A high-quality MLLM alignment dataset.
A strong Critique-Based MLLM reward model and its training algorithm.
A novel… See the full description on the dataset page: https://huggingface.co/datasets/yifanzhang114/MM-RLHF-RewardBench.multimodal_rewardbench
Dataset Card for Multimodal RewardBench
🏆 Dataset Attribution
This dataset is created by Yasunaga et al. (2025).
📄 Paper: Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
💻 GitHub Repository: https://github.com/facebookresearch/multimodal_rewardbench
I have downloaded the dataset from the GitHub repo and only modified the "Image" attribute by converting file paths to datasets.Image() for easier integration with 🤗… See the full description on the dataset page: https://huggingface.co/datasets/syhuggingface/multimodal_rewardbench.Agent-RewardBenchVL-RewardBench
Dataset Card for VLRewardBench
Project Page:
https://vl-rewardbench.github.io
Dataset Summary
VLRewardBench is a comprehensive benchmark designed to evaluate vision-language generative reward models (VL-GenRMs) across visual perception, hallucination detection, and reasoning tasks. The benchmark contains 1,250 high-quality examples specifically curated to probe model limitations.
Dataset Structure
Each instance consists of multimodal queries spanning three key… See the full description on the dataset page: https://huggingface.co/datasets/txsgbb/VL-RewardBench.XAIGID-RewardBench
XAIGID-RewardBench
The test split contains 3,988 benchmark triplets with an image, two policy-model responses, and up to three human-response slots. It contains 4,412 populated human annotations. The human_response split contains the 336-image Human-vs-Policy subset, including 36 GPT-5.5 comparisons, in the same image-and-response schema.
See the project repository and paper.
Vl-RewardBench
Dataset Card for VLRewardBench
Project Page:
https://vl-rewardbench.github.io
Dataset Summary
VLRewardBench is a comprehensive benchmark designed to evaluate vision-language generative reward models (VL-GenRMs) across visual perception, hallucination detection, and reasoning tasks. The benchmark contains 1,250 high-quality examples specifically curated to probe model limitations.
Dataset Structure
Each instance consists of multimodal queries spanning three key… See the full description on the dataset page: https://huggingface.co/datasets/Zhihui/Vl-RewardBench.Med-RewardBenchseed-mm
