datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mllm-as-embodied-world-judge
RoboJudge Data
This repository hosts the video assets and canonical metadata for evaluating
Physical Adherence (PA) and Instruction Alignment (IA) in generated
embodied-manipulation videos.
Current RoboJudge release
Use robojudge_release/ for the paper release:
Path
Contents
robojudge_release/train/physical_adherence.json
12,351 PA training records
robojudge_release/train/instruction_alignment.json
11,520 IA training records… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/mllm-as-embodied-world-judge.watercolour-rollouts-judge-led
Watercolour rollouts, judge-led run
Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by
writing p5.brush sketches. 861 paintings, the
sketch that produced each one, and the reward it earned, indexed by training step. This
is the run with the original reward mix from the write-up, where the pairwise judge and
its hand-rated pool carry most… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-judge-led.MLLM-as-a-Judgexprmt-qwen2.5-7b-instruct-multijail-judge-evalJudge-v2table-judge-benchmark
Benchmark design
The benchmark contains 538 paired clean/corrupted examples:
Error type
Cases
Corruption
Content: numeric
90
Change one numeric body-cell value
Content: typo
90
Transpose two adjacent, distinct Unicode letters in one body cell
Formatting
179
Bold and italicize letter-containing cells in one body row
Structure
179
Remove one row
Total
538
One fixed corruption per table
Corruption rules
Text typos never modify headers, tags… See the full description on the dataset page: https://huggingface.co/datasets/reducto/table-judge-benchmark.xprmt-llama-3.1-8b-instruct-multijail-judge-evalsoda-bottles-judged-agree2
Box-overlay preview
Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push.
roadsign-judged-ensemble-agree1
Box-overlay preview
Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push.
soda-bottles-judged-agree1
Box-overlay preview
Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push.
MM-JudgeBias
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
🏠 Project Page |
📄 arXiv |
🤗 Huggingface Dataset |
💻 Code
MM-JudgeBias measures Compositional Bias in MLLM-as-a-Judge — a systematic failure mode in which a judge does not correctly integrate and reason over all components (query, image, and response), and instead relies on partial cues. It contains 1,804 samples drawn from 29 source benchmarks (4 task types… See the full description on the dataset page: https://huggingface.co/datasets/naver-ai/MM-JudgeBias.JudgeMM-JudgeBench
MM-JudgeBench
Dataset Summary
MM-JudgeBench is a multilingual multimodal preference benchmark for evaluating
vision-language judge and reward models. Each row contains an image reference,
a query, two candidate responses, and a preference label.
The dataset includes three configurations:
m-vl-rewardbench
m-opencqa
m-mm-rewardbench
Each configuration provides two splits:
original
reversed
In the reversed split, the response order is swapped and the preference… See the full description on the dataset page: https://huggingface.co/datasets/tahmedge/MM-JudgeBench.judgebench-results
JudgeBench: LLM Cross-Judging Results
Experimental results from a controlled cross-judging study of 5 open-weights 7-9B LLMs acting as judges across 9 conditions × 2 temperatures.
Code repository: https://github.com/samarthraina/judgebench (coming soon)
Contents
summary_T*.csv, per_prompt_T*.csv — aggregated CSVs
full_results_T*.json — per-cell results with justifications (30,900 rows each)
cot_log_T*.jsonl — every individual K-draw with raw output
stat_tests.json —… See the full description on the dataset page: https://huggingface.co/datasets/samarthraina/judgebench-results.docvqa-media3-judged-trainval-agree2
Box-overlay preview
Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push.
spatial_edit_judge
Spatial Edit Judge Dataset
This dataset tests whether a visual judge can decide if a requested spatial camera edit was actually satisfied.
Each row contains a before image, an after image, a text instruction, and the ground-truth judge label. The judge should answer whether the after image correctly follows the instruction.
Dataset Structure
The Hugging Face viewer uses:
data/train-00000-of-00001.parquet
The repository also keeps the raw exported files:… See the full description on the dataset page: https://huggingface.co/datasets/saaduddinM/spatial_edit_judge.xprmt-qwen2.5-7b-instruct-advbench-judge-evalablation-llama-3.1-8b-instruct-multijail-judge-evaldocvqa-media-judged-ensemble
Box-overlay preview
Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push.
xprmt-llama-3.1-8b-instruct-advbench-judge-evalJudgement-Dayjudge_prompt_dataset_v1_dedup
Self-Contained Judge Dataset
Exported at: 20260602T022100Z
This is the Hugging Face Hub-ready form of the judge dataset.
It contains parquet shards with embedded prompt-order images and no separate asset tree.
Layout
train-*-of-*.parquet
validation-*-of-*.parquet
test-*-of-*.parquet
metadata.json
README.md
Load
from datasets import load_dataset
ds = load_dataset("Jsonwu/judge_prompt_dataset_v1_dedup")
Fields
prompt_images stores… See the full description on the dataset page: https://huggingface.co/datasets/ciderlab/judge_prompt_dataset_v1_dedup.docvqa-media3-judged-ensemble-v2-agree1
Box-overlay preview
Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push.
Judgement-Day
Judgement Day — Review Subset
A 5,200-submission subset of the Judgement Day dataset, released for anonymous review of
Judgement Day: An Anatomy of Successful Multimodal Attacks on Safety-Critical AI Systems.
Each record is an attack input (audio, image, video, PDF, email, or text) that a participant
submitted against a multimodal agent in one of eight safety-critical scenarios, and that
caused at least one evaluated model to select an unsafe action.
Scenarios… See the full description on the dataset page: https://huggingface.co/datasets/judgement-day-anon/Judgement-Day.docvqa-media3-judged-trainval-agree1
Box-overlay preview
Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push.
euclid_strong_lens_expert_judgesroadsign-judged-ensemble-agree2
Box-overlay preview
Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push.
docvqa-media3-judged-splits-agree1docvqa-media3-judged-splits-agree2synthetic_vqa_dataset_21.4k_images_vlm_as_judge_qwen_2.5_vl_3b_instruct
