judge
Datasets
All datasets matching “judge”mllm-as-embodied-world-judge
RoboJudge Data
This repository hosts the video assets and canonical metadata for evaluating
Physical Adherence (PA) and Instruction Alignment (IA) in generated
embodied-manipulation videos.
Current RoboJudge release
Use robojudge_release/ for the paper release:
Path
Contents
robojudge_release/train/physical_adherence.json
12,351 PA training records
robojudge_release/train/instruction_alignment.json
11,520 IA training records… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/mllm-as-embodied-world-judge.safedocs-1M-muse-spark-1.3-judged
SafeDocs: Muse Spark 1.3 judge annotations
Incrementally published, one complete shard per commit. All original source columns,
images, complete Paddle JSON, rows and row order are preserved. No language or quality
filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status,
and judge_error. Operational failures retain the original page with a null verdict
and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs.
Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.AnyAudio-Judge-Corpus
AnyAudio-Judge Corpus
An SFT training corpus that powers the AnyAudio-Judge evaluator. Each sample contains:
An audio clip (referenced relatively under audios/).
A multi-turn chat (messages) where the user enumerates a list of decomposed binary rubric items and the assistant answers them in JSON, with per-item evidence (Chain-of-Thought rationale).
A coarse label ("yes" if the caption originally matched the audio, "no" otherwise) and a tag describing how the caption was… See the full description on the dataset page: https://huggingface.co/datasets/cucl2/AnyAudio-Judge-Corpus.RefAny3D-Datasettwo-box-judge-gui
Two-Box Judge GUI Dataset
A multimodal dataset for training GUI element selection models. Given two candidate bounding boxes on a GUI screenshot, the model learns to select the one that better fulfills the user's intent.
Dataset Description
This dataset is designed for training judge models in GUI grounding pipelines. When a visual grounding model produces multiple candidate regions, the judge model determines which candidate best matches the user's command.… See the full description on the dataset page: https://huggingface.co/datasets/THU-BoZhang/two-box-judge-gui.JudgeAnythingThis dataset is described in the paper Judge Anything: MLLM as a Judge Across Any Modality.
