datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mllm-as-embodied-world-judge
RoboJudge Data
This repository hosts the video assets and canonical metadata for evaluating
Physical Adherence (PA) and Instruction Alignment (IA) in generated
embodied-manipulation videos.
Current RoboJudge release
Use robojudge_release/ for the paper release:
Path
Contents
robojudge_release/train/physical_adherence.json
12,351 PA training records
robojudge_release/train/instruction_alignment.json
11,520 IA training records… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/mllm-as-embodied-world-judge.safedocs-1M-muse-spark-1.3-judged
SafeDocs: Muse Spark 1.3 judge annotations
Incrementally published, one complete shard per commit. All original source columns,
images, complete Paddle JSON, rows and row order are preserved. No language or quality
filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status,
and judge_error. Operational failures retain the original page with a null verdict
and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs.
Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.AnyAudio-Judge-Corpus
AnyAudio-Judge Corpus
An SFT training corpus that powers the AnyAudio-Judge evaluator. Each sample contains:
An audio clip (referenced relatively under audios/).
A multi-turn chat (messages) where the user enumerates a list of decomposed binary rubric items and the assistant answers them in JSON, with per-item evidence (Chain-of-Thought rationale).
A coarse label ("yes" if the caption originally matched the audio, "no" otherwise) and a tag describing how the caption was… See the full description on the dataset page: https://huggingface.co/datasets/cucl2/AnyAudio-Judge-Corpus.RefAny3D-Datasettwo-box-judge-gui
Two-Box Judge GUI Dataset
A multimodal dataset for training GUI element selection models. Given two candidate bounding boxes on a GUI screenshot, the model learns to select the one that better fulfills the user's intent.
Dataset Description
This dataset is designed for training judge models in GUI grounding pipelines. When a visual grounding model produces multiple candidate regions, the judge model determines which candidate best matches the user's command.… See the full description on the dataset page: https://huggingface.co/datasets/THU-BoZhang/two-box-judge-gui.JudgeAnythingThis dataset is described in the paper Judge Anything: MLLM as a Judge Across Any Modality.
AnyAudio-Judge-Bench
AnyAudio-Judge Bench
Bilingual (English / Chinese) multi-domain benchmark for instruction-audio alignment evaluation, released alongside the paper "AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following".
7,920 curated samples per language across 7 subsets
Strict 1 : 1 positive : negative ratio per subset
Hard negatives via instruction swapping and attribute perturbation
Each row carries a list of decomposed binary rubric items (yes/no… See the full description on the dataset page: https://huggingface.co/datasets/cucl2/AnyAudio-Judge-Bench.safedocs-cc-2m-paddle-vl-1-6-openrouter-judged
SafeDocs selected corpus: OCR judge annotations
All source rows and columns are preserved, including images and complete Paddle
outputs. Added columns: judge_verdict, judge_reason, judge_status, judge_error.
PERFECT/ERROR are model quality judgments, not verified ground truth.
Operational failures have null verdicts and are distinct from OCR errors.
No pages are filtered. Whole-document filtering and enrichment are downstream.
Muse Spark 1.3 Contributor through OpenRouter, low… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-cc-2m-paddle-vl-1-6-openrouter-judged.ru-llm-judge-dataset
RU-LLM-Judge-Dataset
Датасет собирается инкрементально в течение нескольких сессий бесплатного Google Colab.
Текущий объём: 20,553 суждений (по состоянию на последний запуск).
Прогресс к цели (5,000 суждений)
[████████████████████] 100% (20,553 / 5,000)
История сессий сбора
Сессия
Дата
Добавлено
Итого
1
2026-08-05 08:42
617
617
2
2026-08-06 14:40
583
1,200
3
2026-08-07 19:20
486
1,686
4
2026-08-08 22:34
868
2,554
5
2026-08-13… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-llm-judge-dataset.Indian-Supreme-Court-Judgements-Chunked
Indian Supreme Court Judgements Chunked
Executive Summary
The dataset aims to address the chronic backlog in the Indian judiciary system, particularly in the Supreme Court, by creating a dataset optimized for legal language models (LLMs). The dataset will consist of pre-processed, chunked, and embedded textual data derived from the Supreme Court's judgment PDFs.
Problem and Importance - Motivation
Indian courts are overwhelmed with pending cases, with the… See the full description on the dataset page: https://huggingface.co/datasets/vihaannnn/Indian-Supreme-Court-Judgements-Chunked.MultiTurn-Chat-MT-Bench-Judge
SEA-MT-Bench-Judge
SEA-MT-Bench-Judge expands on the original SEA-MTBench through the use of a criteria-based evaluation framework. We use GPT-OSS-120B as the judge model.
The prompts are based on MT-Bench and was manually translated by native speakers. Furthermore, some prompts were modified to be more suitable for the criteria-based judgments.
Supported Tasks and Leaderboards
SEA-MT-Bench-Judge is designed for evaluating chat or instruction-tuned large language… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/MultiTurn-Chat-MT-Bench-Judge.watercolour-rollouts-judge-led
Watercolour rollouts, judge-led run
Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by
writing p5.brush sketches. 861 paintings, the
sketch that produced each one, and the reward it earned, indexed by training step. This
is the run with the original reward mix from the write-up, where the pairwise judge and
its hand-rated pool carry most… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-judge-led.judged_science_completionsJudgeBench
JudgeBench: A Benchmark for Evaluating LLM-Based Judges
📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard]
JudgeBench is a benchmark aimed at evaluating LLM-based judges for objective correctness on challenging response pairs. For more information on how the response pairs are constructed, please see our paper.
Data Instance and Fields
This release includes two dataset splits. The gpt split includes 350 unique response pairs generated by GPT-4o and the claude… See the full description on the dataset page: https://huggingface.co/datasets/ScalerLab/JudgeBench.MLLM-as-a-Judgeus-court-opinions-dockets-judges-dataset
US Court Opinions Metadata, Dockets & Judges (CourtListener)
10M opinion clusters, 70M dockets and 16K judges from official CourtListener / Free Law Project bulk data as metadata + derived-signals tables — citation graph, company litigation profiles; no opinion full text.
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:
https://data.zalize.com/datasets/us-court-opinions-dockets-judges-dataset… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-court-opinions-dockets-judges-dataset.Linguistic-Diagnostics-Syntax-Judge
LINDSEA Syntax
LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian.
Supported Tasks and Leaderboards
LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs).
Languages
Indonesian (id)
Dataset Details
Data Sources
Data Source
License
Language/s
Split/s
CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax-Judge.Cultural-Evaluation-Kalahi-Judge
Kalahi-Judge
Kalahi evaluates the ability of LLMs to generate responses relevant to Filipino culture in terms of shared knowledge and ethics. This dataset extends the prompts found in Kalahi dataset to use a criteria-based judging metric.
Supported Tasks and Leaderboards
Kalahi is designed for evaluating Filipino cultural representations in instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Cultural-Evaluation-Kalahi-Judge.decomposeRL-tiny-judge
DecomposeRL Tiny-Judge: Distillation Data
Overview
DecomposeRL Tiny-Judge is the distillation dataset used to train DecomposeRL's tiny-judge stack — eight small ModernBERT-large classifier heads that replace a Qwen3-32B LLM judge as the reward model during GRPO training.
Each row is a judgment task instance: a text input (claim / question / answer / evidence, depending on the task) paired with a label distilled from a Qwen/Qwen3-32B judge call… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/decomposeRL-tiny-judge.numina-math-olympiads-judgedjudgeResearch-Intent-Judge
Research Intent — LLM-as-Judge
▶️ Watch the Video
LLM-as-Judge annotations for research paper intent classification, collected
through the Echo-DSRN collaborative platform during the OpenAIRE AI Hackathon 2026.
The dataset has one split per judge model (Gemma_4_E4B_it_GGUF,
Qwen3.6_35B_A3B_GGUF, Bonsai_8B_gguf, ...) plus a human_annotations
split with curator annotations. Split names use underscores in place of the
dashes in model names (HF does not allow dashes in split… See the full description on the dataset page: https://huggingface.co/datasets/ethicalabs/Research-Intent-Judge.VA-Judger-Bench
VA-Judger-Bench
VA-Judger-Bench is a paired audio/video preference benchmark with 1,150 cases:
easy: 400 cases
indomain: 250 cases
outdomain: 500 cases
Layout
Each split contains data.jsonl and a videos/ directory. All video paths are
relative to the split directory.
VA-Judger-Bench/
├── README.md
├── easy/
│ ├── data.jsonl
│ └── videos/
├── indomain/
│ ├── data.jsonl
│ └── videos/
└── outdomain/
├── data.jsonl
└── videos/
Record… See the full description on the dataset page: https://huggingface.co/datasets/ShareLab-SII/VA-Judger-Bench.judge-distillation-medical-interpretability
Judge-Distillation Medical Misalignment Interpretability Dataset
A complete artifact bundle for the Phase 2 judge-distillation experiments
described in
judge_distillation/RESULTS.md.
Includes training datasets, source per-prompt activations (the underlying
drift_pct labels), Gemma Scope SAE feature attributions, hidden-state
captures, and transfer-test corpora & scores across all five versions
(v1–v5).
What's in here
Training datasets
Three versions of… See the full description on the dataset page: https://huggingface.co/datasets/burnssa/judge-distillation-medical-interpretability.r_judge_labelled
R-Judge with LLM-Judge Labels
This dataset augments the R-Judge benchmark with automated safety labels produced by an LLM judge. R-Judge is a benchmark for evaluating the safety judgment capability of LLMs in multi-turn agent scenarios, spanning five application domains.
Files
File
Description
r_judge_data.csv
Base dataset extracted from R-Judge (568 rows, deduplicated)
r_judge_labelled_anthropic_claude-sonnet-4-6.csv
Base dataset augmented with… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/r_judge_labelled.data-zh-legal-judgements
裁判文书数据集 (1985-2021)
说明
整理自马克数据网提供的原始数据(1985年至2021年10月的裁判文书全量数据,原数据共94.36G)。
裁判文书官网:https://wenshu.court.gov.cn/
马克数据网:https://www.macrodatas.cn/
下载来源:
http://github.com/cncases/cases
磁力链接
magnet:?xt=urn:btih:c6aac12ebd697041ba60a8cba9f7326155921fae
magnet:?xt=urn:btih:afa29281baf8ab6a3f5b1e9b9b0799e120611db1
resilio sync
BC76W4N26A3ZCOQEIAQYMMBGY7PRWE6TG
Archive.org
https://archive.org/details/caipan
和鲸社区… See the full description on the dataset page: https://huggingface.co/datasets/qundao/data-zh-legal-judgements.petri-judge-summaries-top50-llama70bgpt4_judge_battles_judged_science_traces_original_DeepSeek-R1-Distill-Qwen-32BJudgeLM-100K
Dataset Card for JudgeLM
Dataset Summary
JudgeLM-100K dataset contains 100,000 judge samples for training and 5,000 judge samples for validation. All the judge samples have the GPT-4-generated high-quality judgements.
This instruction data can be used to conduct instruction-tuning for language models and make the language model has ability to judge open-ended answer pairs.
See more details in the "Dataset" section and the appendix sections of this paper.
This produced a… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/JudgeLM-100K.
