CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFriends /mllm-as-embodied-world-judge RoboJudge Data This repository hosts the video assets and canonical metadata for evaluating Physical Adherence (PA) and Instruction Alignment (IA) in generated embodied-manipulation videos. Current RoboJudge release Use robojudge_release/ for the paper release: Path Contents robojudge_release/train/physical_adherence.json 12,351 PA training records robojudge_release/train/instruction_alignment.json 11,520 IA training records… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/mllm-as-embodied-world-judge.imagevideo-classification2 likes13k downloads10h agoHugging Face02FineEnvs /watercolour-rollouts-judge-led Watercolour rollouts, judge-led run Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one. Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by writing p5.brush sketches. 861 paintings, the sketch that produced each one, and the reward it earned, indexed by training step. This is the run with the original reward mix from the write-up, where the pairwise judge and its hand-rated pool carry most… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-judge-led.imagetext-to-imagen<1K0 likes1.5k downloads24d agoHugging Face03ONE-Lab /MLLM-as-a-Judgeimagequestion-answering1K<n<10K4 likes976 downloads2y agoHugging Face04Turbs /xprmt-qwen2.5-7b-instruct-multijail-judge-evalimage0 likes231 downloads5mo agoHugging Face05Icey444 /Judge-v2image1K<n<10K0 likes199 downloads7mo agoHugging Face06reducto /table-judge-benchmark Benchmark design The benchmark contains 538 paired clean/corrupted examples: Error type Cases Corruption Content: numeric 90 Change one numeric body-cell value Content: typo 90 Transpose two adjacent, distinct Unicode letters in one body cell Formatting 179 Bold and italicize letter-containing cells in one body row Structure 179 Remove one row Total 538 One fixed corruption per table Corruption rules Text typos never modify headers, tags… See the full description on the dataset page: https://huggingface.co/datasets/reducto/table-judge-benchmark.imageimage-to-textn<1K0 likes189 downloads3mo agoHugging Face07Turbs /xprmt-llama-3.1-8b-instruct-multijail-judge-evalimage0 likes166 downloads5mo agoHugging Face08merve /soda-bottles-judged-agree2 Box-overlay preview Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push. image1K<n<10K0 likes123 downloads3mo agoHugging Face09merve /roadsign-judged-ensemble-agree1 Box-overlay preview Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push. image1K<n<10K0 likes117 downloads3mo agoHugging Face10merve /soda-bottles-judged-agree1 Box-overlay preview Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push. image1K<n<10K0 likes116 downloads3mo agoHugging Face11naver-ai /MM-JudgeBias MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge 🏠 Project Page &nbsp;|&nbsp; 📄 arXiv &nbsp;|&nbsp; 🤗 Huggingface Dataset &nbsp;|&nbsp; 💻 Code MM-JudgeBias measures Compositional Bias in MLLM-as-a-Judge — a systematic failure mode in which a judge does not correctly integrate and reason over all components (query, image, and response), and instead relies on partial cues. It contains 1,804 samples drawn from 29 source benchmarks (4 task types… See the full description on the dataset page: https://huggingface.co/datasets/naver-ai/MM-JudgeBias.imagevisual-question-answering1K<n<10K1 likes99 downloads3mo agoHugging Face12Icey444 /Judgeimage1K<n<10K0 likes94 downloads7mo agoHugging Face13tahmedge /MM-JudgeBench MM-JudgeBench Dataset Summary MM-JudgeBench is a multilingual multimodal preference benchmark for evaluating vision-language judge and reward models. Each row contains an image reference, a query, two candidate responses, and a preference label. The dataset includes three configurations: m-vl-rewardbench m-opencqa m-mm-rewardbench Each configuration provides two splits: original reversed In the reversed split, the response order is swapped and the preference… See the full description on the dataset page: https://huggingface.co/datasets/tahmedge/MM-JudgeBench.imagevisual-question-answering100K<n<1M1 likes93 downloads3mo agoHugging Face14samarthraina /judgebench-results JudgeBench: LLM Cross-Judging Results Experimental results from a controlled cross-judging study of 5 open-weights 7-9B LLMs acting as judges across 9 conditions × 2 temperatures. Code repository: https://github.com/samarthraina/judgebench (coming soon) Contents summary_T*.csv, per_prompt_T*.csv — aggregated CSVs full_results_T*.json — per-cell results with justifications (30,900 rows each) cot_log_T*.jsonl — every individual K-draw with raw output stat_tests.json —… See the full description on the dataset page: https://huggingface.co/datasets/samarthraina/judgebench-results.documenttext-classificationn<1K0 likes90 downloads5mo agoHugging Face15merve /docvqa-media3-judged-trainval-agree2 Box-overlay preview Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push. image1K<n<10K0 likes81 downloads3mo agoHugging Face16saaduddinM /spatial_edit_judge Spatial Edit Judge Dataset This dataset tests whether a visual judge can decide if a requested spatial camera edit was actually satisfied. Each row contains a before image, an after image, a text instruction, and the ground-truth judge label. The judge should answer whether the after image correctly follows the instruction. Dataset Structure The Hugging Face viewer uses: data/train-00000-of-00001.parquet The repository also keeps the raw exported files:… See the full description on the dataset page: https://huggingface.co/datasets/saaduddinM/spatial_edit_judge.imageimage-to-imagen<1K0 likes76 downloads3mo agoHugging Face17Turbs /xprmt-qwen2.5-7b-instruct-advbench-judge-evalimage0 likes70 downloads5mo agoHugging Face18Turbs /ablation-llama-3.1-8b-instruct-multijail-judge-evalimage0 likes62 downloads5mo agoHugging Face19merve /docvqa-media-judged-ensemble Box-overlay preview Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push. image1K<n<10K0 likes56 downloads3mo agoHugging Face20Turbs /xprmt-llama-3.1-8b-instruct-advbench-judge-evalimage0 likes52 downloads5mo agoHugging Face21Dasool /Judgement-Daygatedaudio1K<n<10K0 likes52 downloads2mo agoHugging Face22ciderlab /judge_prompt_dataset_v1_dedup Self-Contained Judge Dataset Exported at: 20260602T022100Z This is the Hugging Face Hub-ready form of the judge dataset. It contains parquet shards with embedded prompt-order images and no separate asset tree. Layout train-*-of-*.parquet validation-*-of-*.parquet test-*-of-*.parquet metadata.json README.md Load from datasets import load_dataset ds = load_dataset("Jsonwu/judge_prompt_dataset_v1_dedup") Fields prompt_images stores… See the full description on the dataset page: https://huggingface.co/datasets/ciderlab/judge_prompt_dataset_v1_dedup.image1K<n<10K0 likes49 downloads3mo agoHugging Face23merve /docvqa-media3-judged-ensemble-v2-agree1 Box-overlay preview Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push. image1K<n<10K0 likes49 downloads3mo agoHugging Face24judgement-day-anon /Judgement-Day Judgement Day — Review Subset A 5,200-submission subset of the Judgement Day dataset, released for anonymous review of Judgement Day: An Anatomy of Successful Multimodal Attacks on Safety-Critical AI Systems. Each record is an attack input (audio, image, video, PDF, email, or text) that a participant submitted against a multimodal agent in one of eight safety-critical scenarios, and that caused at least one evaluated model to select an unsafe action. Scenarios… See the full description on the dataset page: https://huggingface.co/datasets/judgement-day-anon/Judgement-Day.audio1K<n<10K0 likes47 downloads2d agoHugging Face25merve /docvqa-media3-judged-trainval-agree1 Box-overlay preview Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push. image1K<n<10K0 likes44 downloads3mo agoHugging Face26mwalmsley /euclid_strong_lens_expert_judgesimage10K<n<100K0 likes40 downloads1y agoHugging Face27merve /roadsign-judged-ensemble-agree2 Box-overlay preview Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push. image1K<n<10K0 likes38 downloads3mo agoHugging Face28merve /docvqa-media3-judged-splits-agree1image1K<n<10K0 likes36 downloads3mo agoHugging Face29merve /docvqa-media3-judged-splits-agree2image1K<n<10K0 likes32 downloads3mo agoHugging Face30cosmo3769 /synthetic_vqa_dataset_21.4k_images_vlm_as_judge_qwen_2.5_vl_3b_instructimage10K<n<100K0 likes31 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.