CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFriends /mllm-as-embodied-world-judge RoboJudge Data This repository hosts the video assets and canonical metadata for evaluating Physical Adherence (PA) and Instruction Alignment (IA) in generated embodied-manipulation videos. Current RoboJudge release Use robojudge_release/ for the paper release: Path Contents robojudge_release/train/physical_adherence.json 12,351 PA training records robojudge_release/train/instruction_alignment.json 11,520 IA training records… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/mllm-as-embodied-world-judge.imagevideo-classification2 likes13k downloads7h agoHugging Face02albertklorer /safedocs-1M-muse-spark-1.3-judged SafeDocs: Muse Spark 1.3 judge annotations Incrementally published, one complete shard per commit. All original source columns, images, complete Paddle JSON, rows and row order are preserved. No language or quality filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status, and judge_error. Operational failures retain the original page with a null verdict and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs. Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.tabular100K<n<1M0 likes12k downloads7d agoHugging Face03cucl2 /AnyAudio-Judge-Corpus AnyAudio-Judge Corpus An SFT training corpus that powers the AnyAudio-Judge evaluator. Each sample contains: An audio clip (referenced relatively under audios/). A multi-turn chat (messages) where the user enumerates a list of decomposed binary rubric items and the assistant answers them in JSON, with per-item evidence (Chain-of-Thought rationale). A coarse label ("yes" if the caption originally matched the audio, "no" otherwise) and a tag describing how the caption was… See the full description on the dataset page: https://huggingface.co/datasets/cucl2/AnyAudio-Judge-Corpus.audioaudio-text-to-text10K<n<100K0 likes11k downloads2mo agoHugging Face04JudgementH /RefAny3D-Dataset0 likes7.8k downloads8mo agoHugging Face05THU-BoZhang /two-box-judge-gui Two-Box Judge GUI Dataset A multimodal dataset for training GUI element selection models. Given two candidate bounding boxes on a GUI screenshot, the model learns to select the one that better fulfills the user's intent. Dataset Description This dataset is designed for training judge models in GUI grounding pipelines. When a visual grounding model produces multiple candidate regions, the judge model determines which candidate best matches the user's command.… See the full description on the dataset page: https://huggingface.co/datasets/THU-BoZhang/two-box-judge-gui.visual-question-answering100K<n<1M2 likes5.9k downloads4mo agoHugging Face06pudashi /JudgeAnythingThis dataset is described in the paper Judge Anything: MLLM as a Judge Across Any Modality. any-to-any1K<n<10K2 likes4.6k downloads1y agoHugging Face07cucl2 /AnyAudio-Judge-Bench AnyAudio-Judge Bench Bilingual (English / Chinese) multi-domain benchmark for instruction-audio alignment evaluation, released alongside the paper "AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following". 7,920 curated samples per language across 7 subsets Strict 1 : 1 positive : negative ratio per subset Hard negatives via instruction swapping and attribute perturbation Each row carries a list of decomposed binary rubric items (yes/no… See the full description on the dataset page: https://huggingface.co/datasets/cucl2/AnyAudio-Judge-Bench.audioaudio-classification10K<n<100K2 likes4.5k downloads3mo agoHugging Face08albertklorer /safedocs-cc-2m-paddle-vl-1-6-openrouter-judged SafeDocs selected corpus: OCR judge annotations All source rows and columns are preserved, including images and complete Paddle outputs. Added columns: judge_verdict, judge_reason, judge_status, judge_error. PERFECT/ERROR are model quality judgments, not verified ground truth. Operational failures have null verdicts and are distinct from OCR errors. No pages are filtered. Whole-document filtering and enrichment are downstream. Muse Spark 1.3 Contributor through OpenRouter, low… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-cc-2m-paddle-vl-1-6-openrouter-judged.0 likes3.2k downloads1m agoHugging Face09AtesiT /ru-llm-judge-dataset RU-LLM-Judge-Dataset Датасет собирается инкрементально в течение нескольких сессий бесплатного Google Colab. Текущий объём: 20,553 суждений (по состоянию на последний запуск). Прогресс к цели (5,000 суждений) [████████████████████] 100% (20,553 / 5,000) История сессий сбора Сессия Дата Добавлено Итого 1 2026-08-05 08:42 617 617 2 2026-08-06 14:40 583 1,200 3 2026-08-07 19:20 486 1,686 4 2026-08-08 22:34 868 2,554 5 2026-08-13… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-llm-judge-dataset.text10K<n<100K0 likes2.9k downloads2h agoHugging Face10vihaannnn /Indian-Supreme-Court-Judgements-Chunked Indian Supreme Court Judgements Chunked Executive Summary The dataset aims to address the chronic backlog in the Indian judiciary system, particularly in the Supreme Court, by creating a dataset optimized for legal language models (LLMs). The dataset will consist of pre-processed, chunked, and embedded textual data derived from the Supreme Court's judgment PDFs. Problem and Importance - Motivation Indian courts are overwhelmed with pending cases, with the… See the full description on the dataset page: https://huggingface.co/datasets/vihaannnn/Indian-Supreme-Court-Judgements-Chunked.textfeature-extraction10K<n<100K6 likes2.4k downloads2y agoHugging Face11aisingapore /MultiTurn-Chat-MT-Bench-Judgegated SEA-MT-Bench-Judge SEA-MT-Bench-Judge expands on the original SEA-MTBench through the use of a criteria-based evaluation framework. We use GPT-OSS-120B as the judge model. The prompts are based on MT-Bench and was manually translated by native speakers. Furthermore, some prompts were modified to be more suitable for the criteria-based judgments. Supported Tasks and Leaderboards SEA-MT-Bench-Judge is designed for evaluating chat or instruction-tuned large language… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/MultiTurn-Chat-MT-Bench-Judge.tabularn<1K0 likes2k downloads2mo agoHugging Face12FineEnvs /watercolour-rollouts-judge-led Watercolour rollouts, judge-led run Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one. Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by writing p5.brush sketches. 861 paintings, the sketch that produced each one, and the reward it earned, indexed by training step. This is the run with the original reward mix from the write-up, where the pairwise judge and its hand-rated pool carry most… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-judge-led.imagetext-to-imagen<1K0 likes1.5k downloads23d agoHugging Face13reasoning-proj /judged_science_completionstabularn<1K2 likes1.4k downloads1y agoHugging Face14ScalerLab /JudgeBench JudgeBench: A Benchmark for Evaluating LLM-Based Judges 📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard] JudgeBench is a benchmark aimed at evaluating LLM-based judges for objective correctness on challenging response pairs. For more information on how the response pairs are constructed, please see our paper. Data Instance and Fields This release includes two dataset splits. The gpt split includes 350 unique response pairs generated by GPT-4o and the claude… See the full description on the dataset page: https://huggingface.co/datasets/ScalerLab/JudgeBench.texttext-classificationn<1K12 likes1.1k downloads2y agoHugging Face15ONE-Lab /MLLM-as-a-Judgeimagequestion-answering1K<n<10K4 likes976 downloads2y agoHugging Face16zalizedata /us-court-opinions-dockets-judges-dataset US Court Opinions Metadata, Dockets & Judges (CourtListener) 10M opinion clusters, 70M dockets and 16K judges from official CourtListener / Free Law Project bulk data as metadata + derived-signals tables — citation graph, company litigation profiles; no opinion full text. Part of the DataForge Open Data program — full production packages, free for academic and personal use. Canonical dataset page: https://data.zalize.com/datasets/us-court-opinions-dockets-judges-dataset… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-court-opinions-dockets-judges-dataset.tabulartext-classification10M<n<100M0 likes882 downloads1mo agoHugging Face17aisingapore /Linguistic-Diagnostics-Syntax-Judgegated LINDSEA Syntax LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian. Supported Tasks and Leaderboards LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs). Languages Indonesian (id) Dataset Details Data Sources Data Source License Language/s Split/s CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax-Judge.textn<1K0 likes659 downloads2mo agoHugging Face18aisingapore /Cultural-Evaluation-Kalahi-Judgegated Kalahi-Judge Kalahi evaluates the ability of LLMs to generate responses relevant to Filipino culture in terms of shared knowledge and ethics. This dataset extends the prompts found in Kalahi dataset to use a criteria-based judging metric. Supported Tasks and Leaderboards Kalahi is designed for evaluating Filipino cultural representations in instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Cultural-Evaluation-Kalahi-Judge.textn<1K0 likes650 downloads2mo agoHugging Face19dipta007 /decomposeRL-tiny-judge DecomposeRL Tiny-Judge: Distillation Data Overview DecomposeRL Tiny-Judge is the distillation dataset used to train DecomposeRL's tiny-judge stack — eight small ModernBERT-large classifier heads that replace a Qwen3-32B LLM judge as the reward model during GRPO training. Each row is a judgment task instance: a text input (claim / question / answer / evidence, depending on the task) paired with a label distilled from a Qwen/Qwen3-32B judge call… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/decomposeRL-tiny-judge.texttext-classification10M<n<100M0 likes532 downloads4mo agoHugging Face20mlfoundations-dev /numina-math-olympiads-judgedtext10K<n<100K0 likes486 downloads2y agoHugging Face21Songama /judge0 likes483 downloads25d agoHugging Face22ethicalabs /Research-Intent-Judge Research Intent — LLM-as-Judge ▶️ Watch the Video LLM-as-Judge annotations for research paper intent classification, collected through the Echo-DSRN collaborative platform during the OpenAIRE AI Hackathon 2026. The dataset has one split per judge model (Gemma_4_E4B_it_GGUF, Qwen3.6_35B_A3B_GGUF, Bonsai_8B_gguf, ...) plus a human_annotations split with curator annotations. Split names use underscores in place of the dashes in model names (HF does not allow dashes in split… See the full description on the dataset page: https://huggingface.co/datasets/ethicalabs/Research-Intent-Judge.texttext-classification100K<n<1M0 likes373 downloads1mo agoHugging Face23ShareLab-SII /VA-Judger-Bench VA-Judger-Bench VA-Judger-Bench is a paired audio/video preference benchmark with 1,150 cases: easy: 400 cases indomain: 250 cases outdomain: 500 cases Layout Each split contains data.jsonl and a videos/ directory. All video paths are relative to the split directory. VA-Judger-Bench/ ├── README.md ├── easy/ │ ├── data.jsonl │ └── videos/ ├── indomain/ │ ├── data.jsonl │ └── videos/ └── outdomain/ ├── data.jsonl └── videos/ Record… See the full description on the dataset page: https://huggingface.co/datasets/ShareLab-SII/VA-Judger-Bench.video1K<n<10K0 likes346 downloads1mo agoHugging Face24burnssa /judge-distillation-medical-interpretability Judge-Distillation Medical Misalignment Interpretability Dataset A complete artifact bundle for the Phase 2 judge-distillation experiments described in judge_distillation/RESULTS.md. Includes training datasets, source per-prompt activations (the underlying drift_pct labels), Gemma Scope SAE feature attributions, hidden-state captures, and transfer-test corpora & scores across all five versions (v1–v5). What's in here Training datasets Three versions of… See the full description on the dataset page: https://huggingface.co/datasets/burnssa/judge-distillation-medical-interpretability.1K<n<10K0 likes300 downloads5mo agoHugging Face25Glide-py /r_judge_labelled R-Judge with LLM-Judge Labels This dataset augments the R-Judge benchmark with automated safety labels produced by an LLM judge. R-Judge is a benchmark for evaluating the safety judgment capability of LLMs in multi-turn agent scenarios, spanning five application domains. Files File Description r_judge_data.csv Base dataset extracted from R-Judge (568 rows, deduplicated) r_judge_labelled_anthropic_claude-sonnet-4-6.csv Base dataset augmented with… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/r_judge_labelled.tabulartext-classificationn<1K0 likes279 downloads4mo agoHugging Face26qundao /data-zh-legal-judgementsgated 裁判文书数据集 (1985-2021) 说明 整理自马克数据网提供的原始数据(1985年至2021年10月的裁判文书全量数据,原数据共94.36G)。 裁判文书官网:https://wenshu.court.gov.cn/ 马克数据网:https://www.macrodatas.cn/ 下载来源: http://github.com/cncases/cases 磁力链接 magnet:?xt=urn:btih:c6aac12ebd697041ba60a8cba9f7326155921fae magnet:?xt=urn:btih:afa29281baf8ab6a3f5b1e9b9b0799e120611db1 resilio sync BC76W4N26A3ZCOQEIAQYMMBGY7PRWE6TG Archive.org https://archive.org/details/caipan 和鲸社区… See the full description on the dataset page: https://huggingface.co/datasets/qundao/data-zh-legal-judgements.texttext-classification10M<n<100M4 likes278 downloads6mo agoHugging Face27auditing-agents /petri-judge-summaries-top50-llama70b0 likes271 downloads3mo agoHugging Face28routellm /gpt4_judge_battlestabular100K<n<1M2 likes258 downloads1y agoHugging Face29reasoning-proj /_judged_science_traces_original_DeepSeek-R1-Distill-Qwen-32Btextn<1K0 likes255 downloads1y agoHugging Face30BAAI /JudgeLM-100K Dataset Card for JudgeLM Dataset Summary JudgeLM-100K dataset contains 100,000 judge samples for training and 5,000 judge samples for validation. All the judge samples have the GPT-4-generated high-quality judgements. This instruction data can be used to conduct instruction-tuning for language models and make the language model has ability to judge open-ended answer pairs. See more details in the "Dataset" section and the appendix sections of this paper. This produced a… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/JudgeLM-100K.text-generation52 likes246 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.