CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gililior /mmlu-prox-eval-predictions MMLU-ProX Multilingual Model Predictions Raw per-sample model predictions on MMLU-ProX across 29 languages and 25 open-weight LLMs, produced with lm-evaluation-harness. This dataset releases the full prediction logs (not just aggregate scores) so that item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling of multilingual benchmarks, error analysis, or per-item difficulty estimation. Repository structure mmlu_prox_<lang>/ └──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.tabularquestion-answering1M<n<10M0 likes33k downloads3mo agoHugging Face02zhouzypaul /auto_evaltext1K<n<10K0 likes16k downloads3mo agoHugging Face03wis-k /instruction-following-evaltextn<1K10 likes4.3k downloads3y agoHugging Face04evalstate /all-defectstabularn<1K2 likes4.1k downloads5mo agoHugging Face05yongxin2020 /TempPerturb-Eval-data TempPerturb-Eval-data Summary TempPerturb-Eval-data is the released output dataset for TempPerturb-Eval, a benchmark for analyzing the robustness of Retrieval-Augmented Generation (RAG) systems under both internal variation and external perturbation. This is an evaluation-artifact dataset: it stores model outputs and experiment metadata for controlled robustness analysis, rather than a new QA training corpus. The release covers: 5 models 11 temperatures from 0.0 to 2.0 4… See the full description on the dataset page: https://huggingface.co/datasets/yongxin2020/TempPerturb-Eval-data.textquestion-answering10K<n<100K1 likes3.6k downloads6mo agoHugging Face06Anthropic /model-written-evals Model-Written Evaluation Datasets This repository includes datasets written by language models, used in our paper on "Discovering Language Model Behaviors with Model-Written Evaluations." We intend the datasets to be useful to: Those who are interested in understanding the quality and properties of model-generated data Those who wish to use our datasets to evaluate other models for the behaviors we examined in our work (e.g., related to model persona, sycophancy, advanced AI risks… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/model-written-evals.textmultiple-choice1K<n<10K68 likes3.4k downloads4y agoHugging Face07MERA-evaluation /MERA MERA (Multimodal Evaluation for Russian-language Architectures) Summary MERA (Multimodal Evaluation for Russian-language Architectures) is a new open independent benchmark for the evaluation of SOTA models for the Russian language. The MERA benchmark unites industry and academic partners in one place to research the capabilities of fundamental models, draw attention to AI-related issues, foster collaboration within the Russian Federation and in the international arena… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/MERA.text10K<n<100K11 likes3.1k downloads2y agoHugging Face08zjunlp /Chat2Workflow-Evaluation Chat2Workflow Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions. Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Repository: zjunlp/Chat2Workflow Overview Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.documenttext-generationn<1K4 likes3.1k downloads4mo agoHugging Face09scimdr /SciMDR-Evalimagequestion-answeringn<1K1 likes3k downloads6mo agoHugging Face10mmathys /openai-moderation-api-evaluation Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection" The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper. Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label. Category Label Definition sexual S Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.tabulartext-classification1K<n<10K38 likes2.8k downloads3y agoHugging Face11NJU-LINK /DR3-EvalDR3-Eval: Towards Realistic and ReproducibleDeep Research Evaluation ✨ Overview DR³-Eval is a realistic, reproducible, and multimodal evaluation benchmark for Deep Research Agents, focusing on multi-file report generation tasks. Existing benchmarks face a fundamental tension between realism, controllability, and reproducibility when evaluating deep research agents. DR³-Eval addresses this through the following design:… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/DR3-Eval.documenttext-generationn<1K2 likes2.7k downloads5mo agoHugging Face12mtabvqa /MTabVQA-Eval Dataset Card for MTabVQA Paper Dataset Description Dataset Summary MTabVQA (Multi-Tabular Visual Question Answering) is a novel benchmark designed to evaluate the ability of Vision-Language Models (VLMs) to perform multi-hop reasoning over multiple tables presented as images. This scenario is common in real-world documents like web pages and PDFs but is critically under-represented in existing benchmarks. The dataset consists of two main parts: MTabVQA-Eval:… See the full description on the dataset page: https://huggingface.co/datasets/mtabvqa/MTabVQA-Eval.imagetable-question-answering1K<n<10K2 likes2.4k downloads11mo agoHugging Face13neuralmagic /quantized-llama-3.1-leaderboard-v2-evals Open LLM Leaderboard v2 Benchmark Results This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models. These evaluations were produced with lm-evaluation-harness by running the following command: lm_eval \ --model vllm \ --model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \ --apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.tabular100K<n<1M0 likes2.1k downloads2y agoHugging Face14evalstate /transformers-merge-experimentstabularn<1K3 likes1.7k downloads5mo agoHugging Face15YWZBrandon /officeqa-checkpoint-eval-data Checkpoint evaluation plot data Snapshot: 2026-09-14T16:26:45.684890+00:00. Aggregate inputs to notes/Sept-2-2026.md performance figures. No model execution, grading, publication, or source-result changes were performed to make this export. Contents checkpoint_evaluations: 454 checkpoint rows, one evaluation per run/iteration/protocol; score, mean output tokens, mean steps, and the existing two-sided 95% confidence bounds. pareto_points: current mean-token/USD… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/officeqa-checkpoint-eval-data.tabular10K<n<100K0 likes1.5k downloads8d agoHugging Face16math-eval /TAL-SCQ5KTAL-SCQ5K Dataset Description Dataset Summary TAL-SCQ5K-EN/TAL-SCQ5K-CN are high quality mathematical competition datasets in English and Chinese language created by TAL Education Group, each consisting of 5K questions(3K training and 2K testing). The questions are in the form of multiple-choice and cover mathematical topics at the primary,junior high and high school levels. In addition, detailed solution steps are provided to facilitate CoT training and all the… See the full description on the dataset page: https://huggingface.co/datasets/math-eval/TAL-SCQ5K.text10K<n<100K60 likes1.4k downloads3y agoHugging Face17IS2Lab /S-Eval S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models 🏆 Leaderboard 🔔 Updates 📣 [2025/10/09]: We update the evaluation for the latest LLMs in 🏆 LeaderBoard, and further release Octopus, an automated LLM safety evaluator, to meet the community’s need for accurate and reproducible safety assessment tools. You can download the model from HuggingFace or ModelScope. 📣 [2025/03/30]: 🎉 Our paper has been accepted by ISSTA 2025. To meet… See the full description on the dataset page: https://huggingface.co/datasets/IS2Lab/S-Eval.texttext-generation100K<n<1M17 likes1.4k downloads7mo agoHugging Face18Anthropic /discrim-eval Dataset Card for Discrim-Eval Dataset Summary The data contains a diverse set of prompts covering 70 hypothetical decision scenarios, ranging from approving a loan to providing press credentials. Each prompt instructs the model to make a binary decision (yes/no) about a particular person described in the prompt. Each person is described in terms of three demographic attributes: age (ranging from 20 to 100 in increments of 10), gender (male, female, non-binary) , and race… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/discrim-eval.tabularquestion-answering10K<n<100K60 likes1.1k downloads3y agoHugging Face19re-align /just-eval-instruct Just Eval Instruct Highlights Data sources: AlpacaEval (covering 5 datasets), LIMA-test, MT-bench, Anthropic red-teaming, and MaliciousInstruct. 1K examples: 1,000 instructions, including 800 for problem-solving test, and 200 specifically for safety test. Category: We tag each example with (one or multiple) labels on its task types and topics.… See the full description on the dataset page: https://huggingface.co/datasets/re-align/just-eval-instruct.text10K<n<100K34 likes1k downloads3y agoHugging Face20glayguo /evalarc-casebook EvalArc Casebook The same 93.75% score can pass one acceptance gate and fail another. Inspect the rules, actual failed checks and original Docker records in a filterable table. This is the data companion to the interactive evidence lab. In the default suite_jobs view, compare support-partial and support-protected. Both use the same frozen defective policy, score 93.75% and fully resolve 0/2 attempts. The deliberately permissive rule accepts partial progress; the rule requiring… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/evalarc-casebook.tabularothern<1K0 likes899 downloads2d agoHugging Face21evalstate /test-traces Test Traces Codex-style rollout JSONL traces exported from fast-agent for validation against the Hugging Face Agent Trace Viewer. tabularn<1K1 likes806 downloads4mo agoHugging Face22mlfoundations /tabula-8b-eval-suiteEvaluation suite used in our paper "Large Scale Transfer Learning for Tabular Data via Language Modeling." This suite includes our preprocessed versions of benchmark datasets except the AutoML Multimodal Benchmark, which can be accessed by following the installation instructions in their repo here. We recommend using rtfm when evaluating models with these datasets. See the rtfm repo for more information on using this data for evaluation. texttabular-classification10K<n<100K6 likes788 downloads2y agoHugging Face23evaluate /conll2003-citextn<1K0 likes730 downloads4y agoHugging Face24nvidia /compute-eval Dataset Card for ComputeEval ComputeEval is a benchmark for evaluating LLM-generated CUDA code on correctness and performance. Each problem provides a self-contained programming challenge — spanning kernels, runtime APIs, memory management, parallel algorithms, GPU libraries, and template/DSL kernel authoring — with a held-out test harness for functional validation and optional performance benchmarks for measuring GPU execution time against a baseline. Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/compute-eval.tabulartext-generation1K<n<10K30 likes720 downloads8d agoHugging Face25nyu-dice-lab /lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.tabular100K<n<1M0 likes683 downloads2y agoHugging Face26sci-m-wang /C4-Eval C4-Eval C4-Eval is the evaluation set for C4 Bench, a Chengyu-based benchmark for measuring whether multimodal language models can understand cross-concept creativity. The release contains the original images, the corresponding idiom answers, and the complete task-specific questions used for evaluation. 221 base items: 37 human-designed seed figures and 184 bridge-controlled synthetic figures. 1,105 evaluation instances: five task formulations for every base item. Language:… See the full description on the dataset page: https://huggingface.co/datasets/sci-m-wang/C4-Eval.imageimage-text-to-text1K<n<10K0 likes674 downloads1mo agoHugging Face27aisa-group /EvalAwareBenchEvalAwareBench Changling Li1,3, Terry Jingchen Zhang6, Jie Zhang1 Zhijing Jin3,5,6, Sahar Abdelnabi2,3,4, Maksym Andriushchenko2,3,4 1ETH Zürich, 2ELLIS Institute Tübingen, 3Max Planck Institute for Intelligent Systems, 4Tübingen AI Center, 5University of Toronto, 6Vector Institute Dataset Summary A factor-controlled benchmark for studying evaluation awareness in language models, where eight psychology-grounded trigger factors can be independently… See the full description on the dataset page: https://huggingface.co/datasets/aisa-group/EvalAwareBench.texttext-generation100K<n<1M7 likes646 downloads4mo agoHugging Face28shisa-ai /eval-IFBench-results IFBench Evaluation Results This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following. Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos: eval-IFBench-results - Model evaluation outputs (this repo) eval-IFBench-prompts - Test prompts/questions (if separated) Dataset Structure Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.texttext-generation10K<n<100K0 likes633 downloads3mo agoHugging Face29PKU-Alignment /BeaverTails-Evaluation Dataset Card for BeaverTails-Evaluation BeaverTails is an AI safety-focused collection comprising a series of datasets. This repository contains test prompts specifically designed for evaluating language model safety. It is important to note that although each prompt can be connected to multiple categories, only one category is labeled for each prompt. The 14 harm categories are defined as follows: Animal Abuse: This involves any form of cruelty or harm inflicted on animals… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails-Evaluation.texttext-classificationn<1K15 likes632 downloads3y agoHugging Face30minhhien0811 /java_evaluation_benchmarkstext1K<n<10K0 likes610 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.