CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gililior /mmlu-prox-eval-predictions MMLU-ProX Multilingual Model Predictions Raw per-sample model predictions on MMLU-ProX across 29 languages and 25 open-weight LLMs, produced with lm-evaluation-harness. This dataset releases the full prediction logs (not just aggregate scores) so that item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling of multilingual benchmarks, error analysis, or per-item difficulty estimation. Repository structure mmlu_prox_<lang>/ └──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.tabularquestion-answering1M<n<10M0 likes34k downloads3mo agoHugging Face02evalstate /all-defectstabularn<1K2 likes4.1k downloads5mo agoHugging Face03mmathys /openai-moderation-api-evaluation Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection" The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper. Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label. Category Label Definition sexual S Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.tabulartext-classification1K<n<10K38 likes2.8k downloads3y agoHugging Face04neuralmagic /quantized-llama-3.1-leaderboard-v2-evals Open LLM Leaderboard v2 Benchmark Results This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models. These evaluations were produced with lm-evaluation-harness by running the following command: lm_eval \ --model vllm \ --model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \ --apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.tabular100K<n<1M0 likes2.1k downloads2y agoHugging Face05evalstate /transformers-merge-experimentstabularn<1K3 likes1.8k downloads5mo agoHugging Face06YWZBrandon /officeqa-checkpoint-eval-data Checkpoint evaluation plot data Snapshot: 2026-09-14T16:26:45.684890+00:00. Aggregate inputs to notes/Sept-2-2026.md performance figures. No model execution, grading, publication, or source-result changes were performed to make this export. Contents checkpoint_evaluations: 454 checkpoint rows, one evaluation per run/iteration/protocol; score, mean output tokens, mean steps, and the existing two-sided 95% confidence bounds. pareto_points: current mean-token/USD… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/officeqa-checkpoint-eval-data.tabular10K<n<100K0 likes1.5k downloads8d agoHugging Face07Anthropic /discrim-eval Dataset Card for Discrim-Eval Dataset Summary The data contains a diverse set of prompts covering 70 hypothetical decision scenarios, ranging from approving a loan to providing press credentials. Each prompt instructs the model to make a binary decision (yes/no) about a particular person described in the prompt. Each person is described in terms of three demographic attributes: age (ranging from 20 to 100 in increments of 10), gender (male, female, non-binary) , and race… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/discrim-eval.tabularquestion-answering10K<n<100K60 likes1.1k downloads3y agoHugging Face08glayguo /evalarc-casebook EvalArc Casebook The same 93.75% score can pass one acceptance gate and fail another. Inspect the rules, actual failed checks and original Docker records in a filterable table. This is the data companion to the interactive evidence lab. In the default suite_jobs view, compare support-partial and support-protected. Both use the same frozen defective policy, score 93.75% and fully resolve 0/2 attempts. The deliberately permissive rule accepts partial progress; the rule requiring… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/evalarc-casebook.tabularothern<1K0 likes900 downloads3d agoHugging Face09evalstate /test-traces Test Traces Codex-style rollout JSONL traces exported from fast-agent for validation against the Hugging Face Agent Trace Viewer. tabularn<1K1 likes791 downloads4mo agoHugging Face10nvidia /compute-eval Dataset Card for ComputeEval ComputeEval is a benchmark for evaluating LLM-generated CUDA code on correctness and performance. Each problem provides a self-contained programming challenge — spanning kernels, runtime APIs, memory management, parallel algorithms, GPU libraries, and template/DSL kernel authoring — with a held-out test harness for functional validation and optional performance benchmarks for measuring GPU execution time against a baseline. Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/compute-eval.tabulartext-generation1K<n<10K30 likes727 downloads8d agoHugging Face11nyu-dice-lab /lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.tabular100K<n<1M0 likes697 downloads2y agoHugging Face12nyu-dice-lab /lm-eval-results-AurelPx-Pegasus-7b-slerp-private Dataset Card for Evaluation run of AurelPx/Pegasus-7b-slerp Dataset automatically created during the evaluation run of model AurelPx/Pegasus-7b-slerp The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-AurelPx-Pegasus-7b-slerp-private.tabular100K<n<1M0 likes564 downloads2y agoHugging Face13akhadangi /NoiseFiT-lm-eval-results Dataset Card for Evaluation run of akhadangi/Mistral-7B-v0.1-0.001 Dataset automatically created during the evaluation run of model akhadangi/Mistral-7B-v0.1-0.001 The dataset is composed of 1536 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 93 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/akhadangi/NoiseFiT-lm-eval-results.tabular1M<n<10M0 likes550 downloads1y agoHugging Face14nyu-dice-lab /lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0 Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.tabular100K<n<1M0 likes515 downloads2y agoHugging Face15giskard-bot /evaluator-leaderboardtabularn<1K0 likes452 downloads2y agoHugging Face16evalitahf /word_in_contextDataset homepage: https://wic-ita.github.io/index.html tabulartext-classification1K<n<10K0 likes450 downloads2y agoHugging Face17nyu-dice-lab /lm-eval-results-s3nh-Severusectum-7B-DPO-private Dataset Card for Evaluation run of s3nh/Severusectum-7B-DPO Dataset automatically created during the evaluation run of model s3nh/Severusectum-7B-DPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-s3nh-Severusectum-7B-DPO-private.tabular100K<n<1M0 likes425 downloads2y agoHugging Face18nyu-dice-lab /lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private Dataset Card for Evaluation run of hkust-nlp/dart-math-llama3-8b-prop2diff Dataset automatically created during the evaluation run of model hkust-nlp/dart-math-llama3-8b-prop2diff The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private.tabular100K<n<1M0 likes416 downloads2y agoHugging Face19nyu-dice-lab /lm-eval-results-shyamieee-JARVIS-v2.0-private Dataset Card for Evaluation run of shyamieee/JARVIS-v2.0 Dataset automatically created during the evaluation run of model shyamieee/JARVIS-v2.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-JARVIS-v2.0-private.tabular100K<n<1M0 likes401 downloads2y agoHugging Face20evalitahf /hatespeech_detection HaSpeeDe2 The HaSpeeDe2 dataset collects 8,012 tweets and 500 news headlines annotated for the presence of hate speech, stereotypes and nominal utterance. The dataset has been used in the context of the HaSpeeDe task (http://www.di.unito.it/~tutreeb/haspeede-evalita20/index.html), organized as part of the EVALITA 2020 evaluation campaign (http://www.evalita.it/2020). In order to meet the GDPR requirements, texts have been pseudonymized replacing all original IDs in both datasets… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/hatespeech_detection.tabulartext-classification10K<n<100K0 likes393 downloads2y agoHugging Face21Rakancorle1 /vtg-bench-frontier-eval Video temporal grounding with frontier models — frame-based eval (Sept 2026) Can a frontier multimodal model, given a video and a question, return the answer and the time interval where it happens? Four models were tested with byte-identical inputs on 351 human-annotated questions over 178 YouTube videos (dashcam, SNL sketches, talk shows, vlogs, soccer and basketball highlights): Claude Fable 5.1, Claude Fable 5, GPT-6 Astra and GPT-5.6 Sol (all via Amazon Bedrock). Neither… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/vtg-bench-frontier-eval.tabularn<1K0 likes388 downloads5d agoHugging Face22nyu-dice-lab /lm-eval-results-yleo-EmertonMonarch-7B-slerp-private Dataset Card for Evaluation run of yleo/EmertonMonarch-7B-slerp Dataset automatically created during the evaluation run of model yleo/EmertonMonarch-7B-slerp The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-yleo-EmertonMonarch-7B-slerp-private.tabular100K<n<1M0 likes379 downloads2y agoHugging Face23nyu-dice-lab /lm-eval-results-shadowml-WestBeagle-7B-private Dataset Card for Evaluation run of shadowml/WestBeagle-7B Dataset automatically created during the evaluation run of model shadowml/WestBeagle-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shadowml-WestBeagle-7B-private.tabular100K<n<1M0 likes352 downloads2y agoHugging Face24nyu-dice-lab /lm-eval-results-bunnycore-SmartToxic-7B-private Dataset Card for Evaluation run of bunnycore/SmartToxic-7B Dataset automatically created during the evaluation run of model bunnycore/SmartToxic-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-bunnycore-SmartToxic-7B-private.tabular100K<n<1M0 likes339 downloads2y agoHugging Face25nyu-dice-lab /lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private.tabular100K<n<1M0 likes337 downloads2y agoHugging Face26nyu-dice-lab /lm-eval-results-automerger-Inex12Yamshadow-7B-private Dataset Card for Evaluation run of automerger/Inex12Yamshadow-7B Dataset automatically created during the evaluation run of model automerger/Inex12Yamshadow-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-automerger-Inex12Yamshadow-7B-private.tabular100K<n<1M0 likes321 downloads2y agoHugging Face27youdotcom /minimax-m3-deepsearchqa-skill-eval MiniMax M3 DeepSearchQA Skill Eval Evaluates minimax/minimax-m3 on google/deepsearchqa using a Pi agent, You.com MCP tools, and a research skill optimized for this harness, model, and tool surface. MiniMax M3 Medium Reasoning with the You.com research skill reached 74.85% adjusted F1 on DeepSearchQA, above the paper's GPT-5 High Reasoning F1 result. Public artifacts are available for inspection and reproduction. Links GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/youdotcom/minimax-m3-deepsearchqa-skill-eval.tabularquestion-answering1K<n<10K1 likes309 downloads12d agoHugging Face28nyu-dice-lab /lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v3.0 Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v3.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private.tabular100K<n<1M0 likes302 downloads2y agoHugging Face29compass-group-tue /sdf_evaluation_traits Models That Know How Evaluations Are Designed Score Safer This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.tabulartext-generation10K<n<100K1 likes302 downloads4mo agoHugging Face30nyu-dice-lab /lm-eval-results-abideen-AlphaMonarch-daser-private Dataset Card for Evaluation run of abideen/AlphaMonarch-daser Dataset automatically created during the evaluation run of model abideen/AlphaMonarch-daser The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-abideen-AlphaMonarch-daser-private.tabular100K<n<1M0 likes289 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.