CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /arena-resultsThis dataset contains the saved results from MTEB-Arena tabular1K<n<10K4 likes9k downloads1y agoHugging Face02transformers-community /circleci-test-resultstextn<1K4 likes2k downloads3mo agoHugging Face03manikandan18ramalingam /agentic-ai-options-resultstextn<1K1 likes2k downloads1h agoHugging Face04JetBrains-Research /lca-results Long Code Arena (raw results) These are the raw results from the Long Code Arena benchmark suite, as well as the corresponding model predictions. Please use the subset dropdown menu to select the necessary data relating to our six benchmarks: 🤗 Library-based code generation 🤗 CI builds repair 🤗 Project-level code completion 🤗 Commit message generation🤗 Bug localization 🤗 Module summarization tabularn<1K2 likes2k downloads1y agoHugging Face05eduagarcia-temp /llm_pt_leaderboard_resultstextn<1K0 likes1.9k downloads1y agoHugging Face06OALL /resultstextn<1K0 likes1.7k downloads2y agoHugging Face07mii-llm /pinocchio-resultstextn<1K0 likes1.4k downloads2y agoHugging Face08MicroAGI-Labs /vlm-info-loss-results VLM Grounding Evaluation Results Grounding evaluation results for vision-language models on robotics manipulation datasets. Part of the vlm-info-loss project studying how VLM connectors transform visual representations. Background Our embedding-level analysis shows VLM connectors perform a compress-then-expand transformation: they sharpen dominant-object representations while compressing secondary-object category identity. All tested models converge to ~83%… See the full description on the dataset page: https://huggingface.co/datasets/MicroAGI-Labs/vlm-info-loss-results.imageobject-detectionn<1K0 likes1k downloads5mo agoHugging Face09mii-llm /resultstextn<1K0 likes999 downloads1y agoHugging Face10AverageMetaheuristicsEnjoyer /moe-routing-drift-results MoE routing drift — results Measurements for a 2x2 experiment: adaptation (none / GEPA / prompt-tuning / prefix-tuning) crossed with router retraining (frozen gate / gate retrained), on inclusionAI/Ling-mini-2.0 and Qwen/Qwen3-30B-A3B-Instruct-2507. The weights are in moe-routing-drift-checkpoints. Content warning. quality/*/*.responses.jsonl contain verbatim comments from civil_comments together with model outputs; the task is toxicity labelling, so the text includes insults… See the full description on the dataset page: https://huggingface.co/datasets/AverageMetaheuristicsEnjoyer/moe-routing-drift-results.tabulartext-classificationn<1K1 likes970 downloads11h agoHugging Face11jakeatx /qwen36-kquant-offload-mtp-swebench-lite100-results Qwen3.6 K-Quant Offload MTP SWE-bench Lite 100 Results This dataset contains the complete 5-model x 100-prompt runtime benchmark artifacts plus a detailed statistical analysis layer. Primary conclusion: hot30/cold30 was the best decode-throughput run, while Q4_K_M had the best total wall clock. The ATX hot30/cold30 quantization significantly outperformed both Q4_K_M and Q3_K_XL on paired decode throughput, but Q4_K_M remains the elapsed-time control. The ATX/K3 hot10, hot20, and… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-kquant-offload-mtp-swebench-lite100-results.imagen<1K0 likes806 downloads4mo agoHugging Face12nyu-dice-lab /lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.tabular100K<n<1M0 likes769 downloads2y agoHugging Face13nyu-dice-lab /lm-eval-results-AurelPx-Pegasus-7b-slerp-private Dataset Card for Evaluation run of AurelPx/Pegasus-7b-slerp Dataset automatically created during the evaluation run of model AurelPx/Pegasus-7b-slerp The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-AurelPx-Pegasus-7b-slerp-private.tabular100K<n<1M0 likes600 downloads2y agoHugging Face14nyu-dice-lab /wildchat-50m-extended-resultstabular10K<n<100K1 likes561 downloads2y agoHugging Face15akhadangi /NoiseFiT-lm-eval-results Dataset Card for Evaluation run of akhadangi/Mistral-7B-v0.1-0.001 Dataset automatically created during the evaluation run of model akhadangi/Mistral-7B-v0.1-0.001 The dataset is composed of 1536 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 93 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/akhadangi/NoiseFiT-lm-eval-results.tabular1M<n<10M0 likes559 downloads2y agoHugging Face16nyu-dice-lab /lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0 Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.tabular100K<n<1M0 likes515 downloads2y agoHugging Face17shisa-ai /eval-IFBench-results IFBench Evaluation Results This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following. Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos: eval-IFBench-results - Model evaluation outputs (this repo) eval-IFBench-prompts - Test prompts/questions (if separated) Dataset Structure Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.texttext-generation10K<n<100K0 likes483 downloads3mo agoHugging Face18nyu-dice-lab /lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private Dataset Card for Evaluation run of hkust-nlp/dart-math-llama3-8b-prop2diff Dataset automatically created during the evaluation run of model hkust-nlp/dart-math-llama3-8b-prop2diff The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private.tabular100K<n<1M0 likes480 downloads2y agoHugging Face19djain95 /sae-jailbreaks-resultsimagen<1K1 likes451 downloads5mo agoHugging Face20prakashknaikade /3DGS_Results3dn<1K2 likes444 downloads2y agoHugging Face21nyu-dice-lab /lm-eval-results-s3nh-Severusectum-7B-DPO-private Dataset Card for Evaluation run of s3nh/Severusectum-7B-DPO Dataset automatically created during the evaluation run of model s3nh/Severusectum-7B-DPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-s3nh-Severusectum-7B-DPO-private.tabular100K<n<1M0 likes408 downloads2y agoHugging Face22nyu-dice-lab /lm-eval-results-shyamieee-JARVIS-v2.0-private Dataset Card for Evaluation run of shyamieee/JARVIS-v2.0 Dataset automatically created during the evaluation run of model shyamieee/JARVIS-v2.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-JARVIS-v2.0-private.tabular100K<n<1M0 likes401 downloads2y agoHugging Face23amphora /bon-resultstext10K<n<100K0 likes391 downloads2y agoHugging Face24nyu-dice-lab /lm-eval-results-yleo-EmertonMonarch-7B-slerp-private Dataset Card for Evaluation run of yleo/EmertonMonarch-7B-slerp Dataset automatically created during the evaluation run of model yleo/EmertonMonarch-7B-slerp The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-yleo-EmertonMonarch-7B-slerp-private.tabular100K<n<1M0 likes379 downloads2y agoHugging Face25allenai /href_resultstabularn<1K0 likes376 downloads1y agoHugging Face26nyu-dice-lab /lm-eval-results-bunnycore-SmartToxic-7B-private Dataset Card for Evaluation run of bunnycore/SmartToxic-7B Dataset automatically created during the evaluation run of model bunnycore/SmartToxic-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-bunnycore-SmartToxic-7B-private.tabular100K<n<1M0 likes339 downloads2y agoHugging Face27nyu-dice-lab /lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private.tabular100K<n<1M0 likes337 downloads2y agoHugging Face28cot-leaderboard /cot-leaderboard-resultstextn<1K0 likes334 downloads2y agoHugging Face29nyu-dice-lab /lm-eval-results-shadowml-WestBeagle-7B-private Dataset Card for Evaluation run of shadowml/WestBeagle-7B Dataset automatically created during the evaluation run of model shadowml/WestBeagle-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shadowml-WestBeagle-7B-private.tabular100K<n<1M0 likes334 downloads2y agoHugging Face30false-facts-finetuning /brittleness-results Adapters copied (2026-09-08). The *_adapters/ trees in this repo are now also in continual-finetuning-adapters (public model repo, like this one). Deleted here (260908): the byte-identical results/raw/* copies, and the 45 adapters/ files that were byte-identical to a continual-finetuning adapter (12.3 GB); both lists are in MIGRATION_260908.md of any new repo. Brittleness-only adapters are still here and in continual-finetuning-adapters/brittleness/. Please prefer the new repo for loading.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/brittleness-results.imagen<1K0 likes322 downloads15d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.