CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.1k downloads2d agoHugging Face02noel7Y /data-agent-benchmarks LongHorizon Full Data-Agent Benchmarks Companion data artifacts for five complete evaluation tracks: DataSciBench full55 / 167 metric entries DABStep full450 DABStep-Research full100 DSBench Modeling full74 LongDS full68 / 2,225 turns The companion GitHub repository contains processed manifests, evaluation code, historical API ReAct baseline code, download/preparation tools, and the frozen source lock. artifact_manifest.json records every uploaded object's size, SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.imagequestion-answering1 likes402 downloads12d agoHugging Face03yshenaw /SkillOpt_Lite_Benchmarks SkillOpt_Lite Benchmarks Train / val / test splits used by the SkillOpt_Lite project. One multi-config repo containing all six benchmarks: Config Rows (train / val / test) Content shipped searchqa 400 / 200 / 1400 Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA. docvqa 107 / 53 / 374 Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.imagequestion-answering1K<n<10K0 likes390 downloads3mo agoHugging Face04omegaprime669 /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.imagetext-generationn<1K0 likes255 downloads2mo agoHugging Face05CopyleftCultivars /qwen3.6-35b-a3b-chemistry-benchmarks Qwen3.6-35B-A3B Chemistry Benchmark Results Raw outputs and scores from running Qwen3.6-35B-A3B (Q8_0 quant) through five published chemistry and biosecurity benchmarks, entirely on local hardware (two secondhand Tesla M40 24GB GPUs, no cloud compute). This is the raw data behind our blog post on locally reproducible AI capability evaluation, including the full per-item outputs, the parsing failures, and the negative results, not just the headline numbers. Results at… See the full description on the dataset page: https://huggingface.co/datasets/CopyleftCultivars/qwen3.6-35b-a3b-chemistry-benchmarks.question-answering0 likes179 downloads1mo agoHugging Face06nyamtulla /benchmarking-the-benchmarks-data Benchmarking the Benchmarks — Raw SLM Safety Evaluation Runs Raw evaluation data for the ESORICS 2026 paper: Nyamtulla Shaik, Fengjun Li, Bo Luo — University of Kansas Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models. ESORICS 2026. arXiv:2608.17183 Code and processed data: https://github.com/nyamtulla/benchmarking-the-benchmarks ⚠️ Content warning This dataset contains adversarial safety prompts and model responses… See the full description on the dataset page: https://huggingface.co/datasets/nyamtulla/benchmarking-the-benchmarks-data.text-generation100K<n<1M2 likes173 downloads19d agoHugging Face07LemOneLabs /OMNIX_Benchmarks_Latest 📊 OMNIX Benchmark Comparison Table Model Overall Score (Grade) Format Adherence Logical Reasoning Knowledge Recall Constraint Following First-Pass SR Eventual SR FCI Avg Latency (ms) qwen-3-4b-q4 92/100 (A) 97/100 77/100 100/100 95/100 89% 95% 0.56 16140 gemma-4-e4b-q4 89/100 (B) 100/100 67/100 100/100 94/100 94% 94% 0.22 34308 qwen-2.5-coder-3b-text 76/100 (C) 94/100 50/100 95/100 68/100 72% 87% 1.11 3601 llama-3.2-3b-q4 75/100 (C) 89/100 37/100 98/100 84/100… See the full description on the dataset page: https://huggingface.co/datasets/LemOneLabs/OMNIX_Benchmarks_Latest.documenttext-generationn<1K0 likes148 downloads3mo agoHugging Face08humanlong /emotion-negotiation-benchmarks Emotion-Aware LLM Negotiation Benchmarks Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency. The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.tabulartext-generationn<1K0 likes137 downloads4mo agoHugging Face09ArsSocratica /egora-benchmarks EgoRA Benchmark Results Comprehensive benchmark results for EgoRA (Entropy-Governed Orthogonality Regularization for Adaptation) across multiple model scales, domains, and architectures. 📦 Package: egora on PyPI 💻 Code: ArsSocratica/EgoRA on GitHub 📄 Paper: arXiv:2602.05192 🔖 DOI: 10.5281/zenodo.19398709 Dataset Structure llama-3.2-1b/, llama-3.2-3b/, llama-3.1-8b/ Fine-tuning results across 3 model scales, 2 domains (Alpaca general, Medical), 4… See the full description on the dataset page: https://huggingface.co/datasets/ArsSocratica/egora-benchmarks.text-generation1K<n<10K0 likes125 downloads6mo agoHugging Face10axjns /llmfit-benchmarks llmfit Real-World LLM Inference Benchmarks An evolving dataset of real-world LLM inference measurements across consumer, workstation, datacenter, and unified-memory hardware. It combines benchmarks from an external community source with community benchmarks contributed directly to llmfit. The initial release contains 1,501 normalized observations: 1,010 unique external-community observations from the repository's 2026-08-10 snapshot. 491 llmfit-community observations from 61… See the full description on the dataset page: https://huggingface.co/datasets/axjns/llmfit-benchmarks.tabulartext-generation1K<n<10K2 likes108 downloads1mo agoHugging Face11RLAIF-V /RLPR-Benchmarks Dataset Card for RLPR-Test GitHub | Paper News: [2025.06.23] 📃 Our paper detailing the RLPR framework and its comprehensive evaluation using this suite is accessible at arXiv! Dataset Summary We include the following seven benchmarks for evaluation of RLPR: Mathematical Reasoning Benchmarks: MATH-500 (Cobbe et al., 2021) Minerva (Lewkowycz et al., 2022) AIME24 General Domain Reasoning Benchmarks: MMLU-Pro (Wang et al., 2024): A multitask language… See the full description on the dataset page: https://huggingface.co/datasets/RLAIF-V/RLPR-Benchmarks.textquestion-answeringn<1K0 likes104 downloads1y agoHugging Face12XReyRobert /smoke24-agentic-benchmarks Smoke24 Agentic Benchmarks Public, reproducible Terminal-Bench 2.0 Smoke24 benchmark artifacts for local RTX 3090-class agentic model serving. Why Smoke24 I created the Smoke24 subset because I needed a relatively quick benchmark that could run locally on an RTX 3090-class machine in a couple of hours. The goal is to get a practical read on model performance and stability under a real agentic Terminal-Bench workload, without paying the turnaround cost of a much… See the full description on the dataset page: https://huggingface.co/datasets/XReyRobert/smoke24-agentic-benchmarks.text-generation0 likes97 downloads24d agoHugging Face13LostGentoo /hf-inference-endpoint-benchmarks Raw benchmark result files Raw JSON outputs from the sessions described in benchmarking-methodology.md. Model: Qwen3.5-4B family, hf-endpoints deployed via the configs documented in cli-and-api.md. Short-prompt decode comparison (valid metric — prompt negligible vs output, see trap #1 in methodology doc) File Setup llamacpp_results.json llama.cpp, GGUF Q8_0, MTP, A10G vllm_results.json vLLM, FP8-dynamic, MTP, A10G vllm_bf16_results.json vLLM, bf16… See the full description on the dataset page: https://huggingface.co/datasets/LostGentoo/hf-inference-endpoint-benchmarks.text-generationn<1K1 likes97 downloads2mo agoHugging Face14manulife /tam-benchmarks Tasks over Application Manuals (TAM) TAM is a benchmark for evaluating long-horizon procedural reasoning: the ability of a language-model system to follow a large application manual, resolve cross-references, apply interdependent constraints, and produce an exact answer. Unlike short-horizon multi-hop tasks, TAM requires systems to maintain consistency across dozens of decisions drawn from manuals containing tens of thousands of rules. An early missed exception or incorrect… See the full description on the dataset page: https://huggingface.co/datasets/manulife/tam-benchmarks.documenttext-generationn<1K0 likes87 downloads11d agoHugging Face15Abdulrahmankalil /enterprise-llm-inference-benchmarks-2026 🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide) A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments. 🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation) Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.tabulartext-generationn<1K1 likes82 downloads5d agoHugging Face16RISys-Lab /Benchmarks_CyberSec_SECURE Dataset Card for SECURE (RISys-Lab Mirror) ⚠️ Disclaimer: > This repository is a mirror/re-host of the original SECURE benchmark.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below. Repository Intent This Hugging Face dataset is a re-host of the original SECURE benchmark. It has been converted… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SECURE.texttext-classification1K<n<10K0 likes71 downloads8mo agoHugging Face17G3nadh /dgx-spark-benchmarks DGX Spark LLM Benchmarks First comprehensive benchmark suite for NVIDIA DGX Spark (GB10 Blackwell). Hardware GPU: NVIDIA GB10 Blackwell (1 PFLOP FP4) Memory: 128GB unified LPDDR5x (273 GB/s) CPU: 20-core ARM (10x Cortex-X925 + 10x Cortex-A725) Storage: 4TB NVMe Framework: Ollama 0.18.3 CUDA: 13.0 | Driver: 580.142 Benchmark Results Run 1 — General Inference (11 models) Model Size Prompt tok/s Gen tok/s Load Time Llama 3.1 8B 4.9 GB… See the full description on the dataset page: https://huggingface.co/datasets/G3nadh/dgx-spark-benchmarks.texttext-generationn<1K1 likes68 downloads6mo agoHugging Face18lllouo /BD-benchmarks BD-benchmarks: Denoised Benchmark Datasets Dataset Description BD-benchmarks is a comprehensive collection of denoised versions of popular NLP benchmark datasets. This repository contains original and cleaned versions of 11 widely-used benchmarks processed using two state-of-the-art denoising methods: DeepSeek-R1 and WAC-GEC (Whitespace Anomaly Correction - Grammar Error Correction). Dataset Summary This dataset addresses the critical issue of noise in… See the full description on the dataset page: https://huggingface.co/datasets/lllouo/BD-benchmarks.textquestion-answering10K<n<100K0 likes65 downloads8mo agoHugging Face19LemOneLabs /OMNIX_Benchmarks_7-6-26 📊 OMNIX Benchmark Comparison Table Model Overall Score (Grade) Format Adherence Logical Reasoning Knowledge Recall Constraint Following First-Pass SR Eventual SR FCI Avg Latency (ms) qwen-3-4b-q4 92/100 (A) 97/100 77/100 100/100 95/100 89% 95% 0.56 16140 llama-3.2-3b-q4 75/100 (C) 89/100 37/100 98/100 84/100 78% 84% 0.89 5559 bonsai-8b-q4 63/100 (D) 89/100 13/100 88/100 76/100 75% 76% 1.11 9721 Note: SR = Success Rate. FCI (Friction Correction Index)… See the full description on the dataset page: https://huggingface.co/datasets/LemOneLabs/OMNIX_Benchmarks_7-6-26.documenttext-generationn<1K0 likes59 downloads3mo agoHugging Face20callensxavier /runux-tpu-v5e-benchmarks ⚡ RunuX-AI — TPU v5e Inference Benchmarks Achieving 3× Throughput & 3× Energy Reduction on Google TPU v5e Xavier Callens · Socrate AI Lab (Non-Profit) Reproducible benchmark data & scripts — No proprietary code included 🎯 What Is This? This repository contains benchmark results and Apache-2.0 reproduction scripts for comparing LLM inference performance across 5 frameworks on Google TPU v5e. The goal is to enable independent verification of our claims… See the full description on the dataset page: https://huggingface.co/datasets/callensxavier/runux-tpu-v5e-benchmarks.texttext-generationn<1K0 likes57 downloads4mo agoHugging Face21studioburnside /mlx-local-inference-benchmarks MLX local-inference benchmarks — Qwen3.6 & Laguna-S/XS families Raw results, harnesses and methodology for an 8-axis benchmark of four MLX checkpoints on a 128 GB M5 Max. Everything a person would need to check my numbers or disagree with them. Companion model repos: Tess-4-27B-MLX-Q8 — with a working MTP head Tess-4-27B-MLX-Q4 — same, at 4-bit NEW (2026-07-24): the Laguna chapter — REPORT-LAGUNA.md + results-laguna/ Five-way same-engine bake-off (Laguna-S… See the full description on the dataset page: https://huggingface.co/datasets/studioburnside/mlx-local-inference-benchmarks.text-generation2 likes53 downloads2mo agoHugging Face22Hzou9 /PCBSchemaGen-Benchmarks PCBSchemaGen Benchmarks Two benchmark suites for LLM-driven PCB schematic synthesis, from the paper PCBSchemaGen: Reward-Guided LLM Code Synthesis for Printed Circuit Board (PCB) Schematic Design with Structured Verification. Correctness in this domain is not defined by unit tests: there are no per-task golden references, and SPICE does not validate schematic-level correctness. Instead, each task is scored by a deterministic structural verifier against real-IC pin- and… See the full description on the dataset page: https://huggingface.co/datasets/Hzou9/PCBSchemaGen-Benchmarks.tabulartext-generationn<1K0 likes45 downloads2mo agoHugging Face23zhangdw /Anchor-benchmarks 🧠 Anchor Benchmarks A curated long-term memory benchmark bundle for LLM and agent evaluation &nbsp;&nbsp;&nbsp;&nbsp; Anchor Benchmarks packages three public long-term memory evaluation resources for studying factual recall, temporal reasoning, knowledge update, multi-hop inference, and multimodal conversational memory. Quick Start · At a Glance · Benchmarks · Evaluation · Citation [!IMPORTANT] This repository is a benchmark bundle, not a new… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/Anchor-benchmarks.question-answering1K<n<10K0 likes40 downloads4mo agoHugging Face24hmnshudhmn24 /llm-benchmarks-capabilities-2020-2026 📊 LLM Benchmarks & Capabilities 2020–2026 The most comprehensive open dataset tracking the evolution of Large Language Models — from GPT-3 to GPT-5.5, Claude Opus 4.7, Gemini 3.5, and beyond. 🧭 Overview This dataset captures the complete LLM landscape from 2020 to 2026 across five dimensions: 🤖 113 models from 25+ organizations 📈 17 benchmarks tracking capability growth over time 💰 Monthly API pricing showing 100x+ cost reductions ⚙️ Training compute… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/llm-benchmarks-capabilities-2020-2026.tabulartext-generation1 likes40 downloads3mo agoHugging Face25kevo666 /packrat-benchmarks PackRat v2 Benchmarks Version: 2.0.0 Date: 2026-04-10 Tokenizer: tiktoken cl100k_base (GPT-4 / Claude compatible) Platform: Node.js v25.6.1, Windows 11 Summary Metric Result Round-trip accuracy 100% (144/144 tests) Token savings (avg) 2.4% Token savings (best) 17.3% (path/URL-heavy files) Byte savings (avg) 2.5% Search speedup 12.03x Codebook entries 72 (auto-learned) Negative-savings entries 0 Comparison: PackRat vs MemPalace… See the full description on the dataset page: https://huggingface.co/datasets/kevo666/packrat-benchmarks.texttext-generationn<1K0 likes39 downloads5mo agoHugging Face26zevatov /nra-benchmarks 🧬 NRA Benchmark Datasets All benchmark datasets for Neural Ready Archive (NRA) — the Rust-native streaming format for ML training. Train on gigabytes of real data without downloading a single byte. NRA replaces tar.gz and zip for the AI era. 📦 Available Datasets File Domain Source Files Size food-101.nra 🖼️ Vision ethz/food101 101,000 images 4.7 GB wikitext.nra 📝 Text Salesforce/wikitext 23,767 text files 7.6 MB pokemon.nra 🎨 Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/zevatov/nra-benchmarks.image-classification100K<n<1M1 likes39 downloads5mo agoHugging Face27LemOneLabs /OMNIX_Benchmarks_6-30-26 📊 OMNIX Benchmark Comparison Table Model Overall Eventual Score (Grade) Format Adherence (Eventual) Logical Reasoning (Eventual) Knowledge Recall (Eventual) Constraint Following (Eventual) First-Pass Success Rate Eventual Success Rate Friction Correction Index (FCI) Average Request Latency gemma-3 1B 65/100 (D) 97/100 23/100 90/100 59/100 60% 74% 1.56 9545ms gemma-4-e2b-q4 73/100 (C) 100/100 33/100 88/100 83/100 72% 83% 0.78 20575ms gemma-4-e4b-q4 89/100 (B)… See the full description on the dataset page: https://huggingface.co/datasets/LemOneLabs/OMNIX_Benchmarks_6-30-26.documenttext-generationn<1K0 likes39 downloads3mo agoHugging Face28liaialley /latent-sft-eval-benchmarks Latent-SFT Evaluation Benchmarks Processed evaluation datasets for Latent-SFT trajectory generation and model diagnostics. These files are packaged for batch CoT trajectory generation. The prompt should put the final boxed answer only in the generated response / cot_answer; the reasoning-only part should not contain an extra boxed answer. Files file rows keys mmlu_pro_validation.jsonl 70 answer, problem mmlu_pro_validation_audit.jsonl 70 answer… See the full description on the dataset page: https://huggingface.co/datasets/liaialley/latent-sft-eval-benchmarks.question-answering0 likes36 downloads2mo agoHugging Face29ChuGyouk /Qwen3.5-4B-nothink-benchmarks Qwen3.5-4B (non-thinking) — 13 benchmarks, multi-sample outputs with pass@k All sampled outputs of Qwen/Qwen3.5-4B in non-thinking mode (enable_thinking=False) on 13 benchmarks, with per-response correctness and pass@k / avg@n metrics. One row per problem; every row carries the exact prompt that was sent to the model, all sampled responses, their scores, and the benchmark-level metrics. Generation setup (identical for every benchmark) Model… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/Qwen3.5-4B-nothink-benchmarks.texttext-generation10K<n<100K0 likes32 downloads1d agoHugging Face30odyn-network /odyn-benchmarks Odyn Benchmarks Inference benchmark datasets and results for the Odyn Network — a distributed, OpenAI-compatible AI inference platform built on vLLM, Ray Serve, and FastAPI. Dataset Structure Prompt Profiles (data/) Four load profiles covering the full input/output token distribution space, sourced from real Odyn traffic and augmented with ShareGPT Vicuna Unfiltered: Profile Description Input tokens Output tokens Rows A Short input, Long output avg… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/odyn-benchmarks.text-generation1K<n<10K0 likes30 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.