datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.data-agent-benchmarks
LongHorizon Full Data-Agent Benchmarks
Companion data artifacts for five complete evaluation tracks:
DataSciBench full55 / 167 metric entries
DABStep full450
DABStep-Research full100
DSBench Modeling full74
LongDS full68 / 2,225 turns
The companion GitHub repository contains processed manifests, evaluation code,
historical API ReAct baseline code, download/preparation tools, and the frozen
source lock. artifact_manifest.json records every uploaded object's size,
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.SkillOpt_Lite_Benchmarks
SkillOpt_Lite Benchmarks
Train / val / test splits used by the SkillOpt_Lite project.
One multi-config repo containing all six benchmarks:
Config
Rows (train / val / test)
Content shipped
searchqa
400 / 200 / 1400
Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA.
docvqa
107 / 53 / 374
Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.qwen3.6-35b-a3b-chemistry-benchmarks
Qwen3.6-35B-A3B Chemistry Benchmark Results
Raw outputs and scores from running Qwen3.6-35B-A3B (Q8_0 quant) through five published chemistry and biosecurity benchmarks, entirely on local hardware (two secondhand Tesla M40 24GB GPUs, no cloud compute). This is the raw data behind our blog post on locally reproducible AI capability evaluation, including the full per-item outputs, the parsing failures, and the negative results, not just the headline numbers.
Results at… See the full description on the dataset page: https://huggingface.co/datasets/CopyleftCultivars/qwen3.6-35b-a3b-chemistry-benchmarks.benchmarking-the-benchmarks-data
Benchmarking the Benchmarks — Raw SLM Safety Evaluation Runs
Raw evaluation data for the ESORICS 2026 paper:
Nyamtulla Shaik, Fengjun Li, Bo Luo — University of Kansas
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models.
ESORICS 2026. arXiv:2608.17183
Code and processed data: https://github.com/nyamtulla/benchmarking-the-benchmarks
⚠️ Content warning
This dataset contains adversarial safety prompts and model responses… See the full description on the dataset page: https://huggingface.co/datasets/nyamtulla/benchmarking-the-benchmarks-data.OMNIX_Benchmarks_Latest
📊 OMNIX Benchmark Comparison Table
Model
Overall Score (Grade)
Format Adherence
Logical Reasoning
Knowledge Recall
Constraint Following
First-Pass SR
Eventual SR
FCI
Avg Latency (ms)
qwen-3-4b-q4
92/100 (A)
97/100
77/100
100/100
95/100
89%
95%
0.56
16140
gemma-4-e4b-q4
89/100 (B)
100/100
67/100
100/100
94/100
94%
94%
0.22
34308
qwen-2.5-coder-3b-text
76/100 (C)
94/100
50/100
95/100
68/100
72%
87%
1.11
3601
llama-3.2-3b-q4
75/100 (C)
89/100
37/100
98/100
84/100… See the full description on the dataset page: https://huggingface.co/datasets/LemOneLabs/OMNIX_Benchmarks_Latest.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.egora-benchmarks
EgoRA Benchmark Results
Comprehensive benchmark results for EgoRA (Entropy-Governed Orthogonality Regularization for Adaptation) across multiple model scales, domains, and architectures.
📦 Package: egora on PyPI
💻 Code: ArsSocratica/EgoRA on GitHub
📄 Paper: arXiv:2602.05192
🔖 DOI: 10.5281/zenodo.19398709
Dataset Structure
llama-3.2-1b/, llama-3.2-3b/, llama-3.1-8b/
Fine-tuning results across 3 model scales, 2 domains (Alpaca general, Medical), 4… See the full description on the dataset page: https://huggingface.co/datasets/ArsSocratica/egora-benchmarks.llmfit-benchmarks
llmfit Real-World LLM Inference Benchmarks
An evolving dataset of real-world LLM inference measurements across consumer, workstation, datacenter, and unified-memory hardware. It combines benchmarks from an external community source with community benchmarks contributed directly to llmfit.
The initial release contains 1,501 normalized observations:
1,010 unique external-community observations from the repository's 2026-08-10 snapshot.
491 llmfit-community observations from 61… See the full description on the dataset page: https://huggingface.co/datasets/axjns/llmfit-benchmarks.RLPR-Benchmarks
Dataset Card for RLPR-Test
GitHub | Paper
News:
[2025.06.23] 📃 Our paper detailing the RLPR framework and its comprehensive evaluation using this suite is accessible at arXiv!
Dataset Summary
We include the following seven benchmarks for evaluation of RLPR:
Mathematical Reasoning Benchmarks:
MATH-500 (Cobbe et al., 2021)
Minerva (Lewkowycz et al., 2022)
AIME24
General Domain Reasoning Benchmarks:
MMLU-Pro (Wang et al., 2024): A multitask language… See the full description on the dataset page: https://huggingface.co/datasets/RLAIF-V/RLPR-Benchmarks.smoke24-agentic-benchmarks
Smoke24 Agentic Benchmarks
Public, reproducible Terminal-Bench 2.0 Smoke24 benchmark artifacts for local RTX 3090-class agentic model serving.
Why Smoke24
I created the Smoke24 subset because I needed a relatively quick benchmark that
could run locally on an RTX 3090-class machine in a couple of hours. The goal is
to get a practical read on model performance and stability under a real agentic
Terminal-Bench workload, without paying the turnaround cost of a much… See the full description on the dataset page: https://huggingface.co/datasets/XReyRobert/smoke24-agentic-benchmarks.hf-inference-endpoint-benchmarks
Raw benchmark result files
Raw JSON outputs from the sessions described in benchmarking-methodology.md. Model: Qwen3.5-4B family, hf-endpoints deployed via the configs documented in cli-and-api.md.
Short-prompt decode comparison (valid metric — prompt negligible vs output, see trap #1 in methodology doc)
File
Setup
llamacpp_results.json
llama.cpp, GGUF Q8_0, MTP, A10G
vllm_results.json
vLLM, FP8-dynamic, MTP, A10G
vllm_bf16_results.json
vLLM, bf16… See the full description on the dataset page: https://huggingface.co/datasets/LostGentoo/hf-inference-endpoint-benchmarks.tam-benchmarks
Tasks over Application Manuals (TAM)
TAM is a benchmark for evaluating long-horizon procedural reasoning: the ability of a language-model system to follow a large application manual, resolve cross-references, apply interdependent constraints, and produce an exact answer. Unlike short-horizon multi-hop tasks, TAM requires systems to maintain consistency across dozens of decisions drawn from manuals containing tens of thousands of rules. An early missed exception or incorrect… See the full description on the dataset page: https://huggingface.co/datasets/manulife/tam-benchmarks.enterprise-llm-inference-benchmarks-2026
🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide)
A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments.
🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation)
Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.Benchmarks_CyberSec_SECURE
Dataset Card for SECURE (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original SECURE benchmark.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original SECURE benchmark. It has been converted… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SECURE.dgx-spark-benchmarks
DGX Spark LLM Benchmarks
First comprehensive benchmark suite for NVIDIA DGX Spark (GB10 Blackwell).
Hardware
GPU: NVIDIA GB10 Blackwell (1 PFLOP FP4)
Memory: 128GB unified LPDDR5x (273 GB/s)
CPU: 20-core ARM (10x Cortex-X925 + 10x Cortex-A725)
Storage: 4TB NVMe
Framework: Ollama 0.18.3
CUDA: 13.0 | Driver: 580.142
Benchmark Results
Run 1 — General Inference (11 models)
Model
Size
Prompt tok/s
Gen tok/s
Load Time
Llama 3.1 8B
4.9 GB… See the full description on the dataset page: https://huggingface.co/datasets/G3nadh/dgx-spark-benchmarks.BD-benchmarks
BD-benchmarks: Denoised Benchmark Datasets
Dataset Description
BD-benchmarks is a comprehensive collection of denoised versions of popular NLP benchmark datasets. This repository contains original and cleaned versions of 11 widely-used benchmarks processed using two state-of-the-art denoising methods: DeepSeek-R1 and WAC-GEC (Whitespace Anomaly Correction - Grammar Error Correction).
Dataset Summary
This dataset addresses the critical issue of noise in… See the full description on the dataset page: https://huggingface.co/datasets/lllouo/BD-benchmarks.OMNIX_Benchmarks_7-6-26
📊 OMNIX Benchmark Comparison Table
Model
Overall Score (Grade)
Format Adherence
Logical Reasoning
Knowledge Recall
Constraint Following
First-Pass SR
Eventual SR
FCI
Avg Latency (ms)
qwen-3-4b-q4
92/100 (A)
97/100
77/100
100/100
95/100
89%
95%
0.56
16140
llama-3.2-3b-q4
75/100 (C)
89/100
37/100
98/100
84/100
78%
84%
0.89
5559
bonsai-8b-q4
63/100 (D)
89/100
13/100
88/100
76/100
75%
76%
1.11
9721
Note: SR = Success Rate. FCI (Friction Correction Index)… See the full description on the dataset page: https://huggingface.co/datasets/LemOneLabs/OMNIX_Benchmarks_7-6-26.runux-tpu-v5e-benchmarks
⚡ RunuX-AI — TPU v5e Inference Benchmarks
Achieving 3× Throughput & 3× Energy Reduction on Google TPU v5e
Xavier Callens · Socrate AI Lab (Non-Profit)
Reproducible benchmark data & scripts — No proprietary code included
🎯 What Is This?
This repository contains benchmark results and Apache-2.0 reproduction scripts for comparing LLM inference performance across 5 frameworks on Google TPU v5e. The goal is to enable independent verification of our claims… See the full description on the dataset page: https://huggingface.co/datasets/callensxavier/runux-tpu-v5e-benchmarks.mlx-local-inference-benchmarks
MLX local-inference benchmarks — Qwen3.6 & Laguna-S/XS families
Raw results, harnesses and methodology for an 8-axis benchmark of four MLX
checkpoints on a 128 GB M5 Max. Everything a person would need to check my numbers
or disagree with them.
Companion model repos:
Tess-4-27B-MLX-Q8 — with a working MTP head
Tess-4-27B-MLX-Q4 — same, at 4-bit
NEW (2026-07-24): the Laguna chapter — REPORT-LAGUNA.md + results-laguna/
Five-way same-engine bake-off (Laguna-S… See the full description on the dataset page: https://huggingface.co/datasets/studioburnside/mlx-local-inference-benchmarks.PCBSchemaGen-Benchmarks
PCBSchemaGen Benchmarks
Two benchmark suites for LLM-driven PCB schematic synthesis, from the paper
PCBSchemaGen: Reward-Guided LLM Code Synthesis for Printed Circuit Board (PCB) Schematic Design with Structured Verification.
Correctness in this domain is not defined by unit tests: there are no per-task golden references,
and SPICE does not validate schematic-level correctness. Instead, each task is scored by a
deterministic structural verifier against real-IC pin- and… See the full description on the dataset page: https://huggingface.co/datasets/Hzou9/PCBSchemaGen-Benchmarks.Anchor-benchmarks
🧠 Anchor Benchmarks
A curated long-term memory benchmark bundle for LLM and agent evaluation
Anchor Benchmarks packages three public long-term memory evaluation resources for studying factual recall, temporal reasoning, knowledge update, multi-hop inference, and multimodal conversational memory.
Quick Start ·
At a Glance ·
Benchmarks ·
Evaluation ·
Citation
[!IMPORTANT]
This repository is a benchmark bundle, not a new… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/Anchor-benchmarks.llm-benchmarks-capabilities-2020-2026
📊 LLM Benchmarks & Capabilities 2020–2026
The most comprehensive open dataset tracking the evolution of Large Language Models — from GPT-3 to GPT-5.5, Claude Opus 4.7, Gemini 3.5, and beyond.
🧭 Overview
This dataset captures the complete LLM landscape from 2020 to 2026 across five dimensions:
🤖 113 models from 25+ organizations
📈 17 benchmarks tracking capability growth over time
💰 Monthly API pricing showing 100x+ cost reductions
⚙️ Training compute… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/llm-benchmarks-capabilities-2020-2026.packrat-benchmarks
PackRat v2 Benchmarks
Version: 2.0.0
Date: 2026-04-10
Tokenizer: tiktoken cl100k_base (GPT-4 / Claude compatible)
Platform: Node.js v25.6.1, Windows 11
Summary
Metric
Result
Round-trip accuracy
100% (144/144 tests)
Token savings (avg)
2.4%
Token savings (best)
17.3% (path/URL-heavy files)
Byte savings (avg)
2.5%
Search speedup
12.03x
Codebook entries
72 (auto-learned)
Negative-savings entries
0
Comparison: PackRat vs MemPalace… See the full description on the dataset page: https://huggingface.co/datasets/kevo666/packrat-benchmarks.nra-benchmarks
🧬 NRA Benchmark Datasets
All benchmark datasets for Neural Ready Archive (NRA) — the Rust-native streaming format for ML training.
Train on gigabytes of real data without downloading a single byte. NRA replaces tar.gz and zip for the AI era.
📦 Available Datasets
File
Domain
Source
Files
Size
food-101.nra
🖼️ Vision
ethz/food101
101,000 images
4.7 GB
wikitext.nra
📝 Text
Salesforce/wikitext
23,767 text files
7.6 MB
pokemon.nra
🎨 Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/zevatov/nra-benchmarks.OMNIX_Benchmarks_6-30-26
📊 OMNIX Benchmark Comparison Table
Model
Overall Eventual Score (Grade)
Format Adherence (Eventual)
Logical Reasoning (Eventual)
Knowledge Recall (Eventual)
Constraint Following (Eventual)
First-Pass Success Rate
Eventual Success Rate
Friction Correction Index (FCI)
Average Request Latency
gemma-3 1B
65/100 (D)
97/100
23/100
90/100
59/100
60%
74%
1.56
9545ms
gemma-4-e2b-q4
73/100 (C)
100/100
33/100
88/100
83/100
72%
83%
0.78
20575ms
gemma-4-e4b-q4
89/100 (B)… See the full description on the dataset page: https://huggingface.co/datasets/LemOneLabs/OMNIX_Benchmarks_6-30-26.latent-sft-eval-benchmarks
Latent-SFT Evaluation Benchmarks
Processed evaluation datasets for Latent-SFT trajectory generation and model diagnostics.
These files are packaged for batch CoT trajectory generation. The prompt should put the final boxed answer only in the generated response / cot_answer; the reasoning-only part should not contain an extra boxed answer.
Files
file
rows
keys
mmlu_pro_validation.jsonl
70
answer, problem
mmlu_pro_validation_audit.jsonl
70
answer… See the full description on the dataset page: https://huggingface.co/datasets/liaialley/latent-sft-eval-benchmarks.Qwen3.5-4B-nothink-benchmarks
Qwen3.5-4B (non-thinking) — 13 benchmarks, multi-sample outputs with pass@k
All sampled outputs of Qwen/Qwen3.5-4B in non-thinking mode (enable_thinking=False)
on 13 benchmarks, with per-response correctness and pass@k / avg@n metrics. One row per problem; every row carries the exact prompt that was
sent to the model, all sampled responses, their scores, and the benchmark-level metrics.
Generation setup (identical for every benchmark)
Model… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/Qwen3.5-4B-nothink-benchmarks.odyn-benchmarks
Odyn Benchmarks
Inference benchmark datasets and results for the Odyn Network — a distributed, OpenAI-compatible AI inference platform built on vLLM, Ray Serve, and FastAPI.
Dataset Structure
Prompt Profiles (data/)
Four load profiles covering the full input/output token distribution space, sourced from real Odyn traffic and augmented with ShareGPT Vicuna Unfiltered:
Profile
Description
Input tokens
Output tokens
Rows
A
Short input, Long output
avg… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/odyn-benchmarks.
