datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
daytrader-benchmarkswhat-ai-benchmarks-actually-measure
What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models
Item-level model responses and scores for 53 language models across the
56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting
Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
(Desai et al., 2026,
arxiv.org/abs/2609.08812).
We do not release the prompts from the benchmark datasets, but instead refer to them by
item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.benchmarks
OpenChainBench Crypto Infrastructure Benchmarks
Daily snapshots of every public benchmark on
openchainbench.com, released as
Hive-partitioned Parquet under CC-BY-4.0.
OCB measures latency, cost, coverage and accuracy of crypto
infrastructure (RPCs, oracles, bridges, data APIs, Polymarket adapters,
Hyperliquid builders). Every snapshot here mirrors the
/api/citable,
/api/stat/<slug>,
and /api/series/<slug>
JSON feeds at the time of capture.
Latest snapshot: 2026-09-22 (captured… See the full description on the dataset page: https://huggingface.co/datasets/OpenChainBench/benchmarks.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.mlx-benchmarks
MLX Benchmarks
Structured benchmark results for MLX-quantized and other locally-hosted
LLMs on Apple Silicon. Covers throughput, time-to-first-token, tool-calling,
code generation, reasoning, knowledge, and math suites.
Results are produced by a sweep harness that wires upstream evaluation tools
against a local vllm-mlx inference server:
EleutherAI/lm-evaluation-harness — coding, reasoning, knowledge, math
linusvwe/MLXBench — throughput and time-to-first-token
vllm… See the full description on the dataset page: https://huggingface.co/datasets/JacobPEvans/mlx-benchmarks.benchmarks
BitRouter Benchmarks
This dataset release contains the private-router Terminal-Bench 2.1 routing study and its audit evidence. July 2026 experiments are preserved under v0/; the current study, raw Harbor traces, run-scoped BitRouter SQLite files, sanitized production-usage backfill, reproducible joins, and Parquet tables are under terminal-bench-2.1/router-random-study/.
Main result
All valid-case metrics below use the same strict intersection of 80 tasks. “Harbor… See the full description on the dataset page: https://huggingface.co/datasets/BitRouterAI/benchmarks.demo-tabular-benchmarks
📊 Carla HQ Tabular Foundation Model Benchmarks
Centralized benchmark repository of canonical tabular datasets curated for Carla HQ and TabICL (In-Context Learning foundation models for tabular data).
Each dataset is hosted as an independent subset/config with native Parquet storage, schema qualities, OpenML source links, and synchronized Google Sheets for live spreadsheet experimentation.
🚀 Quickstart & Download Options
Option 1: Using… See the full description on the dataset page: https://huggingface.co/datasets/carlahq/demo-tabular-benchmarks.LEM-benchmarks
LEM-benchmarks
Canonical 8-PAC benchmark results for the Lemma model family.
This dataset is an aggregated store of per-round evaluation data produced by
lthn/LEM-Eval. Every row
represents one model's answer to one question in one round of a paired A/B
run against its unmodified base, and the dataset grows monotonically as more
workers contribute — different machines, different sampling states, different
hardware paths — which is the whole point of 8-PAC: multiple independent… See the full description on the dataset page: https://huggingface.co/datasets/lthn/LEM-benchmarks.genomic-benchmarks
Genomic Benchmark
In this repository, we collect benchmarks for classification of genomic sequences. It is shipped as a Python package, together with functions helping to download & manipulate datasets and train NN models.
Citing Genomic Benchmarks
If you use Genomic Benchmarks in your research, please cite it as follows.
Text
GRESOVA, Katarina, et al. Genomic Benchmarks: A Collection of Datasets for Genomic Sequence Classification. bioRxiv, 2022.… See the full description on the dataset page: https://huggingface.co/datasets/katielink/genomic-benchmarks.Benchmarks_CyberSec_RedSageMCQ
Dataset Card for RedSage-MCQ
Dataset Summary
RedSage-MCQ is a large-scale, high-quality multiple-choice question (MCQ) benchmark designed to evaluate the cybersecurity knowledge, skills, and tool proficiency of Large Language Models (LLMs). It is a component of the RedSage-Bench suite introduced in the paper "RedSage: A Cybersecurity Generalist LLM".
The dataset comprises 30,000 questions derived from RedSage-Seed, a curated collection of authoritative… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_RedSageMCQ.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.post-training-benchmarks-viewerbenchmarks
Welcome to 🤗 Diffusers Benchmarks!
This is dataset where we keep track of the inference latency and memory information of the core models in the diffusers library.
Currently, the core models are:
Flux
Wan
LTX
SDXL
Note that we will continue to extend this list based on their usage.
You can analyze the results in this demo.
[!IMPORTANT]
Instead of benchmarking the entire diffusion pipelines, we only benchmark the forward passes
of the diffusion networks under different settings… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/benchmarks.dgx-spark-benchmarks
DGX Spark LLM Arena benchmarks
Reproducible LLM inference benchmarks on an NVIDIA DGX Spark (GB10, 128 GB unified memory). The suite defines eleven tests: six closed-loop (llama-benchy) and five open-loop (vllm bench serve). Results cover all eleven: the ten throughput tests under results, and the rate sweep under rateSweep. Raw results remain inspectable, but only complete runs without a failed sanity check count toward rankings and aggregate throughput. Open-loop tests must… See the full description on the dataset page: https://huggingface.co/datasets/Djangodevreng/dgx-spark-benchmarks.ldr-benchmarks
LDR Community Benchmarks (Leaderboards)
Aggregated leaderboards for Local Deep Research (LDR) community benchmark
runs against SimpleQA, BrowseComp, and xbench-DeepSearch.
👉 Submit results, read raw YAMLs, open PRs:
github.com/LearningCircuit/ldr-benchmarks
This Hugging Face dataset hosts only the aggregated CSV leaderboards.
It is regenerated automatically on every merge to main in the GitHub
repo above. Each CSV row represents one benchmark run (one strategy… See the full description on the dataset page: https://huggingface.co/datasets/local-deep-research/ldr-benchmarks.optical-neuromorphic-eikonal-benchmarks
Optical Neuromorphic Eikonal Solver - Benchmark Datasets
Overview
Benchmark datasets for evaluating the Optical Neuromorphic Eikonal Solver, a GPU-accelerated pathfinding algorithm achieving 30-300× speedup over CPU Dijkstra.
🎯 Key Results
134.9× average speedup vs CPU Dijkstra
0.64% mean error (sub-1% accuracy)
1.025× path length (near-optimal paths)
2-4ms per query on 512×512 grids
📊 Dataset Content
5 synthetic pathfinding test cases covering… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/optical-neuromorphic-eikonal-benchmarks.m5-retail-demand-forecasting-benchmarks
M5 Retail Demand Forecasting & Inventory Risk Benchmarks
This dataset contains the heavily processed artifacts, extracted time-series features, baseline benchmarks, and model artifacts for the M5 Retail Demand Forecasting dataset.
It includes:
Over 1GB of highly engineered temporal, pricing, and calendar features.
Volatility and shortfall risk metrics for 42,840 time series.
XGBoost, Prophet, and SARIMA predictions (point + 95% intervals).
Isolation Forest anomaly detection… See the full description on the dataset page: https://huggingface.co/datasets/snchakri/m5-retail-demand-forecasting-benchmarks.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.benchmarks
py-feat benchmarks
Live benchmark data for py-feat and a
cross-tool comparison against OpenFace 3.0, LibreFace, and PyAFAR.
Powers the py-feat live dashboard. Updated by scheduled
benchmark runs.
Files
File
What
accuracy.csv
Tidy long table: one row per (tool, dataset, modality, metric). Covers AU F1 (DISFA+), 7-class emotion (AffectNet-val, RAF-DB), valence/arousal CCC (AffectNet-val), and gaze angular error (Columbia).
throughput.csv
py-feat… See the full description on the dataset page: https://huggingface.co/datasets/py-feat/benchmarks.maple-preview-cuda-benchmarks
Maple Preview TQ2_0 CUDA Benchmarks
Reproducibility data for the TQ2_0 CUDA patches in
PascalAI2024/maple-preview-windows-cuda.
This repository contains benchmark data, patch files, hashes, and raw validation
evidence. It does not duplicate the Maple model weights.
Result
The fresh local A/B/B/A validation on an RTX 4080 SUPER reproduced the fused-MMQ
prompt-processing gain:
Variant
pp512 mean
pp512 median
tg128 mean
tg128 median
Correctness
MMQ enabled… See the full description on the dataset page: https://huggingface.co/datasets/x0me/maple-preview-cuda-benchmarks.ai-benchmarks-v2-2026
ai-benchmarks-v2-2026
AI data collected daily by Legion API.
🔑 API Access — Updated Daily
Live data via Legion AI API
Free: 100 req/day · Pro €29/month: 50K req/day + full fields
curl "https://api.legion-api.com/incidents?limit=10" -H "X-API-Key: YOUR_PRO_KEY"
Premium archive (1,285 incidents, full analysis): AISI Intelligence Pack €299
📦 Install
pip install legion-intel
from legion_intel import LegionClient
c = LegionClient()… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-benchmarks-v2-2026.llmfit-benchmarks
llmfit Real-World LLM Inference Benchmarks
An evolving dataset of real-world LLM inference measurements across consumer, workstation, datacenter, and unified-memory hardware. It combines benchmarks from an external community source with community benchmarks contributed directly to llmfit.
The initial release contains 1,501 normalized observations:
1,010 unique external-community observations from the repository's 2026-08-10 snapshot.
491 llmfit-community observations from 61… See the full description on the dataset page: https://huggingface.co/datasets/axjns/llmfit-benchmarks.benchmarks-by-vramUpdated on: 21 Sep 2026
Data contains: runs from the last 30 days
Minimum runs: model/hardware combos with fewer than 3 runs are excluded
llm-bench.io — Community LLM Benchmark Leaderboard by Hardware
Per-model community benchmark data for local LLMs, curated from llm-bench.io and grouped by hardware and available VRAM.
This dataset contains only aggregated statistics derived from individual benchmark submissions. It does not contain raw submissions, prompts, model responses… See the full description on the dataset page: https://huggingface.co/datasets/llmbenchio/benchmarks-by-vram.press-release-benchmarks
TechBullion Press Release Builder 📰🚀
TechBullion Press Release Builder helps businesses create professional press releases, technology announcements, startup news, fintech updates, AI stories, and blockchain content ready for publication. Built by GetOnTechBullion.com.
Features
Press Release Quality Score — evaluates structure, clarity, and journalistic standards
Publication Readiness Score — checks formatting and editorial compliance
SEO Optimization Score —… See the full description on the dataset page: https://huggingface.co/datasets/get-on-techbullion/press-release-benchmarks.whisper-browser-benchmarks
whisper-browser-benchmarks
Measurements from a Whisper transcription pipeline running entirely inside a
browser tab: which audio and video containers the browser will actually decode,
how accurate the smallest usable Whisper size is on clean synthetic speech, how
long transcription takes relative to the length of the clip, what the first
load pulls over the wire, and what happens to clips longer than the model's
30-second window.
Everything here was measured, not quoted from a… See the full description on the dataset page: https://huggingface.co/datasets/ruanjiange/whisper-browser-benchmarks.enterprise-llm-inference-benchmarks-2026
🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide)
A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments.
🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation)
Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.soul-benchmarks-locomo
soul.py LoCoMo Benchmark Results
Benchmark results for soul.py on the LoCoMo long-conversation memory benchmark.
Benchmarks repo: github.com/menonpg/soul-benchmarksInteractive results: menonpg.github.io/soul-benchmarks
What is soul.py?
soul.py is an open-source conversational memory layer for LLM agents. It provides multiple retrieval backends (BM25, Qdrant vector search, Relational Learning Model) and an auto-router that selects the best strategy per query.… See the full description on the dataset page: https://huggingface.co/datasets/pgmenon/soul-benchmarks-locomo.sme-valuation-benchmarks-2026
SME Valuation Benchmarks 2026
Reference dataset for small and medium-sized enterprise (SME) valuation: discount rates (WACC), unlevered sector betas and EV/EBITDA multiple ranges for 11 industry sectors across 12 countries (France, Spain, Germany, United Kingdom, United States, Australia, Singapore, India, New Zealand, Ireland, Canada, South Africa). 110 rows.
Columns
Column
Description
sector
Sector key (e.g. software-saas, construction)
sector_label… See the full description on the dataset page: https://huggingface.co/datasets/ValorSME/sme-valuation-benchmarks-2026.octoagent-benchmarks-results
E2E v3 results (Hub splits)
Layout baisbench_celltype (path prefix baisbench_celltype/)
Prepared: datasets.load_dataset(repo, 'baisbench_celltype_prepared', split="task_type_N")
Results: datasets.load_dataset(repo, 'baisbench_celltype', split="task_type_N")
Layout baisbench_codex_gpt55 (path prefix baisbench_codex_gpt55/)
Prepared: datasets.load_dataset(repo, 'baisbench_codex_gpt55_prepared', split="task_type_N")
Results: datasets.load_dataset(repo… See the full description on the dataset page: https://huggingface.co/datasets/ixprzemyslawpietrzak/octoagent-benchmarks-results.app-review-complaint-benchmarks
App Review Complaint Benchmarks — Category Baselines + App Complaint Profiles
Complaint-topic benchmarks for 5,592 apps across 47 app-store categories, computed from real Apple App Store and Google Play review data: complaint rates by topic (crashes, ads, billing, UX…) per app vs. category baseline, plus per-topic worst-offender leaderboards.
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/app-review-complaint-benchmarks.
