datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.benchmarks
Welcome to 🤗 Diffusers Benchmarks!
This is dataset where we keep track of the inference latency and memory information of the core models in the diffusers library.
Currently, the core models are:
Flux
Wan
LTX
SDXL
Note that we will continue to extend this list based on their usage.
You can analyze the results in this demo.
[!IMPORTANT]
Instead of benchmarking the entire diffusion pipelines, we only benchmark the forward passes
of the diffusion networks under different settings… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/benchmarks.dgx-spark-benchmarks
DGX Spark LLM Arena benchmarks
Reproducible LLM inference benchmarks on an NVIDIA DGX Spark (GB10, 128 GB unified memory). The suite defines eleven tests: six closed-loop (llama-benchy) and five open-loop (vllm bench serve). Results cover all eleven: the ten throughput tests under results, and the rate sweep under rateSweep. Raw results remain inspectable, but only complete runs without a failed sanity check count toward rankings and aggregate throughput. Open-loop tests must… See the full description on the dataset page: https://huggingface.co/datasets/Djangodevreng/dgx-spark-benchmarks.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.optical-neuromorphic-eikonal-benchmarks
Optical Neuromorphic Eikonal Solver - Benchmark Datasets
Overview
Benchmark datasets for evaluating the Optical Neuromorphic Eikonal Solver, a GPU-accelerated pathfinding algorithm achieving 30-300× speedup over CPU Dijkstra.
🎯 Key Results
134.9× average speedup vs CPU Dijkstra
0.64% mean error (sub-1% accuracy)
1.025× path length (near-optimal paths)
2-4ms per query on 512×512 grids
📊 Dataset Content
5 synthetic pathfinding test cases covering… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/optical-neuromorphic-eikonal-benchmarks.benchmarks
py-feat benchmarks
Live benchmark data for py-feat and a
cross-tool comparison against OpenFace 3.0, LibreFace, and PyAFAR.
Powers the py-feat live dashboard. Updated by scheduled
benchmark runs.
Files
File
What
accuracy.csv
Tidy long table: one row per (tool, dataset, modality, metric). Covers AU F1 (DISFA+), 7-class emotion (AffectNet-val, RAF-DB), valence/arousal CCC (AffectNet-val), and gaze angular error (Columbia).
throughput.csv
py-feat… See the full description on the dataset page: https://huggingface.co/datasets/py-feat/benchmarks.benchmarks-by-vramUpdated on: 21 Sep 2026
Data contains: runs from the last 30 days
Minimum runs: model/hardware combos with fewer than 3 runs are excluded
llm-bench.io — Community LLM Benchmark Leaderboard by Hardware
Per-model community benchmark data for local LLMs, curated from llm-bench.io and grouped by hardware and available VRAM.
This dataset contains only aggregated statistics derived from individual benchmark submissions. It does not contain raw submissions, prompts, model responses… See the full description on the dataset page: https://huggingface.co/datasets/llmbenchio/benchmarks-by-vram.whisper-browser-benchmarks
whisper-browser-benchmarks
Measurements from a Whisper transcription pipeline running entirely inside a
browser tab: which audio and video containers the browser will actually decode,
how accurate the smallest usable Whisper size is on clean synthetic speech, how
long transcription takes relative to the length of the clip, what the first
load pulls over the wire, and what happens to clips longer than the model's
30-second window.
Everything here was measured, not quoted from a… See the full description on the dataset page: https://huggingface.co/datasets/ruanjiange/whisper-browser-benchmarks.press-release-benchmarks
TechBullion Press Release Builder 📰🚀
TechBullion Press Release Builder helps businesses create professional press releases, technology announcements, startup news, fintech updates, AI stories, and blockchain content ready for publication. Built by GetOnTechBullion.com.
Features
Press Release Quality Score — evaluates structure, clarity, and journalistic standards
Publication Readiness Score — checks formatting and editorial compliance
SEO Optimization Score —… See the full description on the dataset page: https://huggingface.co/datasets/get-on-techbullion/press-release-benchmarks.ninfer-benchmarks
NInfer on one RTX 5090 — recorded benchmark evidence
Historical results from 19–20 September 2026, published by dima0000. This is a collection of benchmark evidence, not a model checkpoint or a live inference service. It preserves successful measurements, failed checks and incomplete experiments.
Hardware: one NVIDIA RTX 5090 with 32 GB VRAM per trial. Models: regular and uncensored Qwen3.8-27B with NVFP4 weights. Engine: NInfer commit 9e163eee4b8acec21ab0ac765107b6a3f287b217… See the full description on the dataset page: https://huggingface.co/datasets/dima0000/ninfer-benchmarks.sme-valuation-benchmarks-2026
SME Valuation Benchmarks 2026
Reference dataset for small and medium-sized enterprise (SME) valuation: discount rates (WACC), unlevered sector betas and EV/EBITDA multiple ranges for 11 industry sectors across 12 countries (France, Spain, Germany, United Kingdom, United States, Australia, Singapore, India, New Zealand, Ireland, Canada, South Africa). 110 rows.
Columns
Column
Description
sector
Sector key (e.g. software-saas, construction)
sector_label… See the full description on the dataset page: https://huggingface.co/datasets/ValorSME/sme-valuation-benchmarks-2026.quantization-benchmarksgenomic-benchmarks
genomic-benchmarks, curated
genomic-benchmarks, as published in quality-curated genomic benchmarks - one format, fixed row order, a permanent ID on every row. 7 datasets, 14 files, 600,434 rows, one gzipped CSV per split.
Getting the data
Two packages are the way in: genomic-benchmarks-data for people, genomic-benchmarks-data4agents for agents, the same functions either way. They resolve the URL, check the checksum, and carry each dataset's QC results, which this… See the full description on the dataset page: https://huggingface.co/datasets/genomic-benchmarks/genomic-benchmarks.apple-silicon-llm-benchmarks
Apple Silicon Local LLM Benchmarks — M2 Max 32GB
Measurements taken while trying to get Qwen3.8-27B usable locally on a 32GB M2 Max. Most of
the popular speedup advice did not transfer from CUDA, so these are mostly negative results.
Everything here was measured on one machine. Treat it as a datapoint, not a law.
Hardware and software
Chip
Apple M2 Max
Unified memory
32 GB (~21.8 GB wireable to the GPU)
macOS
26.5.2
llama.cpp
build c1d0e7a00… See the full description on the dataset page: https://huggingface.co/datasets/RaynarDM/apple-silicon-llm-benchmarks.factanker-us-industry-benchmarks
US Industry Benchmarks — SEC Filing Data, evidence-linked
Percentile distributions (p10/p25/median/p75/p90, mean, min, max) of key
financial metrics — revenue, net income, operating income, total assets,
EBITDA margin — across US-listed companies, grouped by SIC industry code
and fiscal year, computed from SEC EDGAR filings as filed.
Every row carries its evidence. The cite_as column contains the
citation string, page_url the stable public page, and
evidence_fact_urls links to… See the full description on the dataset page: https://huggingface.co/datasets/factanker/factanker-us-industry-benchmarks.liquidity-intelligence-benchmarks
VOIDTRACE AI Liquidity Intelligence Benchmarks
Benchmark dataset of 20 crypto liquidity intelligence cases with individual scores for liquidity flow, stablecoin intelligence, capital rotation, DEX activity, bridge activity, and ecosystem momentum across 8 blockchain networks.
Built by VOIDTRACE AI.
Dataset Description
This dataset contains benchmark data for the VOIDTRACE AI Crypto Liquidity Intelligence Engine — a blockchain intelligence software concept… See the full description on the dataset page: https://huggingface.co/datasets/voidtrace-ai/liquidity-intelligence-benchmarks.distribution-benchmarks
PressRelease Distribution Engine Benchmarks
Benchmark dataset of 20 press release distribution cases with individual scores for release quality, media match, distribution reach, publication rate, syndication, and AI visibility across 6 distribution channels.
Built by PressRelease.fyi.
Dataset Description
This dataset contains benchmark data for the PressRelease Distribution Engine — a software toolkit for preparing, organizing, and managing press release… See the full description on the dataset page: https://huggingface.co/datasets/pressrelease-fyi/distribution-benchmarks.apple-m5-pro-24gb-llm-tps-benchmarks
Apple M5 Pro (24GB) — Local LLM tokens/sec benchmarks
Real-world tokens/sec numbers for 12 local coding/agentic LLMs, measured on a specific,
commonly-owned but previously unbenchmarked configuration: Apple M5 Pro, 24GB unified
memory. At the time of writing, no public benchmark existed for this exact chip + RAM
combination for these models. Includes both the official newly-open-sourced Qwen3.8-27B and
a popular uncensored/abliterated finetune of it, for comparison.… See the full description on the dataset page: https://huggingface.co/datasets/abhisheksharma0994/apple-m5-pro-24gb-llm-tps-benchmarks.deindexing-automation-benchmarks
Deindexing Automation Engine Benchmarks
Benchmark dataset of 20 deindexing automation cases with individual scores for deindex strength, removal rate, review issue handling, reputation health, platform coverage, and workflow efficiency across major removal types and industries.
Built by Deindexing.Services.
Dataset Description
This dataset contains benchmark data for the Deindexing Automation Engine — an automation engine for managing search deindexing, content… See the full description on the dataset page: https://huggingface.co/datasets/deindexing-services/deindexing-automation-benchmarks.miRBench
miRBench, curated
miRBench, as published in quality-curated genomic benchmarks - one format, fixed row order, a permanent ID on every row. 3 datasets, 6 files, 2,849,872 rows, one gzipped CSV per split, all Homo sapiens.
Getting the data
Two packages are the way in: genomic-benchmarks-data for people, genomic-benchmarks-data4agents for agents, the same functions either way. They resolve the URL, check the checksum, and carry each dataset's QC results, which this… See the full description on the dataset page: https://huggingface.co/datasets/genomic-benchmarks/miRBench.us-50state-macro-benchmarks-2026
US 50-State Macro, Real Estate & Energy Benchmarks (2026)
Overview
This dataset provides an empirical, state-by-state economic and infrastructure benchmark across all 50 United States and the District of Columbia for 2026.
Compiled and maintained by the Groundwork Research Desk, it serves as the foundational data layer for open interactive utilities on gworky.com/tools, including mortgage break-even analysis, rooftop solar payback calculations, and regional… See the full description on the dataset page: https://huggingface.co/datasets/elenagroundwork/us-50state-macro-benchmarks-2026.factanker-us-bank-benchmarks
US Bank Benchmarks — FFIEC Call Report Data, evidence-linked
Quarterly percentile distributions (p10/p25/median/p75/p90, mean, min,
max) of the core banking metrics — total assets, deposits, loans, net
income, equity ratio, loan-to-deposit, loan-loss reserve, ROA, ROE, net
interest margin, efficiency ratio — across every US bank that files a
quarterly FFIEC Call Report, by asset-size peer group and by state,
2001Q1 to today. 58 peer groups × 11 metrics × 100+ quarters.
Every row… See the full description on the dataset page: https://huggingface.co/datasets/factanker/factanker-us-bank-benchmarks.visibility-benchmarks
StreetInsider Visibility Engine 📡🗞️
StreetInsider Visibility Engine is a lightweight content visibility and publication workflow tool designed to help businesses, brands, and publishers organize, optimize, and monitor their media content for greater online discoverability. Built by GetOnStreetInsider.com.
Key Functions
Content Preparation — Press release and content preparation for publication readiness
Metadata Validation — Publication metadata validation… See the full description on the dataset page: https://huggingface.co/datasets/streetinsider/visibility-benchmarks.website-project-cost-benchmarks
Website & App Project Cost Benchmarks 2026
Website & app project cost benchmarks - calibrated on 600+ project quotes and public rate benchmarks. CC-BY 4.0. Source and methodology: https://projectcostestimator.com
This is the dataset behind Project Cost Estimator, an independent website cost estimator. The canonical machine-readable source is the live endpoint https://projectcostestimator.com/api/cost-data (no auth, CORS open). The files here are a published snapshot of that… See the full description on the dataset page: https://huggingface.co/datasets/zimzum1984/website-project-cost-benchmarks.compile-benchmarksexperts-backendsglobal-vps-benchmarks-2026
🚀 2026 Global Cloud VPS Performance, Pricing & Latency Benchmark Dataset
A curated benchmark dataset covering global cloud VPS (Virtual Private Server) providers, verified hardware specs, standardized Geekbench 6 synthetic performance, monthly pricing, and network latency back to East Asia (China routing).
Maintained by the independent benchmarking project at VPS Rankings.
📊 Dataset Summary
This dataset compiles standardized hardware and pricing telemetry… See the full description on the dataset page: https://huggingface.co/datasets/VPSRankings/global-vps-benchmarks-2026.webgpu-compute-benchmarks
WebGPU compute benchmarks across browsers and GPU vendors
Every run submitted to gpubench.dev between 2026-03-30
and 2026-08-14, exported from the live table. WebGPU compute-shader throughput
measured in real browsers on whatever hardware visitors happened to have.
This exists because two preprints cite cross-vendor results from this table and
the table is live and mutable, so those claims could not be checked by a reader.
This snapshot is the checkable version.… See the full description on the dataset page: https://huggingface.co/datasets/abgunaydin/webgpu-compute-benchmarks.market-insight-benchmarks
Financial News Market Insight Bridge Benchmarks
Benchmark dataset of 20 financial news article cases with individual scores for AI visibility, content discovery, topic matching, search visibility, article organization, and finance topic mapping.
Built by FinancialNews.it.com.
Dataset Description
This dataset contains benchmark data for a content visibility and AI discovery bot helping Financial News articles gain greater discoverability across AI platforms… See the full description on the dataset page: https://huggingface.co/datasets/financialnews/market-insight-benchmarks.
