CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01leoschneider /daytrader-benchmarkstabular10M<n<100M0 likes7.6k downloads5mo agoHugging Face02madesai /what-ai-benchmarks-actually-measure What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models Item-level model responses and scores for 53 language models across the 56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks (Desai et al., 2026, arxiv.org/abs/2609.08812). We do not release the prompts from the benchmark datasets, but instead refer to them by item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.tabular1M<n<10M0 likes1.3k downloads12d agoHugging Face03OpenChainBench /benchmarks OpenChainBench Crypto Infrastructure Benchmarks Daily snapshots of every public benchmark on openchainbench.com, released as Hive-partitioned Parquet under CC-BY-4.0. OCB measures latency, cost, coverage and accuracy of crypto infrastructure (RPCs, oracles, bridges, data APIs, Polymarket adapters, Hyperliquid builders). Every snapshot here mirrors the /api/citable, /api/stat/<slug>, and /api/series/<slug> JSON feeds at the time of capture. Latest snapshot: 2026-09-22 (captured… See the full description on the dataset page: https://huggingface.co/datasets/OpenChainBench/benchmarks.tabulartime-series-forecasting1M<n<10M1 likes1.2k downloads2h agoHugging Face04witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.2k downloads2d agoHugging Face05JacobPEvans /mlx-benchmarks MLX Benchmarks Structured benchmark results for MLX-quantized and other locally-hosted LLMs on Apple Silicon. Covers throughput, time-to-first-token, tool-calling, code generation, reasoning, knowledge, and math suites. Results are produced by a sweep harness that wires upstream evaluation tools against a local vllm-mlx inference server: EleutherAI/lm-evaluation-harness — coding, reasoning, knowledge, math linusvwe/MLXBench — throughput and time-to-first-token vllm… See the full description on the dataset page: https://huggingface.co/datasets/JacobPEvans/mlx-benchmarks.tabular1K<n<10K1 likes532 downloads12d agoHugging Face06BitRouterAI /benchmarks BitRouter Benchmarks This dataset release contains the private-router Terminal-Bench 2.1 routing study and its audit evidence. July 2026 experiments are preserved under v0/; the current study, raw Harbor traces, run-scoped BitRouter SQLite files, sanitized production-usage backfill, reproducible joins, and Parquet tables are under terminal-bench-2.1/router-random-study/. Main result All valid-case metrics below use the same strict intersection of 80 tasks. “Harbor… See the full description on the dataset page: https://huggingface.co/datasets/BitRouterAI/benchmarks.tabular100K<n<1M1 likes500 downloads13d agoHugging Face07carlahq /demo-tabular-benchmarks 📊 Carla HQ Tabular Foundation Model Benchmarks Centralized benchmark repository of canonical tabular datasets curated for Carla HQ and TabICL (In-Context Learning foundation models for tabular data). Each dataset is hosted as an independent subset/config with native Parquet storage, schema qualities, OpenML source links, and synchronized Google Sheets for live spreadsheet experimentation. 🚀 Quickstart & Download Options Option 1: Using… See the full description on the dataset page: https://huggingface.co/datasets/carlahq/demo-tabular-benchmarks.tabular100K<n<1M0 likes400 downloads26d agoHugging Face08lthn /LEM-benchmarks LEM-benchmarks Canonical 8-PAC benchmark results for the Lemma model family. This dataset is an aggregated store of per-round evaluation data produced by lthn/LEM-Eval. Every row represents one model's answer to one question in one round of a paired A/B run against its unmodified base, and the dataset grows monotonically as more workers contribute — different machines, different sampling states, different hardware paths — which is the whole point of 8-PAC: multiple independent… See the full description on the dataset page: https://huggingface.co/datasets/lthn/LEM-benchmarks.tabularquestion-answering10K<n<100K3 likes345 downloads5mo agoHugging Face09katielink /genomic-benchmarks Genomic Benchmark In this repository, we collect benchmarks for classification of genomic sequences. It is shipped as a Python package, together with functions helping to download & manipulate datasets and train NN models. Citing Genomic Benchmarks If you use Genomic Benchmarks in your research, please cite it as follows. Text GRESOVA, Katarina, et al. Genomic Benchmarks: A Collection of Datasets for Genomic Sequence Classification. bioRxiv, 2022.… See the full description on the dataset page: https://huggingface.co/datasets/katielink/genomic-benchmarks.tabular100K<n<1M7 likes317 downloads3y agoHugging Face10RISys-Lab /Benchmarks_CyberSec_RedSageMCQ Dataset Card for RedSage-MCQ Dataset Summary RedSage-MCQ is a large-scale, high-quality multiple-choice question (MCQ) benchmark designed to evaluate the cybersecurity knowledge, skills, and tool proficiency of Large Language Models (LLMs). It is a component of the RedSage-Bench suite introduced in the paper "RedSage: A Cybersecurity Generalist LLM". The dataset comprises 30,000 questions derived from RedSage-Seed, a curated collection of authoritative… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_RedSageMCQ.tabularquestion-answering10K<n<100K0 likes256 downloads8mo agoHugging Face11omegaprime669 /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.imagetext-generationn<1K0 likes255 downloads2mo agoHugging Face12HuggingFaceTB /post-training-benchmarks-viewertabularn<1K3 likes188 downloads11mo agoHugging Face13diffusers /benchmarks Welcome to 🤗 Diffusers Benchmarks! This is dataset where we keep track of the inference latency and memory information of the core models in the diffusers library. Currently, the core models are: Flux Wan LTX SDXL Note that we will continue to extend this list based on their usage. You can analyze the results in this demo. [!IMPORTANT] Instead of benchmarking the entire diffusion pipelines, we only benchmark the forward passes of the diffusion networks under different settings… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/benchmarks.tabularn<1K15 likes174 downloads7d agoHugging Face14Djangodevreng /dgx-spark-benchmarks DGX Spark LLM Arena benchmarks Reproducible LLM inference benchmarks on an NVIDIA DGX Spark (GB10, 128 GB unified memory). The suite defines eleven tests: six closed-loop (llama-benchy) and five open-loop (vllm bench serve). Results cover all eleven: the ten throughput tests under results, and the rate sweep under rateSweep. Raw results remain inspectable, but only complete runs without a failed sanity check count toward rankings and aggregate throughput. Open-loop tests must… See the full description on the dataset page: https://huggingface.co/datasets/Djangodevreng/dgx-spark-benchmarks.tabularn<1K1 likes162 downloads4d agoHugging Face15local-deep-research /ldr-benchmarks LDR Community Benchmarks (Leaderboards) Aggregated leaderboards for Local Deep Research (LDR) community benchmark runs against SimpleQA, BrowseComp, and xbench-DeepSearch. 👉 Submit results, read raw YAMLs, open PRs: github.com/LearningCircuit/ldr-benchmarks This Hugging Face dataset hosts only the aggregated CSV leaderboards. It is regenerated automatically on every merge to main in the GitHub repo above. Each CSV row represents one benchmark run (one strategy… See the full description on the dataset page: https://huggingface.co/datasets/local-deep-research/ldr-benchmarks.tabularquestion-answeringn<1K13 likes155 downloads4mo agoHugging Face16Agnuxo /optical-neuromorphic-eikonal-benchmarks Optical Neuromorphic Eikonal Solver - Benchmark Datasets Overview Benchmark datasets for evaluating the Optical Neuromorphic Eikonal Solver, a GPU-accelerated pathfinding algorithm achieving 30-300× speedup over CPU Dijkstra. 🎯 Key Results 134.9× average speedup vs CPU Dijkstra 0.64% mean error (sub-1% accuracy) 1.025× path length (near-optimal paths) 2-4ms per query on 512×512 grids 📊 Dataset Content 5 synthetic pathfinding test cases covering… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/optical-neuromorphic-eikonal-benchmarks.tabularothern<1K0 likes138 downloads11mo agoHugging Face17snchakri /m5-retail-demand-forecasting-benchmarks M5 Retail Demand Forecasting & Inventory Risk Benchmarks This dataset contains the heavily processed artifacts, extracted time-series features, baseline benchmarks, and model artifacts for the M5 Retail Demand Forecasting dataset. It includes: Over 1GB of highly engineered temporal, pricing, and calendar features. Volatility and shortfall risk metrics for 42,840 time series. XGBoost, Prophet, and SARIMA predictions (point + 95% intervals). Isolation Forest anomaly detection… See the full description on the dataset page: https://huggingface.co/datasets/snchakri/m5-retail-demand-forecasting-benchmarks.tabulartime-series-forecasting100K<n<1M0 likes131 downloads1mo agoHugging Face18humanlong /emotion-negotiation-benchmarks Emotion-Aware LLM Negotiation Benchmarks Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency. The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.tabulartext-generationn<1K0 likes130 downloads4mo agoHugging Face19py-feat /benchmarks py-feat benchmarks Live benchmark data for py-feat and a cross-tool comparison against OpenFace 3.0, LibreFace, and PyAFAR. Powers the py-feat live dashboard. Updated by scheduled benchmark runs. Files File What accuracy.csv Tidy long table: one row per (tool, dataset, modality, metric). Covers AU F1 (DISFA+), 7-class emotion (AffectNet-val, RAF-DB), valence/arousal CCC (AffectNet-val), and gaze angular error (Columbia). throughput.csv py-feat… See the full description on the dataset page: https://huggingface.co/datasets/py-feat/benchmarks.tabularn<1K0 likes124 downloads3mo agoHugging Face20x0me /maple-preview-cuda-benchmarks Maple Preview TQ2_0 CUDA Benchmarks Reproducibility data for the TQ2_0 CUDA patches in PascalAI2024/maple-preview-windows-cuda. This repository contains benchmark data, patch files, hashes, and raw validation evidence. It does not duplicate the Maple model weights. Result The fresh local A/B/B/A validation on an RTX 4080 SUPER reproduced the fused-MMQ prompt-processing gain: Variant pp512 mean pp512 median tg128 mean tg128 median Correctness MMQ enabled… See the full description on the dataset page: https://huggingface.co/datasets/x0me/maple-preview-cuda-benchmarks.tabularn<1K0 likes108 downloads1mo agoHugging Face21gemmozero /ai-benchmarks-v2-2026 ai-benchmarks-v2-2026 AI data collected daily by Legion API. 🔑 API Access — Updated Daily Live data via Legion AI API Free: 100 req/day · Pro €29/month: 50K req/day + full fields curl "https://api.legion-api.com/incidents?limit=10" -H "X-API-Key: YOUR_PRO_KEY" Premium archive (1,285 incidents, full analysis): AISI Intelligence Pack €299 📦 Install pip install legion-intel from legion_intel import LegionClient c = LegionClient()… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-benchmarks-v2-2026.tabularn<1K0 likes100 downloads11h agoHugging Face22axjns /llmfit-benchmarks llmfit Real-World LLM Inference Benchmarks An evolving dataset of real-world LLM inference measurements across consumer, workstation, datacenter, and unified-memory hardware. It combines benchmarks from an external community source with community benchmarks contributed directly to llmfit. The initial release contains 1,501 normalized observations: 1,010 unique external-community observations from the repository's 2026-08-10 snapshot. 491 llmfit-community observations from 61… See the full description on the dataset page: https://huggingface.co/datasets/axjns/llmfit-benchmarks.tabulartext-generation1K<n<10K2 likes97 downloads1mo agoHugging Face23llmbenchio /benchmarks-by-vramUpdated on: 21 Sep 2026 Data contains: runs from the last 30 days Minimum runs: model/hardware combos with fewer than 3 runs are excluded llm-bench.io — Community LLM Benchmark Leaderboard by Hardware Per-model community benchmark data for local LLMs, curated from llm-bench.io and grouped by hardware and available VRAM. This dataset contains only aggregated statistics derived from individual benchmark submissions. It does not contain raw submissions, prompts, model responses… See the full description on the dataset page: https://huggingface.co/datasets/llmbenchio/benchmarks-by-vram.tabularn<1K0 likes91 downloads9h agoHugging Face24get-on-techbullion /press-release-benchmarks TechBullion Press Release Builder 📰🚀 TechBullion Press Release Builder helps businesses create professional press releases, technology announcements, startup news, fintech updates, AI stories, and blockchain content ready for publication. Built by GetOnTechBullion.com. Features Press Release Quality Score — evaluates structure, clarity, and journalistic standards Publication Readiness Score — checks formatting and editorial compliance SEO Optimization Score —… See the full description on the dataset page: https://huggingface.co/datasets/get-on-techbullion/press-release-benchmarks.tabularn<1K0 likes84 downloads2mo agoHugging Face25ruanjiange /whisper-browser-benchmarks whisper-browser-benchmarks Measurements from a Whisper transcription pipeline running entirely inside a browser tab: which audio and video containers the browser will actually decode, how accurate the smallest usable Whisper size is on clean synthetic speech, how long transcription takes relative to the length of the clip, what the first load pulls over the wire, and what happens to clips longer than the model's 30-second window. Everything here was measured, not quoted from a… See the full description on the dataset page: https://huggingface.co/datasets/ruanjiange/whisper-browser-benchmarks.audion<1K0 likes82 downloads5d agoHugging Face26Abdulrahmankalil /enterprise-llm-inference-benchmarks-2026 🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide) A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments. 🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation) Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.tabulartext-generationn<1K1 likes77 downloads4d agoHugging Face27pgmenon /soul-benchmarks-locomo soul.py LoCoMo Benchmark Results Benchmark results for soul.py on the LoCoMo long-conversation memory benchmark. Benchmarks repo: github.com/menonpg/soul-benchmarksInteractive results: menonpg.github.io/soul-benchmarks What is soul.py? soul.py is an open-source conversational memory layer for LLM agents. It provides multiple retrieval backends (BM25, Qdrant vector search, Relational Learning Model) and an auto-router that selects the best strategy per query.… See the full description on the dataset page: https://huggingface.co/datasets/pgmenon/soul-benchmarks-locomo.tabularquestion-answering1K<n<10K0 likes73 downloads4mo agoHugging Face28ValorSME /sme-valuation-benchmarks-2026 SME Valuation Benchmarks 2026 Reference dataset for small and medium-sized enterprise (SME) valuation: discount rates (WACC), unlevered sector betas and EV/EBITDA multiple ranges for 11 industry sectors across 12 countries (France, Spain, Germany, United Kingdom, United States, Australia, Singapore, India, New Zealand, Ireland, Canada, South Africa). 110 rows. Columns Column Description sector Sector key (e.g. software-saas, construction) sector_label… See the full description on the dataset page: https://huggingface.co/datasets/ValorSME/sme-valuation-benchmarks-2026.tabularn<1K0 likes67 downloads10d agoHugging Face29ixprzemyslawpietrzak /octoagent-benchmarks-results E2E v3 results (Hub splits) Layout baisbench_celltype (path prefix baisbench_celltype/) Prepared: datasets.load_dataset(repo, 'baisbench_celltype_prepared', split="task_type_N") Results: datasets.load_dataset(repo, 'baisbench_celltype', split="task_type_N") Layout baisbench_codex_gpt55 (path prefix baisbench_codex_gpt55/) Prepared: datasets.load_dataset(repo, 'baisbench_codex_gpt55_prepared', split="task_type_N") Results: datasets.load_dataset(repo… See the full description on the dataset page: https://huggingface.co/datasets/ixprzemyslawpietrzak/octoagent-benchmarks-results.tabular1K<n<10K0 likes64 downloads5mo agoHugging Face30zalizedata /app-review-complaint-benchmarks App Review Complaint Benchmarks — Category Baselines + App Complaint Profiles Complaint-topic benchmarks for 5,592 apps across 47 app-store categories, computed from real Apple App Store and Google Play review data: complaint rates by topic (crashes, ads, billing, UX…) per app vs. category baseline, plus per-topic worst-offender leaderboards. Part of the DataForge Open Data program — full production packages, free for academic and personal use. Canonical dataset page:… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/app-review-complaint-benchmarks.tabulartext-classification1K<n<10K0 likes63 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.