datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Video-MME-v2
🔥 News
2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch.
2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups.
🤗 About This Repo
This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.daytrader-benchmarksAwesome_Spatial_VQA_Benchmarkswhat-ai-benchmarks-actually-measure
What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models
Item-level model responses and scores for 53 language models across the
56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting
Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
(Desai et al., 2026,
arxiv.org/abs/2609.08812).
We do not release the prompts from the benchmark datasets, but instead refer to them by
item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.benchmarks
OpenChainBench Crypto Infrastructure Benchmarks
Daily snapshots of every public benchmark on
openchainbench.com, released as
Hive-partitioned Parquet under CC-BY-4.0.
OCB measures latency, cost, coverage and accuracy of crypto
infrastructure (RPCs, oracles, bridges, data APIs, Polymarket adapters,
Hyperliquid builders). Every snapshot here mirrors the
/api/citable,
/api/stat/<slug>,
and /api/series/<slug>
JSON feeds at the time of capture.
Latest snapshot: 2026-09-22 (captured… See the full description on the dataset page: https://huggingface.co/datasets/OpenChainBench/benchmarks.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.Awesome_Spatial_VQA_Benchmarks_ViewSpatial-Benchjava_evaluation_benchmarksramanv-image-vqa-benchmarksmlx-benchmarks
MLX Benchmarks
Structured benchmark results for MLX-quantized and other locally-hosted
LLMs on Apple Silicon. Covers throughput, time-to-first-token, tool-calling,
code generation, reasoning, knowledge, and math suites.
Results are produced by a sweep harness that wires upstream evaluation tools
against a local vllm-mlx inference server:
EleutherAI/lm-evaluation-harness — coding, reasoning, knowledge, math
linusvwe/MLXBench — throughput and time-to-first-token
vllm… See the full description on the dataset page: https://huggingface.co/datasets/JacobPEvans/mlx-benchmarks.benchmarks
BitRouter Benchmarks
This dataset release contains the private-router Terminal-Bench 2.1 routing study and its audit evidence. July 2026 experiments are preserved under v0/; the current study, raw Harbor traces, run-scoped BitRouter SQLite files, sanitized production-usage backfill, reproducible joins, and Parquet tables are under terminal-bench-2.1/router-random-study/.
Main result
All valid-case metrics below use the same strict intersection of 80 tasks. “Harbor… See the full description on the dataset page: https://huggingface.co/datasets/BitRouterAI/benchmarks.Health_Benchmarks
LLM Health Benchmarks Dataset by Yesil Science
The LLM Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties. It provides structured question-answer pairs designed to test the performance of AI models in understanding and generating domain-specific knowledge.
Primary Purpose
This dataset is built to:
Benchmark LLMs in medical specialties and subfields.
Assess the accuracy and contextual… See the full description on the dataset page: https://huggingface.co/datasets/yesilhealth/Health_Benchmarks.data-agent-benchmarks
LongHorizon Full Data-Agent Benchmarks
Companion data artifacts for five complete evaluation tracks:
DataSciBench full55 / 167 metric entries
DABStep full450
DABStep-Research full100
DSBench Modeling full74
LongDS full68 / 2,225 turns
The companion GitHub repository contains processed manifests, evaluation code,
historical API ReAct baseline code, download/preparation tools, and the frozen
source lock. artifact_manifest.json records every uploaded object's size,
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.UI-Grounding-Benchmarks
UI-Grounding-Benchmarks
This is a collection of UI grounding benchmarks:
ScreenSpot
ScreenSpot-V2
ScreenSpot-Pro
OS-World-G
UI-Vision
Thanks for their great work!
This benchmark collection is used in the paper:
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
🖼️ Project Page: https://showlab.github.io/FocusUI/
🏠 Github Repo: https://github.com/showlab/FocusUI
📝 Paper: https://arxiv.org/pdf/2601.03928
Model Zoo
Model
Backbone
🤗… See the full description on the dataset page: https://huggingface.co/datasets/yyyang/UI-Grounding-Benchmarks.ttpd-benchmarks
TTP-D benchmarks, datasets and results
Problem instances, behaviour-cloning datasets, and solver result tables for the
study Fly, Pack, Drive: the Travelling Thief Problem with Drone.
Paper: Drive, Pack, Fly: The Travelling Thief Problem with Drone
A capacitated truck and a single-package drone operate from a common depot on a
collection route. The truck's velocity decreases affinely with its accumulated
load, so an early pickup penalises every subsequent arc. The drone launches… See the full description on the dataset page: https://huggingface.co/datasets/Murjani/ttpd-benchmarks.demo-tabular-benchmarks
📊 Carla HQ Tabular Foundation Model Benchmarks
Centralized benchmark repository of canonical tabular datasets curated for Carla HQ and TabICL (In-Context Learning foundation models for tabular data).
Each dataset is hosted as an independent subset/config with native Parquet storage, schema qualities, OpenML source links, and synchronized Google Sheets for live spreadsheet experimentation.
🚀 Quickstart & Download Options
Option 1: Using… See the full description on the dataset page: https://huggingface.co/datasets/carlahq/demo-tabular-benchmarks.SkillOpt_Lite_Benchmarks
SkillOpt_Lite Benchmarks
Train / val / test splits used by the SkillOpt_Lite project.
One multi-config repo containing all six benchmarks:
Config
Rows (train / val / test)
Content shipped
searchqa
400 / 200 / 1400
Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA.
docvqa
107 / 53 / 374
Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.Genomic_Benchmarks_human_nontata_promoters
Dataset Card for "Genomic_Benchmarks_human_nontata_promoters"
More Information needed
benchmark-scifactOCSR-Benchmarks
OCSR Benchmarks
A collection of ten benchmark datasets for Optical Chemical Structure Recognition (OCSR) —
the task of converting chemical structure diagram images into machine-readable SMILES strings.
These benchmarks were used to evaluate the COMO model
(Closed-Loop Optical Molecule Recognition).
Subsets
Config
Split
Size
Domain
CLEF
test
992
Real
JPO
test
449
Real
UOB
test
5,740
Real
USPTO
test
5719
Real
USPTO-10K
test
9,999
Real
Staker
test
50,000… See the full description on the dataset page: https://huggingface.co/datasets/Keylab/OCSR-Benchmarks.LEM-benchmarks
LEM-benchmarks
Canonical 8-PAC benchmark results for the Lemma model family.
This dataset is an aggregated store of per-round evaluation data produced by
lthn/LEM-Eval. Every row
represents one model's answer to one question in one round of a paired A/B
run against its unmodified base, and the dataset grows monotonically as more
workers contribute — different machines, different sampling states, different
hardware paths — which is the whole point of 8-PAC: multiple independent… See the full description on the dataset page: https://huggingface.co/datasets/lthn/LEM-benchmarks.Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs
Dataset Card for "Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs"
More Information needed
Genomic_Benchmarks_human_enhancers_cohn
Dataset Card for "Genomic_Benchmarks_human_enhancers_cohn"
More Information needed
genomic-benchmarks
Genomic Benchmark
In this repository, we collect benchmarks for classification of genomic sequences. It is shipped as a Python package, together with functions helping to download & manipulate datasets and train NN models.
Citing Genomic Benchmarks
If you use Genomic Benchmarks in your research, please cite it as follows.
Text
GRESOVA, Katarina, et al. Genomic Benchmarks: A Collection of Datasets for Genomic Sequence Classification. bioRxiv, 2022.… See the full description on the dataset page: https://huggingface.co/datasets/katielink/genomic-benchmarks.Genomic_Benchmarks_human_ocr_ensembl
Dataset Card for "Genomic_Benchmarks_human_ocr_ensembl"
More Information needed
screen-benchmarksGenomic_Benchmarks_human_ensembl_regulatory
Dataset Card for "Genomic_Benchmarks_human_ensembl_regulatory"
More Information needed
dna_benchmarks
DNA Benchmarks
Dataset Description
DNA Benchmarks is a collection of genomic datasets organized for benchmarking DNA foundation models, genomic representation learning methods, and multimodal genomic learning frameworks.
The collection covers a wide range of sequence scales, from short regulatory sequences of a few hundred base pairs to genome-scale tiling datasets derived from whole genomes.
This repository is designed as a general-purpose benchmark hub rather… See the full description on the dataset page: https://huggingface.co/datasets/hxxiang/dna_benchmarks.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.Benchmarks_CyberSec_RedSageMCQ
Dataset Card for RedSage-MCQ
Dataset Summary
RedSage-MCQ is a large-scale, high-quality multiple-choice question (MCQ) benchmark designed to evaluate the cybersecurity knowledge, skills, and tool proficiency of Large Language Models (LLMs). It is a component of the RedSage-Bench suite introduced in the paper "RedSage: A Cybersecurity Generalist LLM".
The dataset comprises 30,000 questions derived from RedSage-Seed, a curated collection of authoritative… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_RedSageMCQ.
