datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformersVideo-MME-v2
🔥 News
2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch.
2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups.
🤗 About This Repo
This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.daytrader-benchmarksspeculator_benchmarksThis dataset contains dataset splits for evaluating speculative decoding algorithms on different tasks.
File:
Coding: HumanEval.jsonl
Math: math_reasoning.jsonl
Question Answering: qa.jsonl
MT_bench: question.jsonl
Retrieval-Augmented Generation: rag.jsonl
Summarization: summarization.jsonl
Translation (German to English): translation.jsonl
Writing: writing.jsonl
The data comes from two sources:
https://github.com/openai/human-eval (1). (The MIT License)… See the full description on the dataset page: https://huggingface.co/datasets/RedHatAI/speculator_benchmarks.Awesome_Spatial_VQA_BenchmarksMiroFlow-BenchmarksThese are the benchmarking datasets used for MiroFlow Framework. More information: https://github.com/MiroMindAI/MiroThinker
graphmemix-benchmarks
GraphMemix Benchmarks
Unified multimodal memory benchmark bundles used by
GraphMemix
(arXiv:2608.26983) — four long-term
personalized memory benchmarks with their raw media assets, packaged together
for reproducible evaluation.
Benchmark
Questions
Memories
Track
Upstream license
ATM-Bench (default + hard)
1,044
11,034
memory QA over one multimodal archive
MIT
Mem-Gallery
1,711
7,944
multimodal gallery memory QA
MIT
MemEye
1,855
3,392
comics-derived memory QA… See the full description on the dataset page: https://huggingface.co/datasets/oking0197/graphmemix-benchmarks.audio_benchmarksSBI-benchmarkswhat-ai-benchmarks-actually-measure
What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models
Item-level model responses and scores for 53 language models across the
56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting
Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
(Desai et al., 2026,
arxiv.org/abs/2609.08812).
We do not release the prompts from the benchmark datasets, but instead refer to them by
item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.benchmarks
OpenChainBench Crypto Infrastructure Benchmarks
Daily snapshots of every public benchmark on
openchainbench.com, released as
Hive-partitioned Parquet under CC-BY-4.0.
OCB measures latency, cost, coverage and accuracy of crypto
infrastructure (RPCs, oracles, bridges, data APIs, Polymarket adapters,
Hyperliquid builders). Every snapshot here mirrors the
/api/citable,
/api/stat/<slug>,
and /api/series/<slug>
JSON feeds at the time of capture.
Latest snapshot: 2026-09-21 (captured… See the full description on the dataset page: https://huggingface.co/datasets/OpenChainBench/benchmarks.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.genvsr-video-benchmarksAwesome_Spatial_VQA_Benchmarks_ViewSpatial-BenchCTTA-AD-Benchmarks
CTTA-AD Benchmarks
Dataset collection for CTTA-AD: Continual Test-Time Adaptation for Unified Few-Shot Visual Anomaly Detection (AAAI 2027 submission).
Datasets
Dataset
Domain
Categories
Train Normal
License
MVTec-AD
Industrial
15
209–391 per category
CC BY-NC-SA 4.0
VisA
Industrial
12
400–905 per category
CC BY-NC-SA 4.0
MVTec-LOCO
Logical
5
varies
CC BY-NC-SA 4.0
BrainMRI
Medical
1
7,500
Research only
LiverCT
Medical
1
1,542
Research only… See the full description on the dataset page: https://huggingface.co/datasets/Hammadhaideerr/CTTA-AD-Benchmarks.Squrve-BenchmarksolmOCR-mix-0225-benchmarksetThis is just 10,000 PDFs randomly sampled from https://huggingface.co/datasets/allenai/olmOCR-mix-0225 that we use internally at AI2 for speed benchmarking and also quantization calibration.
benchmarksfastmemory-supremacy-benchmarks
FastMemory: Beyond A Million (BEAM) 10M Audit
Auditing Architectural Integrity at Scale (30 SOTA Wins)
This repository contains the official evaluation logs, simulation code, and technical whitepapers for FastMemory’s 10 Million Token BEAM Benchmark Study.
FastMemory is a sovereign, local-first memory architecture for agentic AI. Unlike traditional vector-based RAG, FastMemory utilizes Topological Isolation to achieve 100% precision in mission-critical reasoning tasks across… See the full description on the dataset page: https://huggingface.co/datasets/fastbuilderai/fastmemory-supremacy-benchmarks.java_evaluation_benchmarksgenvsr-video-benchmarks
GenVSR Video Benchmarks
Public video inputs and restoration outputs used by the GenVSR Visual Comparator.
Each immutable experiment version contains aligned MP4 sources, posters, and a manifest.
Please consult the individual version manifest for source labels and comparison defaults.
mlx-benchmarks
MLX Benchmarks
Structured benchmark results for MLX-quantized and other locally-hosted
LLMs on Apple Silicon. Covers throughput, time-to-first-token, tool-calling,
code generation, reasoning, knowledge, and math suites.
Results are produced by a sweep harness that wires upstream evaluation tools
against a local vllm-mlx inference server:
EleutherAI/lm-evaluation-harness — coding, reasoning, knowledge, math
linusvwe/MLXBench — throughput and time-to-first-token
vllm… See the full description on the dataset page: https://huggingface.co/datasets/JacobPEvans/mlx-benchmarks.ramanv-image-vqa-benchmarksbenchmarks
BitRouter Benchmarks
This dataset release contains the private-router Terminal-Bench 2.1 routing study and its audit evidence. July 2026 experiments are preserved under v0/; the current study, raw Harbor traces, run-scoped BitRouter SQLite files, sanitized production-usage backfill, reproducible joins, and Parquet tables are under terminal-bench-2.1/router-random-study/.
Main result
All valid-case metrics below use the same strict intersection of 80 tasks. “Harbor… See the full description on the dataset page: https://huggingface.co/datasets/BitRouterAI/benchmarks.NT-benchmarks
Dataset Card for Dataset Name
The nucleotide_transformer_downstream_tasks dataset features the 18 downstream tasks presented in the Nucleotide Transformer paper. They consist of both binary and multi-class classification tasks that aim at providing a consistent genomics benchmark.
We note that this is an updated version of this benchmark after the paper has been through peer-review. We highly encourage to move to this version in detriment of the older version.Keypoints about the… See the full description on the dataset page: https://huggingface.co/datasets/mtapiapacheco/NT-benchmarks.perturb-seq-pseudo-pairing-benchmarks
Perturb-seq Pseudo-pairing Benchmarks
Dataset Summary
This repository provides processed single-cell perturbation transcriptomic datasets and representative pseudo-control pairings used to study how pseudo-control construction affects perturbation modeling.
Single-cell perturbation assays are destructive: the same cell cannot be observed before and after perturbation. Cell-level modeling therefore requires an estimated or sampled unperturbed counterpart, referred… See the full description on the dataset page: https://huggingface.co/datasets/JFLa/perturb-seq-pseudo-pairing-benchmarks.Health_Benchmarks
LLM Health Benchmarks Dataset by Yesil Science
The LLM Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties. It provides structured question-answer pairs designed to test the performance of AI models in understanding and generating domain-specific knowledge.
Primary Purpose
This dataset is built to:
Benchmark LLMs in medical specialties and subfields.
Assess the accuracy and contextual… See the full description on the dataset page: https://huggingface.co/datasets/yesilhealth/Health_Benchmarks.SkillOpt_Lite_Benchmarks
SkillOpt_Lite Benchmarks
Train / val / test splits used by the SkillOpt_Lite project.
One multi-config repo containing all six benchmarks:
Config
Rows (train / val / test)
Content shipped
searchqa
400 / 200 / 1400
Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA.
docvqa
107 / 53 / 374
Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.ttpd-benchmarks
TTP-D benchmarks, datasets and results
Problem instances, behaviour-cloning datasets, and solver result tables for the
study Fly, Pack, Drive: the Travelling Thief Problem with Drone.
Paper: Drive, Pack, Fly: The Travelling Thief Problem with Drone
A capacitated truck and a single-package drone operate from a common depot on a
collection route. The truck's velocity decreases affinely with its accumulated
load, so an early pickup penalises every subsequent arc. The drone launches… See the full description on the dataset page: https://huggingface.co/datasets/Murjani/ttpd-benchmarks.data-agent-benchmarks
LongHorizon Full Data-Agent Benchmarks
Companion data artifacts for five complete evaluation tracks:
DataSciBench full55 / 167 metric entries
DABStep full450
DABStep-Research full100
DSBench Modeling full74
LongDS full68 / 2,225 turns
The companion GitHub repository contains processed manifests, evaluation code,
historical API ReAct baseline code, download/preparation tools, and the frozen
source lock. artifact_manifest.json records every uploaded object's size,
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.
