CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MME-Benchmarks /Video-MME-v2 🔥 News 2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch. 2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups. 🤗 About This Repo This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.textvideo-text-to-text1K<n<10K48 likes9.1k downloads1mo agoHugging Face02leoschneider /daytrader-benchmarkstabular10M<n<100M0 likes8.6k downloads5mo agoHugging Face03LLDDSS /Awesome_Spatial_VQA_Benchmarksimage10K<n<100K1 likes4.1k downloads1y agoHugging Face04madesai /what-ai-benchmarks-actually-measure What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models Item-level model responses and scores for 53 language models across the 56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks (Desai et al., 2026, arxiv.org/abs/2609.08812). We do not release the prompts from the benchmark datasets, but instead refer to them by item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.tabular1M<n<10M0 likes1.3k downloads13d agoHugging Face05OpenChainBench /benchmarks OpenChainBench Crypto Infrastructure Benchmarks Daily snapshots of every public benchmark on openchainbench.com, released as Hive-partitioned Parquet under CC-BY-4.0. OCB measures latency, cost, coverage and accuracy of crypto infrastructure (RPCs, oracles, bridges, data APIs, Polymarket adapters, Hyperliquid builders). Every snapshot here mirrors the /api/citable, /api/stat/<slug>, and /api/series/<slug> JSON feeds at the time of capture. Latest snapshot: 2026-09-22 (captured… See the full description on the dataset page: https://huggingface.co/datasets/OpenChainBench/benchmarks.tabulartime-series-forecasting1M<n<10M1 likes1.2k downloads16h agoHugging Face06witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.1k downloads2d agoHugging Face07LLDDSS /Awesome_Spatial_VQA_Benchmarks_ViewSpatial-Benchimage1K<n<10K0 likes751 downloads1y agoHugging Face08minhhien0811 /java_evaluation_benchmarkstext1K<n<10K0 likes612 downloads4mo agoHugging Face09lingamvamshikrishnareddy /ramanv-image-vqa-benchmarksgatedtext100K<n<1M0 likes515 downloads24d agoHugging Face10JacobPEvans /mlx-benchmarks MLX Benchmarks Structured benchmark results for MLX-quantized and other locally-hosted LLMs on Apple Silicon. Covers throughput, time-to-first-token, tool-calling, code generation, reasoning, knowledge, and math suites. Results are produced by a sweep harness that wires upstream evaluation tools against a local vllm-mlx inference server: EleutherAI/lm-evaluation-harness — coding, reasoning, knowledge, math linusvwe/MLXBench — throughput and time-to-first-token vllm… See the full description on the dataset page: https://huggingface.co/datasets/JacobPEvans/mlx-benchmarks.tabular1K<n<10K1 likes507 downloads13d agoHugging Face11BitRouterAI /benchmarks BitRouter Benchmarks This dataset release contains the private-router Terminal-Bench 2.1 routing study and its audit evidence. July 2026 experiments are preserved under v0/; the current study, raw Harbor traces, run-scoped BitRouter SQLite files, sanitized production-usage backfill, reproducible joins, and Parquet tables are under terminal-bench-2.1/router-random-study/. Main result All valid-case metrics below use the same strict intersection of 80 tasks. “Harbor… See the full description on the dataset page: https://huggingface.co/datasets/BitRouterAI/benchmarks.tabular100K<n<1M1 likes500 downloads13d agoHugging Face12yesilhealth /Health_Benchmarks LLM Health Benchmarks Dataset by Yesil Science The LLM Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties. It provides structured question-answer pairs designed to test the performance of AI models in understanding and generating domain-specific knowledge. Primary Purpose This dataset is built to: Benchmark LLMs in medical specialties and subfields. Assess the accuracy and contextual… See the full description on the dataset page: https://huggingface.co/datasets/yesilhealth/Health_Benchmarks.textquestion-answering1K<n<10K10 likes446 downloads1y agoHugging Face13noel7Y /data-agent-benchmarks LongHorizon Full Data-Agent Benchmarks Companion data artifacts for five complete evaluation tracks: DataSciBench full55 / 167 metric entries DABStep full450 DABStep-Research full100 DSBench Modeling full74 LongDS full68 / 2,225 turns The companion GitHub repository contains processed manifests, evaluation code, historical API ReAct baseline code, download/preparation tools, and the frozen source lock. artifact_manifest.json records every uploaded object's size, SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.imagequestion-answering1 likes402 downloads12d agoHugging Face14yyyang /UI-Grounding-Benchmarks UI-Grounding-Benchmarks This is a collection of UI grounding benchmarks: ScreenSpot ScreenSpot-V2 ScreenSpot-Pro OS-World-G UI-Vision Thanks for their great work! This benchmark collection is used in the paper: FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection 🖼️ Project Page: https://showlab.github.io/FocusUI/ 🏠 Github Repo: https://github.com/showlab/FocusUI 📝 Paper: https://arxiv.org/pdf/2601.03928 Model Zoo Model Backbone 🤗… See the full description on the dataset page: https://huggingface.co/datasets/yyyang/UI-Grounding-Benchmarks.image1K<n<10K2 likes398 downloads8mo agoHugging Face15Murjani /ttpd-benchmarks TTP-D benchmarks, datasets and results Problem instances, behaviour-cloning datasets, and solver result tables for the study Fly, Pack, Drive: the Travelling Thief Problem with Drone. Paper: Drive, Pack, Fly: The Travelling Thief Problem with Drone A capacitated truck and a single-package drone operate from a common depot on a collection route. The truck's velocity decreases affinely with its accumulated load, so an early pickup penalises every subsequent arc. The drone launches… See the full description on the dataset page: https://huggingface.co/datasets/Murjani/ttpd-benchmarks.textreinforcement-learning1 likes394 downloads1mo agoHugging Face16carlahq /demo-tabular-benchmarks 📊 Carla HQ Tabular Foundation Model Benchmarks Centralized benchmark repository of canonical tabular datasets curated for Carla HQ and TabICL (In-Context Learning foundation models for tabular data). Each dataset is hosted as an independent subset/config with native Parquet storage, schema qualities, OpenML source links, and synchronized Google Sheets for live spreadsheet experimentation. 🚀 Quickstart & Download Options Option 1: Using… See the full description on the dataset page: https://huggingface.co/datasets/carlahq/demo-tabular-benchmarks.tabular100K<n<1M0 likes391 downloads26d agoHugging Face17yshenaw /SkillOpt_Lite_Benchmarks SkillOpt_Lite Benchmarks Train / val / test splits used by the SkillOpt_Lite project. One multi-config repo containing all six benchmarks: Config Rows (train / val / test) Content shipped searchqa 400 / 200 / 1400 Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA. docvqa 107 / 53 / 374 Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.imagequestion-answering1K<n<10K0 likes390 downloads3mo agoHugging Face18katarinagresova /Genomic_Benchmarks_human_nontata_promoters Dataset Card for "Genomic_Benchmarks_human_nontata_promoters" More Information needed text10K<n<100K0 likes374 downloads4y agoHugging Face19harisarang /benchmark-scifacttext1K<n<10K0 likes360 downloads10mo agoHugging Face20Keylab /OCSR-Benchmarks OCSR Benchmarks A collection of ten benchmark datasets for Optical Chemical Structure Recognition (OCSR) — the task of converting chemical structure diagram images into machine-readable SMILES strings. These benchmarks were used to evaluate the COMO model (Closed-Loop Optical Molecule Recognition). Subsets Config Split Size Domain CLEF test 992 Real JPO test 449 Real UOB test 5,740 Real USPTO test 5719 Real USPTO-10K test 9,999 Real Staker test 50,000… See the full description on the dataset page: https://huggingface.co/datasets/Keylab/OCSR-Benchmarks.imageimage-to-text10K<n<100K3 likes355 downloads4mo agoHugging Face21lthn /LEM-benchmarks LEM-benchmarks Canonical 8-PAC benchmark results for the Lemma model family. This dataset is an aggregated store of per-round evaluation data produced by lthn/LEM-Eval. Every row represents one model's answer to one question in one round of a paired A/B run against its unmodified base, and the dataset grows monotonically as more workers contribute — different machines, different sampling states, different hardware paths — which is the whole point of 8-PAC: multiple independent… See the full description on the dataset page: https://huggingface.co/datasets/lthn/LEM-benchmarks.tabularquestion-answering10K<n<100K3 likes347 downloads5mo agoHugging Face22katarinagresova /Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs Dataset Card for "Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs" More Information needed text100K<n<1M4 likes326 downloads3y agoHugging Face23katarinagresova /Genomic_Benchmarks_human_enhancers_cohn Dataset Card for "Genomic_Benchmarks_human_enhancers_cohn" More Information needed text10K<n<100K2 likes325 downloads4y agoHugging Face24katielink /genomic-benchmarks Genomic Benchmark In this repository, we collect benchmarks for classification of genomic sequences. It is shipped as a Python package, together with functions helping to download & manipulate datasets and train NN models. Citing Genomic Benchmarks If you use Genomic Benchmarks in your research, please cite it as follows. Text GRESOVA, Katarina, et al. Genomic Benchmarks: A Collection of Datasets for Genomic Sequence Classification. bioRxiv, 2022.… See the full description on the dataset page: https://huggingface.co/datasets/katielink/genomic-benchmarks.tabular100K<n<1M7 likes313 downloads3y agoHugging Face25katarinagresova /Genomic_Benchmarks_human_ocr_ensembl Dataset Card for "Genomic_Benchmarks_human_ocr_ensembl" More Information needed text100K<n<1M0 likes268 downloads4y agoHugging Face26mtapiapacheco /screen-benchmarkstext1M<n<10M0 likes256 downloads2mo agoHugging Face27katarinagresova /Genomic_Benchmarks_human_ensembl_regulatory Dataset Card for "Genomic_Benchmarks_human_ensembl_regulatory" More Information needed text100K<n<1M2 likes255 downloads2y agoHugging Face28hxxiang /dna_benchmarks DNA Benchmarks Dataset Description DNA Benchmarks is a collection of genomic datasets organized for benchmarking DNA foundation models, genomic representation learning methods, and multimodal genomic learning frameworks. The collection covers a wide range of sequence scales, from short regulatory sequences of a few hundred base pairs to genome-scale tiling datasets derived from whole genomes. This repository is designed as a general-purpose benchmark hub rather… See the full description on the dataset page: https://huggingface.co/datasets/hxxiang/dna_benchmarks.texttext-classificationn<1K1 likes255 downloads2mo agoHugging Face29omegaprime669 /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.imagetext-generationn<1K0 likes255 downloads2mo agoHugging Face30RISys-Lab /Benchmarks_CyberSec_RedSageMCQ Dataset Card for RedSage-MCQ Dataset Summary RedSage-MCQ is a large-scale, high-quality multiple-choice question (MCQ) benchmark designed to evaluate the cybersecurity knowledge, skills, and tool proficiency of Large Language Models (LLMs). It is a component of the RedSage-Bench suite introduced in the paper "RedSage: A Cybersecurity Generalist LLM". The dataset comprises 30,000 questions derived from RedSage-Seed, a curated collection of authoritative… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_RedSageMCQ.tabularquestion-answering10K<n<100K0 likes253 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.