CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alexshpunt /explicit-edit-benchmark Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.tabulartext-generationn<1K2 likes7.8k downloads2d agoHugging Face02leoschneider /daytrader-benchmarkstabular10M<n<100M0 likes7.6k downloads5mo agoHugging Face03BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes6.3k downloads1y agoHugging Face04YuePanEdward /regx-benchmark RegX Cross-Domain Multi-View Point Cloud Registration Benchmark RegX evaluates multi-view point cloud registration across scales spanning nine orders of magnitude — nanometre-scale microscopy to kilometre-scale airborne maps — and sensors never designed to be compared: clinical colonoscopes, RGB-D cameras, spinning and solid-state LiDAR, terrestrial and airborne laser scanners. Most registration benchmarks fix one sensor and one scale. RegX asks a narrower question instead: does… See the full description on the dataset page: https://huggingface.co/datasets/YuePanEdward/regx-benchmark.3dother1K<n<10K2 likes5.2k downloads17d agoHugging Face05inria-soda /STRABLE-benchmark STRABLE: Benchmarking Tabular Machine Learning with Strings This dataset card describes the STRABLE benchmark, a comprehensive suite designed for evaluating machine learning models on tabular data containing strings. Dataset Description Benchmarking tabular data has revealed the benefit of dedicated architectures, pushing the state of the art. However, real-world tables often contain string entries beyond pure numbers, a setting that has been understudied due to a… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/STRABLE-benchmark.tabular1M<n<10M1 likes4.9k downloads3mo agoHugging Face06gaia-benchmark /results_public Dataset Card for "resultspublic" More Information needed tabular1K<n<10K26 likes3.8k downloads15h agoHugging Face07VietPhong /kitti-yolo11n-robustness-benchmark KITTI YOLO11n Robustness & Adversarial Benchmark Suite This dataset contains 649,425 benchmark samples evaluating the perception robustness of YOLO11n (Ultralytics YOLOv11 nano in original FP32 precision) on the official KITTI Object Detection train set (3,711 images) under 35 attack & corruption techniques across 5 severity levels. ?? Benchmark Leaderboard (mAP@0.5 Drop on YOLO11n) Clean Baseline AP50: 0.3555 Evaluation Model: YOLO11n (Original weights:… See the full description on the dataset page: https://huggingface.co/datasets/VietPhong/kitti-yolo11n-robustness-benchmark.tabularobject-detection100K<n<1M0 likes3.3k downloads24d agoHugging Face08dacorvo /funes-handoff-recall-benchmark handover-vs-recall A long investigation bloats an agent session until each new turn costs more to carry the context than to do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task, on tasks that genuinely require the prior investigation: arm channel A branch-only switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.tabularn<1K0 likes3.1k downloads19d agoHugging Face09inria-soda /tabular-benchmark Tabular Benchmark Dataset Description This dataset is a curation of various datasets from openML and is curated to benchmark performance of various machine learning algorithms. Repository: https://github.com/LeoGrin/tabular-benchmark/community Paper: https://hal.archives-ouvertes.fr/hal-03723551v2/document Dataset Summary Benchmark made of curation of various tabular data learning tasks, including: Regression from Numerical and Categorical Features… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/tabular-benchmark.tabulartabular-classification10M<n<100M51 likes2.4k downloads3y agoHugging Face10ibm-research /data-product-benchmark DPDisc Dataset Paper | Code Dataset Description This dataset provides a benchmark for automatic data product creation. The task is framed as follows: given a natural language data product request and a corpus of text and tables, the objective is to identify the relevant tables and text documents that should be included in the resulting data product which would useful to the given data product request. The benchmark brings together three variants: HybridQA, TAT-QA, and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/data-product-benchmark.texttable-question-answering10K<n<100K3 likes2.1k downloads6mo agoHugging Face11physicl /lighting-invariant-bedroom-perception-robustness-benchmark Lighting-Invariant Bedroom Perception & Robustness Benchmark Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/lighting-invariant-bedroom-perception-robustness-benchmark.imagen<1K0 likes2.1k downloads3mo agoHugging Face12LEXam-Benchmark /LEXam LEXam: Benchmarking Legal Reasoning on 340 Law Exams A diverse, rigorous evaluation suite for legal AI from Swiss, EU, and international law examinations. Paper | Website & Leaderboard | GitHub Repository 🔥 News [2026/01] Our paper has been accepted to ICLR 2026! [2025/12] We reorganized all multiple-choice questions into four separate files, mcq_4_choices (n = 1,655), mcq_8_choices (n = 1,463), mcq_16_choices (n = 1,028), and mcq_32_choices (n = 550), all… See the full description on the dataset page: https://huggingface.co/datasets/LEXam-Benchmark/LEXam.tabulartext-classification1K<n<10K48 likes1.8k downloads4mo agoHugging Face13lerobot /video-benchmark-resultstabular10K<n<100K2 likes1.6k downloads2mo agoHugging Face14brettsp /stan-benchmarktabular1M<n<10M0 likes1.6k downloads4h agoHugging Face15Merserk /Krea-2-Turbo-Checkpoint-Format-Benchmark Krea 2 Turbo ComfyUI Format Fidelity Benchmark This release is a paired, deterministic comparison of eight Krea 2 Turbo checkpoint formats in ComfyUI: BF16, FP8 Scaled, INT8 ConvRot, MXFP8, NVFP4, INT4 ConvRot W4A4, GGUF Q8_0, and GGUF Q4_K_M. It contains 240 scored 1024×1024 images, saved float32 decoded tensors and final latents, every denoising trajectory, raw metric tables, telemetry, statistical comparisons, and reproduction code. Main result BF16 is the… See the full description on the dataset page: https://huggingface.co/datasets/Merserk/Krea-2-Turbo-Checkpoint-Format-Benchmark.imagetext-to-imagen<1K5 likes1.6k downloads2mo agoHugging Face16madesai /what-ai-benchmarks-actually-measure What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models Item-level model responses and scores for 53 language models across the 56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks (Desai et al., 2026, arxiv.org/abs/2609.08812). We do not release the prompts from the benchmark datasets, but instead refer to them by item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.tabular1M<n<10M0 likes1.3k downloads12d agoHugging Face17OpenChainBench /benchmarks OpenChainBench Crypto Infrastructure Benchmarks Daily snapshots of every public benchmark on openchainbench.com, released as Hive-partitioned Parquet under CC-BY-4.0. OCB measures latency, cost, coverage and accuracy of crypto infrastructure (RPCs, oracles, bridges, data APIs, Polymarket adapters, Hyperliquid builders). Every snapshot here mirrors the /api/citable, /api/stat/<slug>, and /api/series/<slug> JSON feeds at the time of capture. Latest snapshot: 2026-09-21 (captured… See the full description on the dataset page: https://huggingface.co/datasets/OpenChainBench/benchmarks.tabulartime-series-forecasting1M<n<10M1 likes1.2k downloads19h agoHugging Face18Keh0t0 /scene-mem-benchmark scene-mem-benchmark A benchmark for scene memory in embodied agents: an agent watches a mobile manipulator work in a house for several minutes, then is asked to retrieve an object it has to remember — one that was moved, dropped, or merely seen along the way — or (resume) to go back and finish the job it was interrupted in, remembering how far it had got — or (routine) to put a new object away where this household keeps that kind of thing, a rule it was never told and can only… See the full description on the dataset page: https://huggingface.co/datasets/Keh0t0/scene-mem-benchmark.tabularrobotics1K<n<10K0 likes1.2k downloads5d agoHugging Face19witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.2k downloads2d agoHugging Face20EleutherAI /hack-ignition-benchmark hack-ignition benchmark — data, v0.1.6 Training trajectories of reinforcement-learning runs on exploitable graders, for studying and predicting when RL comes to produce exploits. Each family is a set of GRPO runs over configurations of (start model, prompt, training set, grader / reward structure, recipe), with one or more seeds per configuration. Every family stores what its training logs contain — per-step exploit, task and reward rates, the item × step exploit record… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/hack-ignition-benchmark.tabular100K<n<1M1 likes1.1k downloads19h agoHugging Face21BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.tabular1M<n<10M0 likes1.1k downloads2y agoHugging Face22superlinked /external-benchmarking Vector Search Benchmarks This repo contains datasets for benchmarking vector search performance, to help Superlinked prioritize integration partners. For performing actual benchmarking on this dataset, see the github repository README. Overview We reviewed a number of publicly available datasets and noted 3 core problems + here is how this dataset fixes them: Problems of other vector search benchmarks How this dataset solves it Not enough metadata of… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/external-benchmarking.image10M<n<100M0 likes1.1k downloads1y agoHugging Face23eduagarcia /portuguese_benchmark Portuguese Benchmark This a collection of datasets in Portuguese initially meant to train and evaluate supervised language models such as BERT, RoBERTa, etc... It contains 10 datasets and 18 Tasks for Classification (CLS), NLI, Semantic Similarity Scoring (STS) and Named-Entity Recognition (NER). NER Classification NLI STS LeNER-Br HateBR_offensive_binary assin2-rte assin2-sts UlyssesNER-Br-PL-coarse HateBR_offensive_level UlyssesNER-Br-C-coarse… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/portuguese_benchmark.tabular10K<n<100K7 likes1.1k downloads2y agoHugging Face24imodels /tabular-benchmark-797-classificationtabular1K<n<10K0 likes1k downloads3y agoHugging Face25Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K6 likes920 downloads11mo agoHugging Face26skyan1002 /flash-flood-benchmark-data TORRENT — CONUS Flash-Flood Benchmark (L1–L3), agent-friendly Traceable, Observation-constrained, Rapid-Response, Episode–gauge–watershed Network of Testbeds: reproducible flash-flood testbeds for hydrological-response analysis and model intercomparison. This mirror carries the paper-matched v1.0 release (companion paper: TORRENT, Earth System Science Data). Archive of record (v1.0): https://doi.org/10.5281/zenodo.22118051 (concept DOI, always the latest version:… See the full description on the dataset page: https://huggingface.co/datasets/skyan1002/flash-flood-benchmark-data.tabular100K<n<1M0 likes861 downloads26d agoHugging Face27lyrain2001 /Auto-Fill-Benchmark Auto-Fill Benchmark Benchmark for predicting missing cell values in real-world tables, introduced in Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models (PVLDB 19(11), 2026 — arXiv:2607.19847). Each case is a real table in which exactly one cell is replaced by [MISSING], together with the ground-truth value. Code: https://github.com/lyrain2001/auto-fill Models: Auto-Fill-Qwen3-8B-Knowledge · Auto-Fill-Qwen3-8B-Reasoning ·… See the full description on the dataset page: https://huggingface.co/datasets/lyrain2001/Auto-Fill-Benchmark.tabulartable-question-answering1K<n<10K0 likes702 downloads24d agoHugging Face28aksh-n /deliberative-monitor-benchmarktabular100K<n<1M0 likes678 downloads5mo agoHugging Face29OALL /AlGhafa-Arabic-LLM-Benchmark-Translatedtabular10K<n<100K2 likes671 downloads2y agoHugging Face30anonymous-structured-agent /structured-file-audit-benchmark Paper Data Release This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them. Contents datasets/ Benchmark data and per-task manifests for the three paper-facing splits. datasets/sc_flat/data SC-Flat is derived from DaBench, augmented with a replayable perturbation injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.texttable-question-answering1 likes658 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.