datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
benchmarking_sbi_runs
Benchmarking SBI Runs
This dataset contains the raw, per-run results underlying the manuscript
"Benchmarking Simulation-Based Inference"
(Lueckmann, Boelts, Greenberg, Goncalves & Macke, AISTATS 2021).
It is a direct migration of the Git LFS data from
mackelab/benchmarking_sbi_runs on GitHub.
For compiled, ready-to-use dataframes built from these raw results (and the code that produced
them), see the companion repository:… See the full description on the dataset page: https://huggingface.co/datasets/mackelab/benchmarking_sbi_runs.PDE_Inverse_Problem_Benchmarking
PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems
This is the official dataset for the paper PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems.
Code: GitHub - ASK-Berkeley/PDEInvBench
Sample Usage
You can use the provided script from the codebase to batch download the data:
pip install huggingface_hub
python3 huggingface_pdeinv_download.py --dataset… See the full description on the dataset page: https://huggingface.co/datasets/DabbyOWL/PDE_Inverse_Problem_Benchmarking.Agri_STT_Benchmarking_Dataset
Agri STT Benchmarking Dataset
10,808 farmer voice queries in Hindi, Telugu and Odia, with reference transcripts, for benchmarking automatic speech recognition in agricultural contexts. The audio is included in this repository.
Every recording is a smallholder farmer speaking a question to Farmer.Chat, an AI advisory service run by Digital Green. Reference transcripts were produced by human annotators. Nothing here is read from a script or recorded in a studio, so the audio… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Agri_STT_Benchmarking_Dataset.external-benchmarking
Vector Search Benchmarks
This repo contains datasets for benchmarking vector search performance, to help Superlinked prioritize integration partners.
For performing actual benchmarking on this dataset, see the github repository README.
Overview
We reviewed a number of publicly available datasets and noted 3 core problems + here is how this dataset fixes them:
Problems of other vector search benchmarks
How this dataset solves it
Not enough metadata of… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/external-benchmarking.benchmark-dk-interaktivt-benchmarkingunivers
Benchmark.dk - interaktivt benchmarkingunivers (komplet data-høst)
Komplet høst af datamaterialet bag Indenrigs- og Sundhedsministeriets
Benchmarkingenheds interaktive benchmarkingunivers:
https://www.benchmark.dk/interaktivt-benchmarkingunivers. Høstet 1. juni 2026.
Del af Silkeborg Skoleatlas - 4 af de deri indeholdte
skole/dagtilbud-tabeller indgår også kurateret i atlassets egen
silkeborg-benchmark-national-datasæt,
men dette repo er den fulde, ukuraterede kilde: alle 6… See the full description on the dataset page: https://huggingface.co/datasets/Skoleatlas/benchmark-dk-interaktivt-benchmarkingunivers.cdsm_benchmarking_data
CDSM Collagen Structure Benchmark — Data
Structures and scores for a benchmark comparing a deterministic collagen
triple-helix builder (CDSM) against four co-folding models — Boltz-2, Chai-1,
Protenix and AlphaFold3, the last in both with-MSA (af3_msa) and no-MSA
(af3_nomsa) conditions — on 80 experimentally resolved collagen triple
helices from the RCSB PDB.
Code: https://github.com/bm-howard/cdsm_benchmarking
Layout
Prefix
Contents
Size
experimental/… See the full description on the dataset page: https://huggingface.co/datasets/CollagenHelixLabs/cdsm_benchmarking_data.map-anything-benchmarking
MapAnything Benchmarking Dataset
Dataset Description
This dataset contains the WAI format data used for benchmarking feed-forward 3D reconstruction models in the MapAnything codebase.
Please see our Data Processing README for more details.
Citation
If you use this dataset in your research, please cite our paper:
@inproceedings{keetha2026mapanything,
title={{MapAnything}: Universal Feed-Forward Metric {3D} Reconstruction},
author={Nikhil Keetha and Norman… See the full description on the dataset page: https://huggingface.co/datasets/facebook/map-anything-benchmarking.groundtruth-dynamic-benchmarking-submissions
Groundtruth Dynamic Benchmarking — Geology — Submissions
Community-submitted evaluation runs against the groundtruth-dynamic-benchmarking geology rubrics, feeding the leaderboard. We are currently running two tracks: model benchmarking (comparing different models with no special harness) and harness benchmarking (comparing different harnesses using a single standard model - GLM 4.7).
Each submission is a pointwise rubric score: one model, scored 0–10 per question against a… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking-submissions.groundtruth-dynamic-benchmarking
Groundtruth Dynamic Benchmarking — Geology
Question sets and grading rubrics for evaluating LLMs on real-world geological
reasoning. Every question is authored from a real source corpus, and every
claim in the grading key carries an evidence locator back to that corpus —
nothing is synthetic. Licensing/redistribution status varies by corpus — see
License.
This dataset holds the questions, grading rubrics, and source corpora.
Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.Enlatics_benchmarking
GAIA-style Evaluation Results (Public)
This dataset contains GAIA-inspired benchmark question results for LLM evaluation.
What is inside
grok_answers.csv: question-by-question outputs, model used, response time (seconds), and run status.
Notes
These tasks are designed in a GAIA-style (multi-hop, web-grounded questions).
Official GAIA leaderboard submission access was restricted for our account, so results are published here for transparency and… See the full description on the dataset page: https://huggingface.co/datasets/enlatics/Enlatics_benchmarking.malayalam_common_voice_benchmarkingmalayalam_msc_benchmarkingDLCspeed_benchmarking
Dataset Card for DLC Speed Benchmarking ZIP
This supports the dlc-live benchmarking zip formally hosted on our Harvard Rowland server.
All information can be found in our publication:
Real-time, low-latency closed-loop feedback using markerless posture tracking
Gary A Kane, Gonçalo Lopes, Jonny L Saunders, Alexander Mathis, Mackenzie W Mathis
https://elifesciences.org/articles/61909
Direct Use
"""
DeepLabCut Toolbox (deeplabcut.org)
© A. & M. Mathis Labs
Licensed… See the full description on the dataset page: https://huggingface.co/datasets/mwmathis/DLCspeed_benchmarking.URSA-benchmarking-sets
URSA benchmarking sets
Benchmark for reaction plausibility, single-step and multistep retrosynthesis from Zagribelnyy et al. (2026).
It bundles two task families: reaction-level plausibility judgment and retrosynthesis target sets.
Reaction plausibility
URSA-reaction-plausibility-bench-2026.csv — 1,000 reactions predicted by different models hand-labeled by expert oragnic chemists (500 plausible / 500 implausible), for benchmarking reaction-level… See the full description on the dataset page: https://huggingface.co/datasets/insilicomedicine/URSA-benchmarking-sets.protein-fitness-datasets-for-benchmarking-ft-esm2-strategies
Protein Fitness Datasets for Benchmarking ESM-2 Fine-Tuning Strategies
Dataset description
This repository contains processed protein sequence–function datasets for CreiLOV, avGFP, and Ube4b.
The variants and experimental measurements were obtained in previously published deep mutational scanning studies:
Chen, Y. et al. Deep Mutational Scanning of an Oxygen-Independent Fluorescent Protein CreiLOV for Comprehensive Profiling of Mutational and Epistatic Effects.… See the full description on the dataset page: https://huggingface.co/datasets/RomeroLab-Duke/protein-fitness-datasets-for-benchmarking-ft-esm2-strategies.benchmarking-the-benchmarks-data
Benchmarking the Benchmarks — Raw SLM Safety Evaluation Runs
Raw evaluation data for the ESORICS 2026 paper:
Nyamtulla Shaik, Fengjun Li, Bo Luo — University of Kansas
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models.
ESORICS 2026. arXiv:2608.17183
Code and processed data: https://github.com/nyamtulla/benchmarking-the-benchmarks
⚠️ Content warning
This dataset contains adversarial safety prompts and model responses… See the full description on the dataset page: https://huggingface.co/datasets/nyamtulla/benchmarking-the-benchmarks-data.minipile_benchmarkingcontext_length_benchmarking
🧠 Context Length - Benchmarking
A Mathematical Framework for Long-Context Attention Evaluation
The Context Length Benchmarking, developed by Sapiens Technology®, is a deterministic and scalable framework designed to evaluate how effectively large language models retain and retrieve information across extremely long contexts, isolating pure attention capability by removing semantic complexity and focusing on distributed anomaly detection; the methodology involves normalizing the… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/context_length_benchmarking.Detecture_ICLR_Benchmarking
Detecture ICLR Benchmarking
The four evaluation routes reported in the ICLR 2027 submission on sub-semantic
image segmentation, bundled so the published numbers can be reproduced from a
single download.
This is a smaller, paper-aligned release. An earlier bundle,
aviadcohz/Detecture_Benchmarking,
carried five datasets at 4.3 GB for a previous version of this work. That one
included routes the current paper does not report. This release carries only what
the paper evaluates on.… See the full description on the dataset page: https://huggingface.co/datasets/aviadcohz/Detecture_ICLR_Benchmarking.fluid-benchmarking
Fluid Language Model Benchmarking
This dataset provides IRT models for ARC Challenge,
GSM8K,
HellaSwag,
MMLU,
TruthfulQA, and
WinoGrande.
Furthermore, it contains
results for pretraining checkpoints of Amber-6.7B,
K2-65B,
OLMo1-7B,
OLMo2-7B,
Pythia-2.8B, and
Pythia-6.9B, evaluated on these six benchmarks.
🚀 Usage
For utilities to use the dataset and to replicate the results from the paper, please see the corresponding GitHub… See the full description on the dataset page: https://huggingface.co/datasets/allenai/fluid-benchmarking.Celldega_Visualization_Benchmarkingos-world-modifiedcross_species_benchmarking**Repository: https://d-script.readthedocs.io/en/stable/data.html
**Reference: Sledzieski, S., Singh, R., Cowen, L. & Berger, B. D-SCRIPT translates genome to phenome with sequence-based, structure-aware, genome-scale predictions of protein-protein interactions. Cell Systems 12, 969-982.e6 (2021).
wikidata_benchmarkingThis dataset was compiled for the purpose of finetuning models in the context of benchmarking for art historical research.
Images scraped from Wikimedia Commons via Wikidata; metadata scraped from
Wikidata (CC0). Image licenses vary per file (predominantly public domain,
some CC-BY-SA) see the license_short_name / license_url columns in the
parquet files for the exact terms of each individual image, and the
commons file page for full details.
UM-DLP-Public-Benchmarking-Dataset
UM DLP Public Benchmarking Dataset
Description
The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement.
This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks:
Financial Data (Account information about… See the full description on the dataset page: https://huggingface.co/datasets/alibustami/UM-DLP-Public-Benchmarking-Dataset.benchmarking-cultures-25
Benchmarking-Cultures-25 Dataset
This dataset accompanies the Unsteady Metrics and Benchmarking Cultures of AI Model Builders paper submitted to FAccT 2026 by Stefan Baack, Christo Buschek and Maty Bohacek.
The dataset contains the following parts:
core: The curated Benchmarking-Cultures-25 dataset.
derived: Datasets that were derived from the core dataset and informed the FAccT submission.
figures: Figures generated from derived data and used in the paper.
docs: Data dictionaries… See the full description on the dataset page: https://huggingface.co/datasets/matybohacek/benchmarking-cultures-25.Benchmarking_tpLMs_dataosworld-benchmarking-gold-filesDetecture_Benchmarking
Detecture Benchmarking Suite
A five-dataset benchmark suite for VLM-guided multi-texture segmentation, released alongside the Detecture architecture (VLM-guided multi-texture segmentation via multiplexed grounding).
This bundle contains the training set, one in-domain test set, and three out-of-domain evaluation benchmarks used to score Detecture against baseline model families (SAM 3 vanilla, Grounded-SAM 3, Sa2VA, and Qwen2SAM zero-shot) in the Detecture paper.
Layout… See the full description on the dataset page: https://huggingface.co/datasets/aviadcohz/Detecture_Benchmarking.Bernett_benchmarking
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Dataset Sources [optional]
**Repository: https://doi.org/10.6084/m9.figshare.21591618.v3
**Reference:
Bernett, J., Blumenthal, D. B. & List, M. Cracking the black box of deep sequence-based protein–protein interaction prediction. Briefings in Bioinformatics 25, bbae076… See the full description on the dataset page: https://huggingface.co/datasets/danliu1226/Bernett_benchmarking.
