datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
daytrader-benchmarksmulti_task_multi_modal_knowledge_retrieval_benchmark_M2KR
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.STRABLE-benchmark
STRABLE: Benchmarking Tabular Machine Learning with Strings
This dataset card describes the STRABLE benchmark, a comprehensive suite designed for evaluating machine learning models on tabular data containing strings.
Dataset Description
Benchmarking tabular data has revealed the benefit of dedicated architectures, pushing the state of the art. However, real-world tables often contain string entries beyond pure numbers, a setting that has been understudied due to a… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/STRABLE-benchmark.results_public
Dataset Card for "resultspublic"
More Information needed
kitti-yolo11n-robustness-benchmark
KITTI YOLO11n Robustness & Adversarial Benchmark Suite
This dataset contains 649,425 benchmark samples evaluating the perception robustness of YOLO11n (Ultralytics YOLOv11 nano in original FP32 precision) on the official KITTI Object Detection train set (3,711 images) under 35 attack & corruption techniques across 5 severity levels.
?? Benchmark Leaderboard (mAP@0.5 Drop on YOLO11n)
Clean Baseline AP50: 0.3555
Evaluation Model: YOLO11n (Original weights:… See the full description on the dataset page: https://huggingface.co/datasets/VietPhong/kitti-yolo11n-robustness-benchmark.stan-benchmarklighting-invariant-bedroom-perception-robustness-benchmark
Lighting-Invariant Bedroom Perception & Robustness Benchmark
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/lighting-invariant-bedroom-perception-robustness-benchmark.LEXam
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
A diverse, rigorous evaluation suite for legal AI from Swiss, EU, and international law examinations.
Paper | Website & Leaderboard | GitHub Repository
🔥 News
[2026/01] Our paper has been accepted to ICLR 2026!
[2025/12] We reorganized all multiple-choice questions into four separate files, mcq_4_choices (n = 1,655), mcq_8_choices (n = 1,463), mcq_16_choices (n = 1,028), and mcq_32_choices (n = 550), all… See the full description on the dataset page: https://huggingface.co/datasets/LEXam-Benchmark/LEXam.video-benchmark-resultsbenchmarks
OpenChainBench Crypto Infrastructure Benchmarks
Daily snapshots of every public benchmark on
openchainbench.com, released as
Hive-partitioned Parquet under CC-BY-4.0.
OCB measures latency, cost, coverage and accuracy of crypto
infrastructure (RPCs, oracles, bridges, data APIs, Polymarket adapters,
Hyperliquid builders). Every snapshot here mirrors the
/api/citable,
/api/stat/<slug>,
and /api/series/<slug>
JSON feeds at the time of capture.
Latest snapshot: 2026-09-24 (captured… See the full description on the dataset page: https://huggingface.co/datasets/OpenChainBench/benchmarks.portuguese_benchmark
Portuguese Benchmark
This a collection of datasets in Portuguese initially meant to train and evaluate supervised language models such as BERT, RoBERTa, etc...
It contains 10 datasets and 18 Tasks for Classification (CLS), NLI, Semantic Similarity Scoring (STS) and Named-Entity Recognition (NER).
NER
Classification
NLI
STS
LeNER-Br
HateBR_offensive_binary
assin2-rte
assin2-sts
UlyssesNER-Br-PL-coarse
HateBR_offensive_level
UlyssesNER-Br-C-coarse… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/portuguese_benchmark.external-benchmarking
Vector Search Benchmarks
This repo contains datasets for benchmarking vector search performance, to help Superlinked prioritize integration partners.
For performing actual benchmarking on this dataset, see the github repository README.
Overview
We reviewed a number of publicly available datasets and noted 3 core problems + here is how this dataset fixes them:
Problems of other vector search benchmarks
How this dataset solves it
Not enough metadata of… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/external-benchmarking.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.deliberative-monitor-benchmarkAlGhafa-Arabic-LLM-Benchmark-Translatedcdsm_benchmarking_data
CDSM Collagen Structure Benchmark — Data
Structures and scores for a benchmark comparing a deterministic collagen
triple-helix builder (CDSM) against four co-folding models — Boltz-2, Chai-1,
Protenix and AlphaFold3, the last in both with-MSA (af3_msa) and no-MSA
(af3_nomsa) conditions — on 80 experimentally resolved collagen triple
helices from the RCSB PDB.
Code: https://github.com/bm-howard/cdsm_benchmarking
Layout
Prefix
Contents
Size
experimental/… See the full description on the dataset page: https://huggingface.co/datasets/CollagenHelixLabs/cdsm_benchmarking_data.jetson-non-reasoning-benchmark-ollama-15w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-07 02:35Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/non-reasoning-models/artifacts/blog-all-20260606-0139-15w
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90
p99
Tok/s
Req/s
E2E avg
p50
p90
p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-15w.Additive-Manufacturing-Benchmark
Additive Manufacturing Benchmark
A benchmark dataset for evaluating knowledge of additive manufacturing (AM) processes, derived from graduate-level coursework at Carnegie Mellon University.
Configurations
general_knowledge_multiple_choice
Multiple-choice questions covering various AM processes with explanations.
Column
Description
source
Source homework assignment (e.g. cmu_24_633_2023/homework_1_exone)
process
AM process type (e.g. Binder Jet… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/Additive-Manufacturing-Benchmark.jetson-non-reasoning-benchmark-ollama-25w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-23 06:04Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/benchmark-jetson-nano-orin-super/non-reasoning-models/artifacts/blog-all-20260622-0159-25w
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90
p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-25w.xplanner-benchmark
XPlanner-Benchmark
XPlanner-Benchmark is the portable release of the X-Planner 1,500-episode evaluation benchmark. It contains synchronized multi-view robot-manipulation videos and the episode-level task, subtask, action, scene, duration, and complexity metadata used by X-Planner.
Contents
1,500 episodes
3,490 MP4 video references
167 source dataset identifiers
525 unique task names
31 task classes
41 inferred atomic action labels
22.70 total hours of episode… See the full description on the dataset page: https://huggingface.co/datasets/x-square-robot/xplanner-benchmark.paper_benchmark
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from all competitions that are incorporated in the MathArena paper.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
problem (string): Problem statement, usually stored as LaTeX source.
answer (string): Gold final… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/paper_benchmark.benchmarks
BitRouter Benchmarks
This dataset release contains the private-router Terminal-Bench 2.1 routing study and its audit evidence. July 2026 experiments are preserved under v0/; the current study, raw Harbor traces, run-scoped BitRouter SQLite files, sanitized production-usage backfill, reproducible joins, and Parquet tables are under terminal-bench-2.1/router-random-study/.
Main result
All valid-case metrics below use the same strict intersection of 80 tasks. “Harbor… See the full description on the dataset page: https://huggingface.co/datasets/BitRouterAI/benchmarks.thinking-benchmark-90
Thinking Benchmark
A calibration pool of 90 competition-mathematics problems assembled to study how output / reasoning-trace length varies with problem difficulty across frontier language models. Part of the Cost of Overthinking research project.
Dataset at a glance
Source
n
Difficulty
Contamination risk
AIME 2026
29
3–5
low
OlymMATH
41
4–6
medium
HMMT February 2026
12
4–5
low
MATH-500
5
2–3
high
FrontierMath-style
3
6
medium
Difficulty is… See the full description on the dataset page: https://huggingface.co/datasets/tyrtleli/thinking-benchmark-90.jetson-non-reasoning-benchmark-ollama-7w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-09 02:38Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/non-reasoning-models/artifacts/blog-all-20260607-0403-7w
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90
p99
Tok/s
Req/s
E2E avg
p50
p90
p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-7w.mlx-benchmarks
MLX Benchmarks
Structured benchmark results for MLX-quantized and other locally-hosted
LLMs on Apple Silicon. Covers throughput, time-to-first-token, tool-calling,
code generation, reasoning, knowledge, and math suites.
Results are produced by a sweep harness that wires upstream evaluation tools
against a local vllm-mlx inference server:
EleutherAI/lm-evaluation-harness — coding, reasoning, knowledge, math
linusvwe/MLXBench — throughput and time-to-first-token
vllm… See the full description on the dataset page: https://huggingface.co/datasets/JacobPEvans/mlx-benchmarks.ci-repair-bench
CI-REPAIR-BENCH
Overview
CI-REPAIR-BENCH is a benchmark dataset for research on Continuous Integration (CI) failures and automated repair in Python repositories.
The dataset contains 567 CI failure instances collected from 105 real-world GitHub repositories, all written in Python.Each instance captures a CI workflow failure, its logs, the corresponding code diff, and repository-level metadata.
Dataset Statistics
Programming language: Python
Number… See the full description on the dataset page: https://huggingface.co/datasets/ci-benchmark-user/ci-repair-bench.variant-benchmark
Variant Benchmark
This benchmark is designed to evaluate how effectively models leverage variant information across diverse biological contexts.
Unlike conventional genomic benchmarks that focus primarily on region classification, our approach extends to a broader range of variant-driven molecular processes.
Existing assessments, such as BEND and the Genomic Long-Range Benchmark (GLRB),
provide valuable insights into specific tasks like noncoding pathogenicity and tissue-specific… See the full description on the dataset page: https://huggingface.co/datasets/m42-health/variant-benchmark.jetson-non-reasoning-benchmark-7w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB (7W / nvpmodel -m 3)
Date: 2026-05-28 18:03Backend: llama.cpp CUDA (-ngl 99) Concurrency: 1Sweep: prompt in {128,512,1024,2048} tok × gen in {64,128,256} tok
Artifacts: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-7w
Note: tok/J computed from per-run start_time/end_time in each aiperf JSON
Full Results
Power = VDD_CPU_GPU_CV average over each aiperf run window (per-run timestamps… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-7w.jetson-non-reasoning-benchmark-maxn
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-05-26 18:18Backend: llama.cpp CUDA (-ngl 99) Concurrency: 1Sweep: prompt in {128,512,1024,2048} tok × gen in {64,128,256} tok
Artifacts: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-maxn
Skipped / Failed Models
gemma3-4b (OOM — server failed to start)
Full Results
Cells marked — = OOM (server crashed or skipped). Power = VDD_CPU_GPU_CV average over… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-maxn.ms_ninespecies_benchmark
Dataset Card for Nine-Species excluding Yeast
Dataset used for the baseline comparison of InstaNovo to other models.
Dataset Summary
Dataset used in the original DeepNovo paper.
The training set contains 8 species excluding yeast
The validation/test set contains the yeast species
Dataset Structure
The dataset is tabular, where each row corresponds to a labelled MS2 spectra.
sequence (string) The target peptide sequence excluding… See the full description on the dataset page: https://huggingface.co/datasets/InstaDeepAI/ms_ninespecies_benchmark.
