CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01leoschneider /daytrader-benchmarkstabular10M<n<100M0 likes10k downloads5mo agoHugging Face02BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes7k downloads1y agoHugging Face03inria-soda /STRABLE-benchmark STRABLE: Benchmarking Tabular Machine Learning with Strings This dataset card describes the STRABLE benchmark, a comprehensive suite designed for evaluating machine learning models on tabular data containing strings. Dataset Description Benchmarking tabular data has revealed the benefit of dedicated architectures, pushing the state of the art. However, real-world tables often contain string entries beyond pure numbers, a setting that has been understudied due to a… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/STRABLE-benchmark.tabular1M<n<10M1 likes5.4k downloads4mo agoHugging Face04gaia-benchmark /results_public Dataset Card for "resultspublic" More Information needed tabular1K<n<10K26 likes3.9k downloads6h agoHugging Face05VietPhong /kitti-yolo11n-robustness-benchmark KITTI YOLO11n Robustness & Adversarial Benchmark Suite This dataset contains 649,425 benchmark samples evaluating the perception robustness of YOLO11n (Ultralytics YOLOv11 nano in original FP32 precision) on the official KITTI Object Detection train set (3,711 images) under 35 attack & corruption techniques across 5 severity levels. ?? Benchmark Leaderboard (mAP@0.5 Drop on YOLO11n) Clean Baseline AP50: 0.3555 Evaluation Model: YOLO11n (Original weights:… See the full description on the dataset page: https://huggingface.co/datasets/VietPhong/kitti-yolo11n-robustness-benchmark.tabularobject-detection100K<n<1M0 likes3.3k downloads26d agoHugging Face06brettsp /stan-benchmarktabular1M<n<10M0 likes2k downloads6h agoHugging Face07physicl /lighting-invariant-bedroom-perception-robustness-benchmark Lighting-Invariant Bedroom Perception & Robustness Benchmark Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/lighting-invariant-bedroom-perception-robustness-benchmark.imagen<1K0 likes2k downloads3mo agoHugging Face08LEXam-Benchmark /LEXam LEXam: Benchmarking Legal Reasoning on 340 Law Exams A diverse, rigorous evaluation suite for legal AI from Swiss, EU, and international law examinations. Paper | Website & Leaderboard | GitHub Repository 🔥 News [2026/01] Our paper has been accepted to ICLR 2026! [2025/12] We reorganized all multiple-choice questions into four separate files, mcq_4_choices (n = 1,655), mcq_8_choices (n = 1,463), mcq_16_choices (n = 1,028), and mcq_32_choices (n = 550), all… See the full description on the dataset page: https://huggingface.co/datasets/LEXam-Benchmark/LEXam.tabulartext-classification1K<n<10K48 likes1.9k downloads4mo agoHugging Face09lerobot /video-benchmark-resultstabular10K<n<100K2 likes1.6k downloads2mo agoHugging Face10OpenChainBench /benchmarks OpenChainBench Crypto Infrastructure Benchmarks Daily snapshots of every public benchmark on openchainbench.com, released as Hive-partitioned Parquet under CC-BY-4.0. OCB measures latency, cost, coverage and accuracy of crypto infrastructure (RPCs, oracles, bridges, data APIs, Polymarket adapters, Hyperliquid builders). Every snapshot here mirrors the /api/citable, /api/stat/<slug>, and /api/series/<slug> JSON feeds at the time of capture. Latest snapshot: 2026-09-24 (captured… See the full description on the dataset page: https://huggingface.co/datasets/OpenChainBench/benchmarks.tabulartime-series-forecasting1M<n<10M1 likes1.2k downloads11h agoHugging Face11eduagarcia /portuguese_benchmark Portuguese Benchmark This a collection of datasets in Portuguese initially meant to train and evaluate supervised language models such as BERT, RoBERTa, etc... It contains 10 datasets and 18 Tasks for Classification (CLS), NLI, Semantic Similarity Scoring (STS) and Named-Entity Recognition (NER). NER Classification NLI STS LeNER-Br HateBR_offensive_binary assin2-rte assin2-sts UlyssesNER-Br-PL-coarse HateBR_offensive_level UlyssesNER-Br-C-coarse… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/portuguese_benchmark.tabular10K<n<100K7 likes1.2k downloads2y agoHugging Face12superlinked /external-benchmarking Vector Search Benchmarks This repo contains datasets for benchmarking vector search performance, to help Superlinked prioritize integration partners. For performing actual benchmarking on this dataset, see the github repository README. Overview We reviewed a number of publicly available datasets and noted 3 core problems + here is how this dataset fixes them: Problems of other vector search benchmarks How this dataset solves it Not enough metadata of… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/external-benchmarking.image10M<n<100M0 likes1.2k downloads1y agoHugging Face13BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.tabular1M<n<10M0 likes1.1k downloads2y agoHugging Face14aksh-n /deliberative-monitor-benchmarktabular100K<n<1M0 likes678 downloads5mo agoHugging Face15OALL /AlGhafa-Arabic-LLM-Benchmark-Translatedtabular10K<n<100K2 likes656 downloads2y agoHugging Face16CollagenHelixLabs /cdsm_benchmarking_data CDSM Collagen Structure Benchmark — Data Structures and scores for a benchmark comparing a deterministic collagen triple-helix builder (CDSM) against four co-folding models — Boltz-2, Chai-1, Protenix and AlphaFold3, the last in both with-MSA (af3_msa) and no-MSA (af3_nomsa) conditions — on 80 experimentally resolved collagen triple helices from the RCSB PDB. Code: https://github.com/bm-howard/cdsm_benchmarking Layout Prefix Contents Size experimental/… See the full description on the dataset page: https://huggingface.co/datasets/CollagenHelixLabs/cdsm_benchmarking_data.tabular10K<n<100K0 likes637 downloads23d agoHugging Face17YuvrajSingh9886 /jetson-non-reasoning-benchmark-ollama-15w Tiny LLM Benchmark — Jetson Orin Nano Super 8GB Date: 2026-06-07 02:35Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/non-reasoning-models/artifacts/blog-all-20260606-0139-15w Full Results — ollama Power = VDD_CPU_GPU_CV avg over aiperf window. Model Quant ISL OSL OSL mis% TTFT avg p50 p90 p99 T2T avg p50 p90 p99 ITL avg p50 p90 p99 Tok/s Req/s E2E avg p50 p90 p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-15w.tabularn<1K0 likes591 downloads7d agoHugging Face18ppak10 /Additive-Manufacturing-Benchmark Additive Manufacturing Benchmark A benchmark dataset for evaluating knowledge of additive manufacturing (AM) processes, derived from graduate-level coursework at Carnegie Mellon University. Configurations general_knowledge_multiple_choice Multiple-choice questions covering various AM processes with explanations. Column Description source Source homework assignment (e.g. cmu_24_633_2023/homework_1_exone) process AM process type (e.g. Binder Jet… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/Additive-Manufacturing-Benchmark.image1K<n<10K2 likes577 downloads7mo agoHugging Face19YuvrajSingh9886 /jetson-non-reasoning-benchmark-ollama-25w Tiny LLM Benchmark — Jetson Orin Nano Super 8GB Date: 2026-06-23 06:04Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/benchmark-jetson-nano-orin-super/non-reasoning-models/artifacts/blog-all-20260622-0159-25w Full Results — ollama Power = VDD_CPU_GPU_CV avg over aiperf window. Model Quant ISL OSL OSL mis% TTFT avg p50 p90 p99 T2T avg p50 p90 p99 ITL avg p50 p90 p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-25w.tabularn<1K0 likes569 downloads7d agoHugging Face20x-square-robot /xplanner-benchmark XPlanner-Benchmark XPlanner-Benchmark is the portable release of the X-Planner 1,500-episode evaluation benchmark. It contains synchronized multi-view robot-manipulation videos and the episode-level task, subtask, action, scene, duration, and complexity metadata used by X-Planner. Contents 1,500 episodes 3,490 MP4 video references 167 source dataset identifiers 525 unique task names 31 task classes 41 inferred atomic action labels 22.70 total hours of episode… See the full description on the dataset page: https://huggingface.co/datasets/x-square-robot/xplanner-benchmark.tabularrobotics1K<n<10K1 likes524 downloads1d agoHugging Face21MathArena /paper_benchmark Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from all competitions that are incorporated in the MathArena paper. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. problem (string): Problem statement, usually stored as LaTeX source. answer (string): Gold final… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/paper_benchmark.tabularn<1K0 likes518 downloads4mo agoHugging Face22BitRouterAI /benchmarks BitRouter Benchmarks This dataset release contains the private-router Terminal-Bench 2.1 routing study and its audit evidence. July 2026 experiments are preserved under v0/; the current study, raw Harbor traces, run-scoped BitRouter SQLite files, sanitized production-usage backfill, reproducible joins, and Parquet tables are under terminal-bench-2.1/router-random-study/. Main result All valid-case metrics below use the same strict intersection of 80 tasks. “Harbor… See the full description on the dataset page: https://huggingface.co/datasets/BitRouterAI/benchmarks.tabular100K<n<1M1 likes498 downloads15d agoHugging Face23tyrtleli /thinking-benchmark-90 Thinking Benchmark A calibration pool of 90 competition-mathematics problems assembled to study how output / reasoning-trace length varies with problem difficulty across frontier language models. Part of the Cost of Overthinking research project. Dataset at a glance Source n Difficulty Contamination risk AIME 2026 29 3–5 low OlymMATH 41 4–6 medium HMMT February 2026 12 4–5 low MATH-500 5 2–3 high FrontierMath-style 3 6 medium Difficulty is… See the full description on the dataset page: https://huggingface.co/datasets/tyrtleli/thinking-benchmark-90.tabularquestion-answeringn<1K0 likes491 downloads1mo agoHugging Face24YuvrajSingh9886 /jetson-non-reasoning-benchmark-ollama-7w Tiny LLM Benchmark — Jetson Orin Nano Super 8GB Date: 2026-06-09 02:38Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/non-reasoning-models/artifacts/blog-all-20260607-0403-7w Full Results — ollama Power = VDD_CPU_GPU_CV avg over aiperf window. Model Quant ISL OSL OSL mis% TTFT avg p50 p90 p99 T2T avg p50 p90 p99 ITL avg p50 p90 p99 Tok/s Req/s E2E avg p50 p90 p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-7w.tabularn<1K0 likes488 downloads7d agoHugging Face25JacobPEvans /mlx-benchmarks MLX Benchmarks Structured benchmark results for MLX-quantized and other locally-hosted LLMs on Apple Silicon. Covers throughput, time-to-first-token, tool-calling, code generation, reasoning, knowledge, and math suites. Results are produced by a sweep harness that wires upstream evaluation tools against a local vllm-mlx inference server: EleutherAI/lm-evaluation-harness — coding, reasoning, knowledge, math linusvwe/MLXBench — throughput and time-to-first-token vllm… See the full description on the dataset page: https://huggingface.co/datasets/JacobPEvans/mlx-benchmarks.tabular1K<n<10K1 likes475 downloads15d agoHugging Face26ci-benchmark-user /ci-repair-bench CI-REPAIR-BENCH Overview CI-REPAIR-BENCH is a benchmark dataset for research on Continuous Integration (CI) failures and automated repair in Python repositories. The dataset contains 567 CI failure instances collected from 105 real-world GitHub repositories, all written in Python.Each instance captures a CI workflow failure, its logs, the corresponding code diff, and repository-level metadata. Dataset Statistics Programming language: Python Number… See the full description on the dataset page: https://huggingface.co/datasets/ci-benchmark-user/ci-repair-bench.tabularn<1K1 likes469 downloads3d agoHugging Face27m42-health /variant-benchmark Variant Benchmark This benchmark is designed to evaluate how effectively models leverage variant information across diverse biological contexts. Unlike conventional genomic benchmarks that focus primarily on region classification, our approach extends to a broader range of variant-driven molecular processes. Existing assessments, such as BEND and the Genomic Long-Range Benchmark (GLRB), provide valuable insights into specific tasks like noncoding pathogenicity and tissue-specific… See the full description on the dataset page: https://huggingface.co/datasets/m42-health/variant-benchmark.tabulartext-classification1M<n<10M3 likes466 downloads1y agoHugging Face28YuvrajSingh9886 /jetson-non-reasoning-benchmark-7w Tiny LLM Benchmark — Jetson Orin Nano Super 8GB (7W / nvpmodel -m 3) Date: 2026-05-28 18:03Backend: llama.cpp CUDA (-ngl 99) Concurrency: 1Sweep: prompt in {128,512,1024,2048} tok × gen in {64,128,256} tok Artifacts: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-7w Note: tok/J computed from per-run start_time/end_time in each aiperf JSON Full Results Power = VDD_CPU_GPU_CV average over each aiperf run window (per-run timestamps… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-7w.tabularn<1K0 likes457 downloads7d agoHugging Face29YuvrajSingh9886 /jetson-non-reasoning-benchmark-maxn Tiny LLM Benchmark — Jetson Orin Nano Super 8GB Date: 2026-05-26 18:18Backend: llama.cpp CUDA (-ngl 99) Concurrency: 1Sweep: prompt in {128,512,1024,2048} tok × gen in {64,128,256} tok Artifacts: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-maxn Skipped / Failed Models gemma3-4b (OOM — server failed to start) Full Results Cells marked — = OOM (server crashed or skipped). Power = VDD_CPU_GPU_CV average over… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-maxn.tabularn<1K0 likes434 downloads7d agoHugging Face30InstaDeepAI /ms_ninespecies_benchmark Dataset Card for Nine-Species excluding Yeast Dataset used for the baseline comparison of InstaNovo to other models. Dataset Summary Dataset used in the original DeepNovo paper. The training set contains 8 species excluding yeast The validation/test set contains the yeast species Dataset Structure The dataset is tabular, where each row corresponds to a labelled MS2 spectra. sequence (string) The target peptide sequence excluding… See the full description on the dataset page: https://huggingface.co/datasets/InstaDeepAI/ms_ninespecies_benchmark.tabular100K<n<1M3 likes427 downloads14d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.