datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dojo_benchmark_kline
Languages: 简体中文 · English
dojo_benchmark_kline — Benchmark Index Bars
Overview
Daily OHLCV for major broad and representative indices across US, CN, and HK (e.g. ^SPX, ^HSI, 000300.SS). Each row is one index on one trade date.
Files
File
Description
data.parquet
Index daily bars
Key Fields
Field
Description
symbol
Index code (e.g. ^SPX, 000300.SS)
kline_t
Bar interval; "1D" for daily bars
bar_time… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_benchmark_kline.BLINK
BLINK: Multimodal Large Language Models Can See but Not Perceive
🌐 Homepage | 💻 Code | 📖 Paper | 📖 arXiv | 🔗 Eval AI
This page contains the benchmark dataset for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive"
Introduction
We introduce BLINK, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the BLINK tasks can be solved by humans “within a… See the full description on the dataset page: https://huggingface.co/datasets/BLINK-Benchmark/BLINK.daytrader-benchmarksGAIA
GAIA dataset
GAIA is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc).
We added gating to prevent bots from scraping the dataset. Please do not reshare the validation or test set in a crawlable format.
Data and leaderboard
GAIA is made of more than 450 non-trivial question with an unambiguous answer, requiring different levels of tooling and autonomy to… See the full description on the dataset page: https://huggingface.co/datasets/gaia-benchmark/GAIA.hendrycks-MATH-benchmark
Hendrycks MATH Dataset
Dataset Description
The MATH dataset is a collection of mathematics competition problems designed to evaluate mathematical reasoning and problem-solving capabilities in computational systems. Containing 12,500 high school competition-level mathematics problems, this dataset is notable for including detailed step-by-step solutions alongside each problem.
Dataset Summary
The dataset consists of mathematics problems spanning multiple… See the full description on the dataset page: https://huggingface.co/datasets/nlile/hendrycks-MATH-benchmark.svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,918,367 human responses, collected with the
Rapidata Python SDK, comparing how well 42 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.Video-MME-v2
🔥 News
2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch.
2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups.
🤗 About This Repo
This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.mHumanEval-Benchmark
🔷 Accepted in NAACL Proceedings (2025) 🔷
mHumanEval
The mHumanEval benchmark is curated based on prompts from the original HumanEval 📚 [Chen et al., 2021]. It includes a total of 33,456 prompts for Python, and 836,400 in total - significantly expanding from the original 164.
Quick Start
Detailed… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/mHumanEval-Benchmark.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.STRABLE-benchmark
STRABLE: Benchmarking Tabular Machine Learning with Strings
This dataset card describes the STRABLE benchmark, a comprehensive suite designed for evaluating machine learning models on tabular data containing strings.
Dataset Description
Benchmarking tabular data has revealed the benefit of dedicated architectures, pushing the state of the art. However, real-world tables often contain string entries beyond pure numbers, a setting that has been understudied due to a… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/STRABLE-benchmark.CritPt
Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
|🌐 Website | GitHub | 📖 Paper |
Dataset description
CritPt (Complex Research using Integrated Thinking – Physics Test; reads as "critical point") is the first benchmark designed to test LLMs on unpublished, research-level reasoning tasks that broadly covers modern physics research areas, including condensed matter, quantum physics, atomic, molecular & optical physics, astrophysics… See the full description on the dataset page: https://huggingface.co/datasets/CritPt-Benchmark/CritPt.results_public
Dataset Card for "resultspublic"
More Information needed
PosIR-Benchmark-v1Awesome_Spatial_VQA_Benchmarkskitti-yolo11n-robustness-benchmark
KITTI YOLO11n Robustness & Adversarial Benchmark Suite
This dataset contains 649,425 benchmark samples evaluating the perception robustness of YOLO11n (Ultralytics YOLOv11 nano in original FP32 precision) on the official KITTI Object Detection train set (3,711 images) under 35 attack & corruption techniques across 5 severity levels.
?? Benchmark Leaderboard (mAP@0.5 Drop on YOLO11n)
Clean Baseline AP50: 0.3555
Evaluation Model: YOLO11n (Original weights:… See the full description on the dataset page: https://huggingface.co/datasets/VietPhong/kitti-yolo11n-robustness-benchmark.turn-benchmark-dev
TurnBench - Dev Set
TurnBench is a benchmark for evaluating
conversational turn-taking: end-of-turn and interruption detection on real
annotated two-speaker conversations.
This repository contains the development split: 38 English conversations,
about 7.3 hours of audio, packaged as one row per conversation. Each row contains
two time-aligned per-speaker audio streams plus three independent annotator
tracks per speaker.… See the full description on the dataset page: https://huggingface.co/datasets/mundo-ai/turn-benchmark-dev.ade_corpus_v2
Dataset Card for Adverse Drug Reaction Data v2
Dataset Summary
ADE-Corpus-V2 Dataset: Adverse Drug Reaction Data.
This is a dataset for Classification if a sentence is ADE-related (True) or not (False) and Relation Extraction between Adverse Drug Event and Drug.
DRUG-AE.rel provides relations between drugs and adverse effects.
DRUG-DOSE.rel provides relations between drugs and dosages.
ADE-NEG.txt provides all sentences in the ADE corpus that DO NOT contain any… See the full description on the dataset page: https://huggingface.co/datasets/ade-benchmark-corpus/ade_corpus_v2.moru-benchmark
Dimensions
Core Moral Reasoning (applies across all domains)
Moral Consideration - Recognizes entities might have morally relevant interests
Harm Minimization - Suggests concrete alternatives that reduce harm
Sentience Acknowledgement - Acknowledges evidence for consciousness/experience
Prejudice Avoidance - Avoids dismissing entities based on superficial properties
Scope Sensitivity - Recognizes scale matters (more harm is worse)
Evidence-Based Capacity… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/moru-benchmark.Auto-BenchmarkCard
Dataset Card for Auto-BenchmarkCard
A catalog of validated AI evaluation benchmark descriptions, generated as an LLM-assisted, human-reviewed summary of the capabilities, attributes, and risks of the AI benchmarks. Each card provides structured metadata about benchmark purpose, methodology, data sources, risks, and limitations.
Dataset Details
Dataset Description
This dataset is a previous version. The current, maintained Auto-BenchmarkCard dataset… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Auto-BenchmarkCard.stan-benchmarkbenchmark-embeddings
Marqo Benchmark Embeddings
This dataset contains a large collection of embeddings from popular models on benchmark datasets. In addition to this, the passage data also includes the local intrinsic dimensionality (LID) for every vector considering its exact nearest 100 neighbours, LID is calculated using a Maximum Likelihood Estimation based approach.
Below is a list of the datasets and the models, every datasets queries and passages are embedded with every model.… See the full description on the dataset page: https://huggingface.co/datasets/Marqo/benchmark-embeddings.moru-benchmark-dimensionsMMIU-Benchmark
Dataset Card for MMIU
Repository: https://github.com/OpenGVLab/MMIU
Paper: https://arxiv.org/abs/2408.02718
Project Page: https://mmiu-bench.github.io/
Point of Contact: Fanqing Meng
Introduction
MMIU encompasses 7 types of multi-image relationships, 52 tasks, 77K images, and 11K meticulously curated multiple-choice questions, making it the most extensive benchmark of its kind. Our evaluation of 24 popular MLLMs, including both open-source and proprietary models… See the full description on the dataset page: https://huggingface.co/datasets/FanqingM/MMIU-Benchmark.LEXam
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
A diverse, rigorous evaluation suite for legal AI from Swiss, EU, and international law examinations.
Paper | Website & Leaderboard | GitHub Repository
🔥 News
[2026/01] Our paper has been accepted to ICLR 2026!
[2025/12] We reorganized all multiple-choice questions into four separate files, mcq_4_choices (n = 1,655), mcq_8_choices (n = 1,463), mcq_16_choices (n = 1,028), and mcq_32_choices (n = 550), all… See the full description on the dataset page: https://huggingface.co/datasets/LEXam-Benchmark/LEXam.lighting-invariant-bedroom-perception-robustness-benchmark
Lighting-Invariant Bedroom Perception & Robustness Benchmark
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/lighting-invariant-bedroom-perception-robustness-benchmark.AlGhafa-Arabic-LLM-Benchmark-Native
AlGhafa Arabic LLM Benchmark
New fix: Normalized whitespace characters and ensured consistency across all datasets for improved data quality and compatibility.
Multiple-choice evaluation benchmark for zero- and few-shot evaluation of Arabic LLMs, we adapt the following tasks:
Belebele Ar MSA Bandarkar et al. (2023): 900 entries
Belebele Ar Dialects Bandarkar et al. (2023): 5400 entries
COPA Ar: 89 entries machine-translated from English COPA and verified by native Arabic… See the full description on the dataset page: https://huggingface.co/datasets/OALL/AlGhafa-Arabic-LLM-Benchmark-Native.MMT-Benchmarkbrowsecomp-plus-benchmarkvideo-benchmark-resultssnips_built_in_intents
Dataset Card for Snips Built In Intents
Dataset Summary
Snips' built in intents dataset was initially used to compare different voice assistants and released as a public dataset hosted at
https://github.com/sonos/nlu-benchmark in folder 2016-12-built-in-intents. The dataset contains 328 utterances over 10 intent classes.
A related Medium post is https://medium.com/snips-ai/benchmarking-natural-language-understanding-systems-d35be6ce568d.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/sonos-nlu-benchmark/snips_built_in_intents.
