CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlphaDojo /dojo_benchmark_kline Languages: 简体中文 · English dojo_benchmark_kline — Benchmark Index Bars Overview Daily OHLCV for major broad and representative indices across US, CN, and HK (e.g. ^SPX, ^HSI, 000300.SS). Each row is one index on one trade date. Files File Description data.parquet Index daily bars Key Fields Field Description symbol Index code (e.g. ^SPX, 000300.SS) kline_t Bar interval; "1D" for daily bars bar_time… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_benchmark_kline.text1K<n<10K1 likes14k downloads8h agoHugging Face02BLINK-Benchmark /BLINK BLINK: Multimodal Large Language Models Can See but Not Perceive 🌐 Homepage | 💻 Code | 📖 Paper | 📖 arXiv | 🔗 Eval AI This page contains the benchmark dataset for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive" Introduction We introduce BLINK, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the BLINK tasks can be solved by humans “within a… See the full description on the dataset page: https://huggingface.co/datasets/BLINK-Benchmark/BLINK.image1K<n<10K49 likes13k downloads1y agoHugging Face03leoschneider /daytrader-benchmarkstabular10M<n<100M0 likes11k downloads5mo agoHugging Face04gaia-benchmark /GAIAgated GAIA dataset GAIA is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc). We added gating to prevent bots from scraping the dataset. Please do not reshare the validation or test set in a crawlable format. Data and leaderboard GAIA is made of more than 450 non-trivial question with an unambiguous answer, requiring different levels of tooling and autonomy to… See the full description on the dataset page: https://huggingface.co/datasets/gaia-benchmark/GAIA.audion<1K846 likes10k downloads11mo agoHugging Face05nlile /hendrycks-MATH-benchmark Hendrycks MATH Dataset Dataset Description The MATH dataset is a collection of mathematics competition problems designed to evaluate mathematical reasoning and problem-solving capabilities in computational systems. Containing 12,500 high school competition-level mathematics problems, this dataset is notable for including detailed step-by-step solutions alongside each problem. Dataset Summary The dataset consists of mathematics problems spanning multiple… See the full description on the dataset page: https://huggingface.co/datasets/nlile/hendrycks-MATH-benchmark.text10K<n<100K33 likes10k downloads2y agoHugging Face06Rapidata /svg-benchmark Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,918,367 human responses, collected with the Rapidata Python SDK, comparing how well 42 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment). The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.imagetext-to-image100K<n<1M35 likes9.9k downloads1mo agoHugging Face07MME-Benchmarks /Video-MME-v2 🔥 News 2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch. 2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups. 🤗 About This Repo This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.textvideo-text-to-text1K<n<10K49 likes9k downloads2mo agoHugging Face08md-nishat-008 /mHumanEval-Benchmark 🔷 Accepted in NAACL Proceedings (2025) 🔷 mHumanEval The mHumanEval benchmark is curated based on prompts from the original HumanEval 📚 [Chen et al., 2021]. It includes a total of 33,456 prompts for Python, and 836,400 in total - significantly expanding from the original 164. Quick Start Detailed… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/mHumanEval-Benchmark.text100K<n<1M4 likes7.6k downloads1y agoHugging Face09BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes7k downloads1y agoHugging Face10inria-soda /STRABLE-benchmark STRABLE: Benchmarking Tabular Machine Learning with Strings This dataset card describes the STRABLE benchmark, a comprehensive suite designed for evaluating machine learning models on tabular data containing strings. Dataset Description Benchmarking tabular data has revealed the benefit of dedicated architectures, pushing the state of the art. However, real-world tables often contain string entries beyond pure numbers, a setting that has been understudied due to a… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/STRABLE-benchmark.tabular1M<n<10M1 likes5.6k downloads4mo agoHugging Face11CritPt-Benchmark /CritPt Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark |🌐 Website | GitHub | 📖 Paper | Dataset description CritPt (Complex Research using Integrated Thinking – Physics Test; reads as "critical point") is the first benchmark designed to test LLMs on unpublished, research-level reasoning tasks that broadly covers modern physics research areas, including condensed matter, quantum physics, atomic, molecular & optical physics, astrophysics… See the full description on the dataset page: https://huggingface.co/datasets/CritPt-Benchmark/CritPt.textn<1K27 likes4.5k downloads10mo agoHugging Face12gaia-benchmark /results_public Dataset Card for "resultspublic" More Information needed tabular1K<n<10K26 likes3.9k downloads1d agoHugging Face13infgrad /PosIR-Benchmark-v1text100K<n<1M3 likes3.4k downloads10mo agoHugging Face14LLDDSS /Awesome_Spatial_VQA_Benchmarksimage10K<n<100K1 likes3.1k downloads1y agoHugging Face15VietPhong /kitti-yolo11n-robustness-benchmark KITTI YOLO11n Robustness & Adversarial Benchmark Suite This dataset contains 649,425 benchmark samples evaluating the perception robustness of YOLO11n (Ultralytics YOLOv11 nano in original FP32 precision) on the official KITTI Object Detection train set (3,711 images) under 35 attack & corruption techniques across 5 severity levels. ?? Benchmark Leaderboard (mAP@0.5 Drop on YOLO11n) Clean Baseline AP50: 0.3555 Evaluation Model: YOLO11n (Original weights:… See the full description on the dataset page: https://huggingface.co/datasets/VietPhong/kitti-yolo11n-robustness-benchmark.tabularobject-detection100K<n<1M0 likes3.1k downloads28d agoHugging Face16mundo-ai /turn-benchmark-devgated TurnBench - Dev Set TurnBench is a benchmark for evaluating conversational turn-taking: end-of-turn and interruption detection on real annotated two-speaker conversations. This repository contains the development split: 38 English conversations, about 7.3 hours of audio, packaged as one row per conversation. Each row contains two time-aligned per-speaker audio streams plus three independent annotator tracks per speaker.… See the full description on the dataset page: https://huggingface.co/datasets/mundo-ai/turn-benchmark-dev.audiovoice-activity-detectionn<1K8 likes2.7k downloads1mo agoHugging Face17ade-benchmark-corpus /ade_corpus_v2 Dataset Card for Adverse Drug Reaction Data v2 Dataset Summary ADE-Corpus-V2 Dataset: Adverse Drug Reaction Data. This is a dataset for Classification if a sentence is ADE-related (True) or not (False) and Relation Extraction between Adverse Drug Event and Drug. DRUG-AE.rel provides relations between drugs and adverse effects. DRUG-DOSE.rel provides relations between drugs and dosages. ADE-NEG.txt provides all sentences in the ADE corpus that DO NOT contain any… See the full description on the dataset page: https://huggingface.co/datasets/ade-benchmark-corpus/ade_corpus_v2.texttext-classification10K<n<100K36 likes2.3k downloads3y agoHugging Face18CompassioninMachineLearning /moru-benchmark Dimensions Core Moral Reasoning (applies across all domains) Moral Consideration - Recognizes entities might have morally relevant interests Harm Minimization - Suggests concrete alternatives that reduce harm Sentience Acknowledgement - Acknowledges evidence for consciousness/experience Prejudice Avoidance - Avoids dismissing entities based on superficial properties Scope Sensitivity - Recognizes scale matters (more harm is worse) Evidence-Based Capacity… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/moru-benchmark.textn<1K0 likes2.1k downloads2mo agoHugging Face19ibm-research /Auto-BenchmarkCard Dataset Card for Auto-BenchmarkCard A catalog of validated AI evaluation benchmark descriptions, generated as an LLM-assisted, human-reviewed summary of the capabilities, attributes, and risks of the AI benchmarks. Each card provides structured metadata about benchmark purpose, methodology, data sources, risks, and limitations. Dataset Details Dataset Description This dataset is a previous version. The current, maintained Auto-BenchmarkCard dataset… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Auto-BenchmarkCard.textn<1K3 likes2k downloads3mo agoHugging Face20brettsp /stan-benchmarktabular1M<n<10M0 likes2k downloads11h agoHugging Face21Marqo /benchmark-embeddings Marqo Benchmark Embeddings This dataset contains a large collection of embeddings from popular models on benchmark datasets. In addition to this, the passage data also includes the local intrinsic dimensionality (LID) for every vector considering its exact nearest 100 neighbours, LID is calculated using a Maximum Likelihood Estimation based approach. Below is a list of the datasets and the models, every datasets queries and passages are embedded with every model.… See the full description on the dataset page: https://huggingface.co/datasets/Marqo/benchmark-embeddings.text100M<n<1B6 likes2k downloads2y agoHugging Face22CompassioninMachineLearning /moru-benchmark-dimensionstextn<1K0 likes2k downloads7mo agoHugging Face23FanqingM /MMIU-Benchmark Dataset Card for MMIU Repository: https://github.com/OpenGVLab/MMIU Paper: https://arxiv.org/abs/2408.02718 Project Page: https://mmiu-bench.github.io/ Point of Contact: Fanqing Meng Introduction MMIU encompasses 7 types of multi-image relationships, 52 tasks, 77K images, and 11K meticulously curated multiple-choice questions, making it the most extensive benchmark of its kind. Our evaluation of 24 popular MLLMs, including both open-source and proprietary models… See the full description on the dataset page: https://huggingface.co/datasets/FanqingM/MMIU-Benchmark.image10K<n<100K11 likes1.9k downloads2y agoHugging Face24LEXam-Benchmark /LEXam LEXam: Benchmarking Legal Reasoning on 340 Law Exams A diverse, rigorous evaluation suite for legal AI from Swiss, EU, and international law examinations. Paper | Website & Leaderboard | GitHub Repository 🔥 News [2026/01] Our paper has been accepted to ICLR 2026! [2025/12] We reorganized all multiple-choice questions into four separate files, mcq_4_choices (n = 1,655), mcq_8_choices (n = 1,463), mcq_16_choices (n = 1,028), and mcq_32_choices (n = 550), all… See the full description on the dataset page: https://huggingface.co/datasets/LEXam-Benchmark/LEXam.tabulartext-classification1K<n<10K48 likes1.9k downloads4mo agoHugging Face25physicl /lighting-invariant-bedroom-perception-robustness-benchmark Lighting-Invariant Bedroom Perception & Robustness Benchmark Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/lighting-invariant-bedroom-perception-robustness-benchmark.imagen<1K0 likes1.9k downloads3mo agoHugging Face26OALL /AlGhafa-Arabic-LLM-Benchmark-Native AlGhafa Arabic LLM Benchmark New fix: Normalized whitespace characters and ensured consistency across all datasets for improved data quality and compatibility. Multiple-choice evaluation benchmark for zero- and few-shot evaluation of Arabic LLMs, we adapt the following tasks: Belebele Ar MSA Bandarkar et al. (2023): 900 entries Belebele Ar Dialects Bandarkar et al. (2023): 5400 entries COPA Ar: 89 entries machine-translated from English COPA and verified by native Arabic… See the full description on the dataset page: https://huggingface.co/datasets/OALL/AlGhafa-Arabic-LLM-Benchmark-Native.text10K<n<100K7 likes1.8k downloads3y agoHugging Face27lmms-lab-encoder /MMT-Benchmarkimage10K<n<100K0 likes1.8k downloads2y agoHugging Face28timchen0618 /browsecomp-plus-benchmarktextn<1K0 likes1.6k downloads4mo agoHugging Face29lerobot /video-benchmark-resultstabular10K<n<100K2 likes1.5k downloads2mo agoHugging Face30sonos-nlu-benchmark /snips_built_in_intents Dataset Card for Snips Built In Intents Dataset Summary Snips' built in intents dataset was initially used to compare different voice assistants and released as a public dataset hosted at https://github.com/sonos/nlu-benchmark in folder 2016-12-built-in-intents. The dataset contains 328 utterances over 10 intent classes. A related Medium post is https://medium.com/snips-ai/benchmarking-natural-language-understanding-systems-d35be6ce568d. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/sonos-nlu-benchmark/snips_built_in_intents.texttext-classificationn<1K14 likes1.3k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.