CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TIGER-Lab /MMLU-Pro MMLU-Pro Dataset MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. |Github | 🏆Leaderboard | 📖Paper | 🚀 What's New [2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro.tabularquestion-answering10K<n<100K516 likes241k downloads5mo agoHugging Face02futurehouse /lab-bench LAB-Bench The Language Agent Biology Benchmark, or LAB-Bench, is an evaluation dataset for AI systems intended to benchmark capabilities foundational to scientific research in biology. The dataset currently consists of 8 broad categories, comprising 30 narrower subtasks, including extracting information from the scientific literature (LitQA2), retrieving information from databases (DbQA) and supplementary information (SuppQA), reasoning about scientific figures (FigQA) and tables… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/lab-bench.imagequestion-answering1K<n<10K51 likes33k downloads1y agoHugging Face03TAUR-Lab /MuSR MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning Creating murder mysteries that require multi-step reasoning with commonsense using ChatGPT! By: Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett. View the dataset on our custom viewer and project website! Check out the paper. Appeared at ICLR 2024 as a spotlight presentation! Git Repo with the source data, how to recreate the dataset (and create new ones!) here textquestion-answeringn<1K24 likes18k downloads2y agoHugging Face04EdisonScientific /labbench2gated LABBench2 LABBench2 is a benchmark for measuring real-world capabilities of AI systems performing scientific research tasks. It is an evolution of the Language Agent Biology Benchmark (LAB-Bench), comprising nearly 1,900 tasks that measure similar capabilities but in more realistic contexts. LABBench2 provides a meaningful jump in difficulty over LAB-Bench (model-specific accuracy differences range from −26% to −46% across subtasks), underscoring continued room for improvement.… See the full description on the dataset page: https://huggingface.co/datasets/EdisonScientific/labbench2.textquestion-answering1K<n<10K60 likes9.7k downloads7mo agoHugging Face05TIGER-Lab /WebInstructSub 🦣 MAmmoTH2: Scaling Instructions from the Web Project Page: https://tiger-ai-lab.github.io/MAmmoTH2/ Paper: https://arxiv.org/pdf/2405.03548 Code: https://github.com/TIGER-AI-Lab/MAmmoTH2 WebInstruct (Subset) This repo contains the partial dataset used in "MAmmoTH2: Scaling Instructions from the Web". This partial data is coming mostly from the forums like stackexchange. This subset contains very high-quality data to boost LLM performance through instruction tuning.… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/WebInstructSub.textquestion-answering1M<n<10M164 likes6.1k downloads2y agoHugging Face06haydn-jones /labbench2-fixed LABBench2 PMID-enriched public mirror This is a public, schema-compatible mirror of EdisonScientific/labbench2, pinned to upstream revision 27d12d72af24e3f70db8a99df63e567366cbdb80. Original columns and source URLs are unchanged. Two columns are added to every configuration: pmids: deduplicated PubMed identifiers resolved for the row's sources. source_pmids: aligned one-to-one with sources; unresolved or non-PubMed sources are null. LABBench2 LABBench2 is a… See the full description on the dataset page: https://huggingface.co/datasets/haydn-jones/labbench2-fixed.textquestion-answering1K<n<10K0 likes3.1k downloads1mo agoHugging Face07TIGER-Lab /TheoremQA Dataset Card for "TheoremQA" Introduction We propose the first question-answering dataset driven by STEM theorems. We annotated 800 QA pairs covering 350+ theorems spanning across Math, EE&CS, Physics and Finance. The dataset is collected by human experts with very high quality. We provide the dataset as a new benchmark to test the limit of large language models to apply theorems to solve challenging university-level questions. We provide a pipeline in the following to… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/TheoremQA.imagequestion-answeringn<1K20 likes2.9k downloads2y agoHugging Face08ssswwwxxx /labbench2 LABBench2 LABBench2 is a benchmark for measuring real-world capabilities of AI systems performing scientific research tasks. It is an evolution of the Language Agent Biology Benchmark (LAB-Bench), comprising nearly 1,900 tasks that measure similar capabilities but in more realistic contexts. LABBench2 provides a meaningful jump in difficulty over LAB-Bench (model-specific accuracy differences range from −26% to −46% across subtasks), underscoring continued room for improvement.… See the full description on the dataset page: https://huggingface.co/datasets/ssswwwxxx/labbench2.textquestion-answering1K<n<10K0 likes2.6k downloads2mo agoHugging Face09lmms-lab-encoder /LLaVA-NeXT-Interleave-Bench LLaVA-Interleave Bench Dataset Card Dataset details Dataset type: LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API. It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs. Dataset date: LLaVA-Interleave Bench was collected in April 2024, and released in June 2024. Paper or resources for more information: Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.imagevisual-question-answering10K<n<100K15 likes2.6k downloads2y agoHugging Face10TIGER-Lab /VisualWebInstruct VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search VisualWebInstruct is a large-scale, diverse multimodal instruction dataset designed to enhance vision-language models' reasoning capabilities. The dataset contains approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs associated with 163,743 unique images, while the remaining 60% are text-only QA pairs. Please also checkout our more recent verified version at Huggingface.… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct.imagequestion-answering1M<n<10M44 likes2.2k downloads8mo agoHugging Face11TIGER-Lab /VisualWebInstruct-Recall Introduction This is the dataset recalled from Google Search from the seed images. Links Github| Paper| Website Citation @article{visualwebinstruct, title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search}, author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu}, journal={arXiv preprint arXiv:2503.10582}, year={2025} } imagequestion-answering100K<n<1M4 likes1.8k downloads2y agoHugging Face12PaDaS-Lab /webfaqWebFAQ Q&A Dataset Overview | Details | Structure | Examples | Considerations | License | Citation | Contact | Acknowledgement Overview The WebFAQ Q&A Dataset is a broad-coverage corpus of 96 million natural question-answer (QA) pairs in 75 languages, gathered from FAQ pages on the web. It leverages structured schema.org FAQPage annotations, making it a unique resource for large-scale Question Answering… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq.textquestion-answering10M<n<100M24 likes1.4k downloads1y agoHugging Face13TIGER-Lab /WebInstruct-verified General-Reasoner: Advancing LLM Reasoning Across All Domains 💻 Code | 📄 Paper | 📊 Dataset | 🤗 Model | 🌐 Project Page Overview Figure: Effectiveness of General-Reasoner trained with diverse verifiable reasoning questions using model-based verifier compared to baseline methods on various reasoning tasks. General-Reasoner is a training paradigm for large language models (LLMs), designed to robustly enhance reasoning abilities across… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/WebInstruct-verified.textquestion-answering100K<n<1M69 likes1.4k downloads10mo agoHugging Face14TIGER-Lab /VisualWebInstruct-verified 🧠 VisualWebInstruct-Verified: High-Confidence Multimodal QA for Reinforcement Learning VisualWebInstruct-Verified is a high-confidence subset of VisualWebInstruct, curated specifically for Reinforcement Learning (RL) and Reward Model training. It contains verified multimodal question–answer pairs where correctness, reasoning quality, and image–text alignment have been explicitly validated. This dataset is ideal for RLVR training pipelines. 📘 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct-verified.imagequestion-answering10K<n<100K7 likes1.1k downloads11mo agoHugging Face15TIGER-Lab /VisualWebInstruct-Seed Introduction This is the seed dataset we used to conduct Google Search. Links Github| Paper| Website Citation @article{visualwebinstruct, title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search}, author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu}, journal={arXiv preprint arXiv:2503.10582}, year={2025} } imagequestion-answering10K<n<100K19 likes1.1k downloads2y agoHugging Face16TIGER-Lab /VISTA-400K VISTA-400K This repo contains all subsets for VISTA-400K. VISTA is a video spatiotemporal augmentation method that generates long-duration and high-resolution video instruction-following data to enhance the video understanding capabilities of video LMMs. This repo is under construction. Please stay tuned. 🌐 Homepage | 📖 arXiv | 💻 GitHub | 🤗 VISTA-400K | 🤗 Models | 🤗 HRVideoBench Video Instruction Data Synthesis Pipeline VISTA leverages insights from… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VISTA-400K.textquestion-answering100K<n<1M6 likes1.1k downloads2y agoHugging Face17TIGER-Lab /Mantis-Eval Overview This is a newly curated dataset to evaluate multimodal language models' capability to reason over multiple images. More details are shown in https://tiger-ai-lab.github.io/Mantis/. Statistics This evaluation dataset contains 217 human-annotated challenging multi-image reasoning problems. Leaderboard We list the current results as follows: Models Size Mantis-Eval LLaVA OneVision 72B 77.60 LLaVA OneVision 7B 64.20 GPT-4V - 62.67… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/Mantis-Eval.imagequestion-answeringn<1K6 likes1k downloads2y agoHugging Face18zou-lab /BioMed-R1-Eval Disentangling Reasoning and Knowledge in Medical Large Language Models This is the evaluation dataset accompanying our paper, comprising 11 publicly available biomedical benchmarks. We disentangle each benchmark question into either medical reasoning or medical knowledge categories. Additionally, we provide a set of adversarial reasoning traces designed to evaluate the robustness of medical reasoning models. For more details, please refer to our GitHub. If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/zou-lab/BioMed-R1-Eval.tabularquestion-answering10K<n<100K1 likes1k downloads1y agoHugging Face19TIGER-Lab /ClawBench ClawBench Dataset ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. |💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website | 🚀 What's New [2026.05.12] Added the V2 corpus (130 newer tasks across 63 platforms) and 7 new models judged with… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/ClawBench.tabulartext-generationn<1K2 likes965 downloads4mo agoHugging Face20TIGER-Lab /SWE-QA-Pro-Bench SWE-QA-Pro Bench (A Repository-level QA Benchmark Built from Diverse Long-tail Repositories) 💻 GitHub | 📖 Paper | 🤗 SWE-QA-Pro 📢 News 🚀 [2026-5-19] The evaluation code is released on GitHub. 🔥 [2026-3-23] SWE-QA-Pro Bench is publicly released! The model and code will be released soon. Introduction SWE-QA-Pro Bench is a repository-level question answering dataset designed to evaluate whether models can perform grounded, agentic reasoning… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/SWE-QA-Pro-Bench.textquestion-answeringn<1K5 likes697 downloads4mo agoHugging Face21TIGER-Lab /AIME25The AIME25 part 1 exam from the website. textquestion-answeringn<1K2 likes541 downloads2y agoHugging Face22ASLP-lab /MSU-Benchmark MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto. Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie** Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/MSU-Benchmark.audioaudio-classification1K<n<10K1 likes541 downloads3mo agoHugging Face23humanlaya-data-lab /OneMillion-Bench $OneMillion-Bench A bilingual (Global/Chinese) realistic expert-level benchmark for evaluating language agents across 5 professional domains. The benchmark contains 400 entries with detailed, weighted rubric-based grading criteria designed for fine-grained evaluation of domain expertise, analytical reasoning, and instruction following. Dataset Structure Each subdirectory is a Hugging Face subset (configuration), and all data is in the test split. $OneMillion-Bench/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/humanlaya-data-lab/OneMillion-Bench.textquestion-answeringn<1K13 likes528 downloads7mo agoHugging Face24TIGER-Lab /ScreenSpot-ProSimplied Version of ScreenSpot-Pro dataset. imagequestion-answering1K<n<10K1 likes435 downloads9mo agoHugging Face25furonghuang-lab /PHTest🌟 PHTest: Evaluating False Refusals in LLMs 🤖 Auto Red-Teaming All prompts are generated automatically using a controllable text-generation technique called AutoDAN. 🌐 Diverse Prompts PHTest introduces false refusal patterns that aren’t present in existing datasets, including prompts that avoid mentioning sensitive words. ⚖️ Harmlessness & Controversial Labeling Controversial prompts are separately labeled to address the… See the full description on the dataset page: https://huggingface.co/datasets/furonghuang-lab/PHTest.texttext-generation1K<n<10K3 likes434 downloads5mo agoHugging Face26twinkle-ai /tw-drug-labels-vision Dataset Card for tw-drug-labels-vision 💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。 Dataset Details Dataset Description 本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段: 下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。 頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。 OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.imageimage-to-text10K<n<100K4 likes430 downloads5mo agoHugging Face27TIGER-Lab /MEGA-Bench MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks [ICLR 2025] 🌐 Homepage | 🏆 Leaderboard | 🤗 Dataset | 🤗 Paper | 🔎 Visualiaztion | 📖 arXiv | GitHub 🔔 News [2025-01]: Paper accepted by ICLR 2025. [2024-10-18]: Initial release of the evaluation code on our Github repo. [2024-10-14]: Paper released on arXiv. ❗❗ Data Information We put the file path of images/videos in HF datasets. Please download the zipped data here. We chose… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MEGA-Bench.imagequestion-answering1K<n<10K23 likes420 downloads1y agoHugging Face28li-lab /HLE-BioMedX HLE-BioMedX — Multilingual HLE Biology/Medicine A multilingual version of the Biology/Medicine subset of Humanity's Last Exam (HLE), released as one subset per language. Source benchmark: Humanity's Last Exam, dataset cais/hle. Subsets Group Languages How the target-language text was produced Source en Original English questions and answers. Machine-translated and expert-verified / revised zh, ja, ko, fr, th Machine translation reviewed by a human… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HLE-BioMedX.textquestion-answering1K<n<10K0 likes396 downloads28d agoHugging Face29jngb-labs /InvoiceBenchmark InvoiceBenchmark 200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number. The Pitch Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.documentquestion-answeringn<1K0 likes379 downloads5mo agoHugging Face30Eureka-Lab /PHYBench PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models [🌐 Project] [📄 Paper] [💻 Code] [🏆 Leaderboard] [🌟 Overview] [🔧 Data Details] [🚩 Citation] New Updates 2025.4.25: We release our code of EED Score. View and star on our github page! 2025.5.15: We have significantly improved the paper and experiments, including diversified experimental discussions and in-depth error analysis. The updated website is now live at… See the full description on the dataset page: https://huggingface.co/datasets/Eureka-Lab/PHYBench.textquestion-answering1K<n<10K62 likes373 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.