CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jat-project /jat-dataset JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.imagereinforcement-learning100M<n<1B78 likes385k downloads3y agoHugging Face02ShadenA /MathNet Quick Start · Overview · Tasks · Comparison · Dataset Stats · Data Sources · Pipeline · Schema · License · Citation This is the official MathNet v0. A larger version v1 will be uploaded soon (more countires, problems and richer metadata). Schema is stable but field values may be revised in v1. Quick start from datasets import load_dataset # Default: all problems ds = load_dataset("ShadenA/MathNet", split="train") # Or a specific country / competition-body config… See the full description on the dataset page: https://huggingface.co/datasets/ShadenA/MathNet.imagequestion-answering10K<n<100K96 likes94k downloads3mo agoHugging Face03MMMU /MMMU MMMU (A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI) 🌐 Homepage | 🏆 Leaderboard | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub 🔔News 🛠️[2026-07-10]: Fixed incorrect ground-truth answer labels in validation_Design_15 and validation_Art_Theory_4. 🛠️[2026-04-21]: Fixed option issue in test_Psychology_15. ‼️[2026-02-12]: We have released the answers for the test set! You can now evaluate your models on the test set… See the full description on the dataset page: https://huggingface.co/datasets/MMMU/MMMU.imagequestion-answering10K<n<100K336 likes76k downloads3mo agoHugging Face04Hothan /OlympiadBench OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems[ACL 2024] 📖 arXiv | GitHub Note: We have made adjustments to the image content in the multimodal portion of the dataset and fixed previous issues where some images in the English physics subset were not displayed properly. If your usage involves images, please re-download the dataset (we recommend all users to download the latest version). Additionally, some entries… See the full description on the dataset page: https://huggingface.co/datasets/Hothan/OlympiadBench.imagequestion-answering1K<n<10K47 likes43k downloads1y agoHugging Face05derek-thomas /ScienceQA Dataset Card Creation Guide Dataset Summary Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering Supported Tasks and Leaderboards Multi-modal Multiple Choice Languages English Dataset Structure Data Instances Explore more samples here. {'image': Image, 'question': 'Which of these states is farthest north?', 'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'], 'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/derek-thomas/ScienceQA.imagemultiple-choice10K<n<100K234 likes39k downloads4y agoHugging Face06futurehouse /lab-bench LAB-Bench The Language Agent Biology Benchmark, or LAB-Bench, is an evaluation dataset for AI systems intended to benchmark capabilities foundational to scientific research in biology. The dataset currently consists of 8 broad categories, comprising 30 narrower subtasks, including extracting information from the scientific literature (LitQA2), retrieving information from databases (DbQA) and supplementary information (SuppQA), reasoning about scientific figures (FigQA) and tables… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/lab-bench.imagequestion-answering1K<n<10K51 likes35k downloads1y agoHugging Face07AI4Math /MathVista Dataset Card for MathVista Dataset Description Paper Information Dataset Examples Leaderboard Dataset Usage Data Downloading Data Format Data Visualization Data Source Automatic Evaluation License Citation Dataset Description MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts. It consists of three newly created datasets, IQTest, FunctionQA, and PaperQA, which address the missing visual domains and are tailored to evaluate logical… See the full description on the dataset page: https://huggingface.co/datasets/AI4Math/MathVista.imagemultiple-choice1K<n<10K226 likes24k downloads3y agoHugging Face08MMMU /MMMU_Pro MMMU-Pro (A More Robust Multi-discipline Multimodal Understanding Benchmark) 🌐 Homepage | 🏆 Leaderboard | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub 🔔News 🛠️[2026-07-10] Fixed incorrect ground-truth answer labels. (validation_Design_15; validation_Art_Theory_4) 🛠️[2026-05-30] Fixed the option augmentation issue in Vision and Standard (10 options) settings. (validation_Diagnostics_and_Laboratory_Medicine_17) 🛠️[2025-03-08] Fixed mismatch between inner image… See the full description on the dataset page: https://huggingface.co/datasets/MMMU/MMMU_Pro.imagequestion-answering1K<n<10K66 likes20k downloads3mo agoHugging Face09moondream /megalith-mdqa Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM. imagequestion-answering1M<n<10M28 likes19k downloads1y agoHugging Face10Lin-Chen /MMStar MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?) 🌐 Homepage | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub Dataset Details As shown in the figure below, existing benchmarks lack consideration of the vision dependency of evaluation samples and potential data leakage from LLMs' and LVLMs' training data. Therefore, we introduce MMStar: an elite vision-indispensible multi-modal benchmark, aiming to ensure each curated sample exhibits… See the full description on the dataset page: https://huggingface.co/datasets/Lin-Chen/MMStar.imagemultiple-choice1K<n<10K53 likes18k downloads2y agoHugging Face11princeton-nlp /CharXiv CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs NeurIPS 2024 🏠Home (🚧Still in construction) | 🤗Data | 🥇Leaderboard | 🖥️Code | 📄Paper This repo contains the full dataset for our paper CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs, which is a diverse and challenging chart understanding benchmark fully curated by human experts. It includes 2,323 high-resolution charts manually sourced from arXiv preprints. Each chart is… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/CharXiv.imagevisual-question-answering1K<n<10K51 likes11k downloads2y agoHugging Face12user9000 /CLEVR-HOPE CLEVR-HOPE The CLEVR Held-Out Pair Evaluation (CLEVR-HOPE) dataset is a diagnostic dataset for testing the systematicity of VQA models. CLEVR-HOPE is a controlled setting to test whether VQA models generalize to pairs of attribute values that were not seen during either training or fine-tuning. Within CLEVR-HOPE, we refer to an unseen pair of attribute values as a Held-Out Pair (HOP). The dataset is composed of 29 sub-datasets, each for a different HOP. For each of the 29 HOPs, we… See the full description on the dataset page: https://huggingface.co/datasets/user9000/CLEVR-HOPE.imagequestion-answering10M<n<100M2 likes11k downloads1y agoHugging Face13MathLLMs /MathVision Measuring Multimodal Mathematical Reasoning with the MATH-Vision Dataset [💻 Github] [🌐 Homepage] [📊 Main Leaderboard ] [📊 Open Source Leaderboard ] [🌿 Wild Leaderboard ] [🔍 Visualization] [📖 Paper] 🌿 NEW: MATH-Vision-Wild MATH-Vision-Wild is a photographic, real-world variant of MATH-Vision. The same testmini problems are physically captured on printed paper, iPads, laptops, and projectors under varying lighting and angles — the conditions VLMs actually… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/MathVision.imagequestion-answering1K<n<10K173 likes9.8k downloads4mo agoHugging Face14sensenova /SenseNova-SI-8MEN | 中文 SenseNova-SI-8M 🚀 This is the official full-scale training dataset of the SenseNova-SI series. SenseNova-SI-8M contains ~8.16 million carefully curated training samples spanning ~2.72 million unique images, organized under a rigorous taxonomy of spatial capabilities. It is the dataset used to train the recommended released model SenseNova-SI-1.1-InternVL3-8B and serves as the canonical training corpus for spatial intelligence research… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/SenseNova-SI-8M.imagevisual-question-answering1M<n<10M23 likes8.7k downloads3mo agoHugging Face15racineai /VDR_MEGA_2 VDR_MEGA_2 Dataset Summary VDR_MEGA_2 is a high-quality multimodal dataset created through the merge of multiple domain-specific datasets with enhanced data processing techniques. This dataset represents our most refined approach to multimodal data generation, incorporating filtering algorithms and improved AI-assisted content generation to deliver superior quality for RAG, DSE, question answering, document search, and vision-language model training tasks.… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_2.imagequestion-answering1M<n<10M16 likes8.3k downloads10mo agoHugging Face16stanford-oval /ccnewsThis dataset is the result of processing all WARC files in the CCNews Corpus, from the beginning (2016) to June of 2024. The data has been cleaned and deduplicated, and language of articles have been detected and added. The process is similar to what HuggingFace's DataTrove does. Overall, it contains about 600 million news articles in more than 100 languages from all around the globe. For license information, please refer to CommonCrawl's Terms of Use. Sample Python code to explore this… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/ccnews.imagetext-classification100M<n<1B36 likes7.1k downloads2y agoHugging Face17JingkunAn /RefSpatial ⚠️ Warning: The Dataset Viewer and Data Studio above are for display only. They show just 800 samples from the full RefSpatial dataset, taken from the "SubsetVisualization" folder in Hugging Face ".parquet" format. ℹ️ Info: The full raw dataset (~357GB) is available in non-HF formats (e.g., images, depth maps, JSON files). RefSpatial: A Large-scale Dataset for teaching a general VLM to achieve spatial referring with reasoning… See the full description on the dataset page: https://huggingface.co/datasets/JingkunAn/RefSpatial.imagequestion-answeringn<1K24 likes6k downloads8mo agoHugging Face18yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M53 likes5.7k downloads7mo agoHugging Face19RUC-NLPIR /Omnimodal-Agent-SFT-2K OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.audioquestion-answering1K<n<10K9 likes5.2k downloads7mo agoHugging Face20bio-nlp-umass /MedThinkVQA MedThinkVQA MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning. Links GitHub: https://github.com/benluwang/MedThinkVQA Leaderboard: https://benluwang.github.io/MedThinkVQA/ Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.imagequestion-answering1K<n<10K11 likes4.8k downloads4mo agoHugging Face21letxbe /BoundingDocs BoundingDocs 🔍 The largest spatially-annotated dataset for Document Question Answering Dataset Description BoundingDocs is a unified dataset for Document Question Answering (QA) that includes spatial annotations. It consolidates multiple public datasets from Document AI and Visually Rich Document Understanding (VRDU) domains. The dataset reformulates Information Extraction (IE) tasks into QA tasks, making it a valuable resource for training and evaluating Large Language… See the full description on the dataset page: https://huggingface.co/datasets/letxbe/BoundingDocs.imagequestion-answering10K<n<100K21 likes4k downloads1y agoHugging Face22stanford-crfm /image2struct-latex-v1 Image2Struct - Latex Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo License: Apache License Version 2.0, January 2004 Dataset description Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images. This subdataset focuses on LaTeX code. The model is given an image of the expected output with the prompt: Please provide the LaTex code used to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/image2struct-latex-v1.imagequestion-answering1K<n<10K12 likes3.8k downloads2y agoHugging Face23OpenMOSS-Team /GameQA-140K [ICLR 2026] Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning 🎊 News [2026/07] 🔥Peking University and Kuaishou Kling Team evaluate their agentic visual reasoning method Beacon on our GameQA benchmark. Beacon learns when tools are truly needed (Mode Adaptiveness) and how tool use extends capability on hard problems (Tool Effect), and achieves the highest accuracy on GameQA among open-source models of the same scale… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-140K.imagequestion-answeringn<1K22 likes3.7k downloads5d agoHugging Face24sensenova /SenseNova-SI-800KEN | 中文 SenseNova-SI-800K 🔥Please check out our newly released SenseNova-SI-8M, official full-scale training dataset of the SenseNova-SI series. SenseNova-SI-8M contains ~8.16 million carefully curated training samples spanning ~2.72 million unique images, organized under a rigorous taxonomy of spatial capabilities.The SenseNova-SI-800K dataset provided here is a downsampled subset of SenseNova-SI-8M, specifically designed for studying scaling… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/SenseNova-SI-800K.imagevisual-question-answering100K<n<1M22 likes3.5k downloads4mo agoHugging Face25jablonkagroup /MaCBench MaCBench A Chemistry and Materials Benchmark for evaluating Vision Large Language Models ⚠️ IMPORTANT NOTICE - NOT FOR TRAINING 🚫 THIS DATASET IS STRICTLY FOR EVALUATION PURPOSES ONLY 🚫 DO NOT USE THIS DATASET FOR TRAINING OR FINE-TUNING MODELS This benchmark is designed exclusively for evaluation and testing of existing models. Using this data for training would compromise the integrity of the benchmark and invalidate evaluation results. Please… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/MaCBench.imagequestion-answering1K<n<10K11 likes3.5k downloads1y agoHugging Face26Salesforce /UniDoc-Bench UNIDOC-BENCH Dataset A unified benchmark for document-centric multimodal retrieval-augmented generation (MM-RAG). Dataset Description UNIDOC-BENCH is the first large-scale, realistic benchmark for multimodal retrieval-augmented generation (MM-RAG) and Visual Question Answering (VQA) built from 70,000 real-world PDF pages across eight domains. The dataset extracts and links evidence from text, tables, and figures, then generates 1,700+ multimodal QA pairs spanning… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/UniDoc-Bench.imagequestion-answering1K<n<10K15 likes3.3k downloads10mo agoHugging Face27JosselinSom /Latex-VLMimagequestion-answering1K<n<10K9 likes3.3k downloads3y agoHugging Face28AI4Math /MathVerse Dataset Card for MathVerse Dataset Description Paper Information Dataset Examples Leaderboard Citation Dataset Description The capabilities of Multi-modal Large Language Models (MLLMs) in visual math problem-solvingremain insufficiently evaluated and understood. We investigate current benchmarks to incorporate excessive visual content within textual questions, which potentially assist MLLMs in deducing answers without truly interpreting the input diagrams. To… See the full description on the dataset page: https://huggingface.co/datasets/AI4Math/MathVerse.imagemultiple-choice1K<n<10K72 likes3.2k downloads1y agoHugging Face29MUIRBENCH /MUIRBENCH MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding 🌐 Homepage | 📖 Paper | 💻 Evaluation Intro MuirBench is a benchmark containing 11,264 images and 2,600 multiple-choice questions, providing robust evaluation on 12 multi-image understanding tasks. MuirBench evaluates on a comprehensive range of 12 multi-image understanding abilities, e.g. geographic understanding, diagram understanding, visual retrieval, ..., etc, while prior benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/MUIRBENCH/MUIRBENCH.imagequestion-answering1K<n<10K17 likes3k downloads2y agoHugging Face30lytang /ChartMuseum [NeurIPS 2025] ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models Authors: Liyan Tang, Grace Kim, Xinyu Zhao, Thom Lake, Wenxuan Ding, Fangcong Yin, Prasann Singhal, Manya Wadhwa, Zeyu Leo Liu, Zayne Sprague, Ramya Namuduri, Bodun Hu, Juan Diego Rodriguez, Puyuan Peng, Greg Durrett Leaderboard 🥇 | Paper 📃 | Code 💻 Overview ChartMuseum is a chart question answering benchmark designed to evaluate reasoning capabilities of large… See the full description on the dataset page: https://huggingface.co/datasets/lytang/ChartMuseum.imagequestion-answering1K<n<10K7 likes3k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.