CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01initiacms /XLRS-Bench_visual_grounding_en 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench_visual_grounding_en.image10K<n<100K0 likes37k downloads11mo agoHugging Face02institutional /institutional-books-hl-visual-elementsgated 📚 Institutional Books: Harvard Library — Visual Elements 22 million visual elements extracted from the volumes that comprise the Institutional Books: Harvard Library dataset. 22,622,060 visual elements extracted from 983,004 volumes 766,992,447 o200k_base tokens in AI-generated captions 6 high-level classes of visual elements organized in splits 5 processing steps: Detection, Classification, Deduplication, Captioning, and Rotation The Institutional Data Initiative at Harvard… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-visual-elements.image10M<n<100M1 likes8.2k downloads1mo agoHugging Face03konpat /visual-jenga-datasets Visual Jenga Datasets This directory contains the original datasets for Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting. Visual Jenga is a novel scene understanding task that involves progressively removing objects from a single image one at a time while keeping the rest of the scene stable. This process reveals object dependencies and provides a new way to evaluate grounded scene understanding by systematically exploring which objects can be removed… See the full description on the dataset page: https://huggingface.co/datasets/konpat/visual-jenga-datasets.imageimage-segmentation1K<n<10K0 likes5k downloads10mo agoHugging Face04xiuhuywh /DRIM-VisualReasonHardThis repository contains the RL training datasets used in the paper Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images image10K<n<100K134 likes4.4k downloads9mo agoHugging Face05visual-layer /imagenet-1k-vl-enriched Visualize on Visual Layer Imagenet-1K-VL-Enriched An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues! With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering. The label issues helps to curate a cleaner and leaner dataset. Description The dataset consists of 6 columns: image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.imageobject-detection1M<n<10M40 likes4.1k downloads2y agoHugging Face06Kyunnilee /visual-puzzlesimagen<1K1 likes3.8k downloads1y agoHugging Face07facebook /cyberseceval3-visual-prompt-injection Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark Dataset Details Dataset Description This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains. Language(s): English License: MIT Dataset Sources Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.imagetext-generation1K<n<10K10 likes2.7k downloads2y agoHugging Face08AhsanBB /Maritime_Visual_Tracking_Dataset_MVTD MVTD: Maritime Visual Tracking Dataset Overview MVTD (Maritime Visual Tracking Dataset) is a large-scale benchmark dataset designed specifically for single-object visual tracking (VOT) in maritime environments.It addresses challenges unique to maritime scenes: such as water reflections, low-contrast objects, dynamic backgrounds, scale variation, and severe illumination changes—which are not adequately covered by generic tracking datasets. The dataset contains 182… See the full description on the dataset page: https://huggingface.co/datasets/AhsanBB/Maritime_Visual_Tracking_Dataset_MVTD.image100K<n<1M2 likes2.4k downloads23h agoHugging Face09TIGER-Lab /VisualWebInstruct VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search VisualWebInstruct is a large-scale, diverse multimodal instruction dataset designed to enhance vision-language models' reasoning capabilities. The dataset contains approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs associated with 163,743 unique images, while the remaining 60% are text-only QA pairs. Please also checkout our more recent verified version at Huggingface.… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct.imagequestion-answering1M<n<10M44 likes2.3k downloads8mo agoHugging Face10TIGER-Lab /VisualWebInstruct-Recall Introduction This is the dataset recalled from Google Search from the seed images. Links Github| Paper| Website Citation @article{visualwebinstruct, title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search}, author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu}, journal={arXiv preprint arXiv:2503.10582}, year={2025} } imagequestion-answering100K<n<1M4 likes2k downloads2y agoHugging Face11initiacms /XLRS-Bench_visual_grounding_zh 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench_visual_grounding_zh.image10K<n<100K0 likes1.7k downloads11mo agoHugging Face12tsakman23 /visual_masked_distracting_metaworld Visual Masked Distracting Meta-World (ground-truth masks) Author: Georgios Tsakoumakis Thesis: Interaction-Masked Latent Action Models for Object-Aware Manipulation under Visual Distractors (MSc, Imperial College London) Expert Meta-World manipulation trajectories rendered with dynamic video-background distractors, augmented with ground-truth segmentation masks and pose for the manipulated object: the agent mask plus two per-frame fields, object_mask and object_state. All… See the full description on the dataset page: https://huggingface.co/datasets/tsakman23/visual_masked_distracting_metaworld.imagerobotics1M<n<10M1 likes1.7k downloads14d agoHugging Face13BoyangZ /VisualGenome_VG_100K_1_and_2 image0 likes1.6k downloads2y agoHugging Face14Mini-o3 /VisualProbe_trainimage1K<n<10K3 likes1.6k downloads1y agoHugging Face15songyiren /visual-reasoning-benchmark-results Visual Reasoning Benchmark Suite v3.3 · 2005 Tasks · 12 Tracks Equal Weight 本版本以用户最新上传的 visual_reasoning_benchmark_suite_v3_修改 为唯一基础版本,不回退、不覆盖用户已经重绘或修改过的既有数据。完整性比对结果:原基础包中 3283 个既有数据文件全部保持字节级不变。 在此基础上新增并整合: Nonogram(数织)150 题:45 Easy / 60 Medium / 45 Hard; Tangram(七巧板)150 题:45 Easy / 60 Medium / 45 Hard; 两个任务的一键生成器、统一生成入口、统一评估入口、雷达图和排行榜支持。 最终总规模:2005 题,12 个 Track。 任务与数量 Task Count figure_completion 394 spatial_generation 56 maze_beginner 64… See the full description on the dataset page: https://huggingface.co/datasets/songyiren/visual-reasoning-benchmark-results.image2 likes1.4k downloads2mo agoHugging Face16Evanwu50020 /visualization-chartsimagen<1K0 likes1.4k downloads1y agoHugging Face17Voxel51 /VisualOverload Dataset Card for VisualOverload This is a FiftyOne dataset with 2,720 samples. It is a FiftyOne-format conversion of the original paulgavrikov/visualoverload dataset (CVPR 2026). All credit for the data, annotations, and benchmark design belongs to the original authors — please see Citation and Dataset Sources. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/VisualOverload.imagevisual-question-answeringn<1K0 likes1.4k downloads3mo agoHugging Face18AdamYao /3D_Visual_Illusion_Depth_Estimation 3D Visual Illusion Depth Estimation Dataset Dataset Summary The 3D Visual Illusion Depth Estimation Dataset is designed for research on stereo and monocular depth estimation in 3D visual illusion scenes.It contains left and right stereo images, depth maps estimated from DepthAnything V2, and illusion-region masks. Dataset Structure Each sample in the dataset includes: left: Left-view RGB image right: Right-view RGB image depth: Monocularly estimated depth… See the full description on the dataset page: https://huggingface.co/datasets/AdamYao/3D_Visual_Illusion_Depth_Estimation.image100K<n<1M2 likes1.4k downloads10mo agoHugging Face19Voxel51 /visual_ai_at_neurips2025 Dataset Card for neurips-2025-vision-papers This is a FiftyOne dataset with 1134 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/visual_ai_at_neurips2025") # Launch the App session = fo.launch_app(dataset) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/visual_ai_at_neurips2025.imageimage-classification1K<n<10K1 likes1.3k downloads11mo agoHugging Face20VisualCloze /Graph200K VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning [Paper]   [Project Page]   [Github] [🤗 Online Demo] [🤗 Full Model Card (Diffusers)]   [🤗 LoRA Model Card (Diffusers)] Graph200k is a large-scale dataset containing a wide range of distinct tasks of image generation. If you find Graph200k is helpful, please consider to star ⭐ the Github Repo. Thanks! 📰 News [2025-5-15] 🤗🤗🤗 VisualCloze has been merged into the… See the full description on the dataset page: https://huggingface.co/datasets/VisualCloze/Graph200K.imageimage-to-image100K<n<1M19 likes1.3k downloads9mo agoHugging Face21AdaptLLM /food-visual-instructions Adapting Multimodal Large Language Models to Domains via Post-Training (EMNLP 2025) This repos contains the food visual instructions for post-training MLLMs in our paper: On Domain-Specific Post-Training for Multimodal Large Language Models. The main project page is: Adapt-MLLM-to-Domains Data Information Using our visual instruction synthesizer, we generate visual instruction tasks based on the image-caption pairs from extended Recipe1M+ dataset. These synthetic… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/food-visual-instructions.imagevisual-question-answering100K<n<1M3 likes1.2k downloads1y agoHugging Face22TIGER-Lab /VisualWebInstruct-verified 🧠 VisualWebInstruct-Verified: High-Confidence Multimodal QA for Reinforcement Learning VisualWebInstruct-Verified is a high-confidence subset of VisualWebInstruct, curated specifically for Reinforcement Learning (RL) and Reward Model training. It contains verified multimodal question–answer pairs where correctness, reasoning quality, and image–text alignment have been explicitly validated. This dataset is ideal for RLVR training pipelines. 📘 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct-verified.imagequestion-answering10K<n<100K7 likes1.2k downloads11mo agoHugging Face23neulab /VisualPuzzles VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge 🏠 Homepage | 📊 VisualPuzzles | 💻 Github | 📄 Arxiv | 📕 PDF | 🖥️ Zeno Model Output Overview VisualPuzzles is a multimodal benchmark specifically designed to evaluate reasoning abilitiesin large models while deliberately minimizing reliance on domain-specific knowledge. Key features: 1168 diverse puzzles 5 reasoning categories: Algorithmic, Analogical, Deductive, Inductive, Spatial… See the full description on the dataset page: https://huggingface.co/datasets/neulab/VisualPuzzles.imagevisual-question-answering1K<n<10K12 likes1.1k downloads6mo agoHugging Face24TIGER-Lab /VisualWebInstruct-Seed Introduction This is the seed dataset we used to conduct Google Search. Links Github| Paper| Website Citation @article{visualwebinstruct, title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search}, author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu}, journal={arXiv preprint arXiv:2503.10582}, year={2025} } imagequestion-answering10K<n<100K19 likes1.1k downloads2y agoHugging Face25VisualSphinx /VisualSphinx-V1-Raw 🦁 VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL VisualSphinx is the largest fully-synthetic open-source dataset providing vision logic puzzles. It consists of over 660K automatically generated logical visual puzzles. Each logical puzzle is grounded with an interpretable rule and accompanied by both correct answers and plausible distractors. 🌐 Project Website - Learn more about VisualSphinx 📖 Technical Report - Discover the methodology and technical details… See the full description on the dataset page: https://huggingface.co/datasets/VisualSphinx/VisualSphinx-V1-Raw.imageimage-text-to-text100K<n<1M7 likes934 downloads1y agoHugging Face26allenai /WildDet3D-visualization-source WildDet3D Visualization Data This repository hosts the visualization data for the WildDet3D-Bench benchmark — a human-annotated evaluation set for monocular 3D object detection in the wild. Dataset Overview WildDet3D-Bench is a validation set of 2,470 images drawn from three source datasets, with 9,256 human-verified 3D bounding box annotations across 2,196 images. Source Images Description COCO Val 424 MS-COCO 2017 validation LVIS Train 1,113 LVIS v1.0 (COCO… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildDet3D-visualization-source.imageobject-detection6 likes893 downloads6mo agoHugging Face27EpicPinkPenguin /visual_distracting_metaworld Visual Distracting Meta-World with Agent Masks Visual Distracting Meta-World with Agent Masks contains successful expert trajectories for all 50 Meta-World MT50 v3 manipulation tasks. Every visual observation includes agent masks: a ground-truth robot-arm mask obtained from the simulator, a SAM 2.1 mask predicted from the clean observation, and a separate SAM 2.1 mask predicted from the distracted observation. Each step pairs these masks with the same underlying simulator state… See the full description on the dataset page: https://huggingface.co/datasets/EpicPinkPenguin/visual_distracting_metaworld.imagereinforcement-learning10M<n<100M0 likes889 downloads18d agoHugging Face28visualwebbench /VisualWebBench VisualWebBench Dataset for the paper: VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding? 🌐 Homepage | 🐍 GitHub | 📖 arXiv Introduction We introduce VisualWebBench, a multimodal benchmark designed to assess the understanding and grounding capabilities of MLLMs in web scenarios. VisualWebBench consists of seven tasks, and comprises 1.5K human-curated instances from 139 real websites, covering 87 sub-domains. We evaluate 14… See the full description on the dataset page: https://huggingface.co/datasets/visualwebbench/VisualWebBench.imageimage-to-text1K<n<10K18 likes839 downloads2y agoHugging Face29Silicon23 /itw_pipeline_visualization ITW Pipeline Visualization — assets Purpose of this repository This repository exists for one reason: to serve media files to the visualization site at https://silicon23.github.io/itw_pipeline_visualization/. A static site cannot host its own heavy media, so the frames, depth maps and overlay renders it streams live here. It is not published for redistribution, and it is not a dataset to train or evaluate on. It is the asset backing of a figure — the equivalent… See the full description on the dataset page: https://huggingface.co/datasets/Silicon23/itw_pipeline_visualization.image10K<n<100K0 likes777 downloads14d agoHugging Face30multimodal-reasoning-lab /Visual-Searchimage10K<n<100K5 likes756 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.