CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01institutional /institutional-books-hl-visual-elementsgated 📚 Institutional Books: Harvard Library — Visual Elements 22 million visual elements extracted from the volumes that comprise the Institutional Books: Harvard Library dataset. 22,622,060 visual elements extracted from 983,004 volumes 766,992,447 o200k_base tokens in AI-generated captions 6 high-level classes of visual elements organized in splits 5 processing steps: Detection, Classification, Deduplication, Captioning, and Rotation The Institutional Data Initiative at Harvard… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-visual-elements.image10M<n<100M1 likes8.2k downloads1mo agoHugging Face02visual-layer /imagenet-1k-vl-enriched Visualize on Visual Layer Imagenet-1K-VL-Enriched An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues! With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering. The label issues helps to curate a cleaner and leaner dataset. Description The dataset consists of 6 columns: image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.imageobject-detection1M<n<10M40 likes4.5k downloads2y agoHugging Face03TIGER-Lab /VisualWebInstruct VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search VisualWebInstruct is a large-scale, diverse multimodal instruction dataset designed to enhance vision-language models' reasoning capabilities. The dataset contains approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs associated with 163,743 unique images, while the remaining 60% are text-only QA pairs. Please also checkout our more recent verified version at Huggingface.… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct.imagequestion-answering1M<n<10M44 likes2.3k downloads8mo agoHugging Face04TIGER-Lab /VisualWebInstruct-Recall Introduction This is the dataset recalled from Google Search from the seed images. Links Github| Paper| Website Citation @article{visualwebinstruct, title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search}, author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu}, journal={arXiv preprint arXiv:2503.10582}, year={2025} } imagequestion-answering100K<n<1M4 likes1.7k downloads2y agoHugging Face05tsakman23 /visual_masked_distracting_metaworld Visual Masked Distracting Meta-World (ground-truth masks) Author: Georgios Tsakoumakis Thesis: Interaction-Masked Latent Action Models for Object-Aware Manipulation under Visual Distractors (MSc, Imperial College London) Expert Meta-World manipulation trajectories rendered with dynamic video-background distractors, augmented with ground-truth segmentation masks and pose for the manipulated object: the agent mask plus two per-frame fields, object_mask and object_state. All… See the full description on the dataset page: https://huggingface.co/datasets/tsakman23/visual_masked_distracting_metaworld.imagerobotics1M<n<10M1 likes1.7k downloads18d agoHugging Face06neulab /VisualPuzzles VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge 🏠 Homepage | 📊 VisualPuzzles | 💻 Github | 📄 Arxiv | 📕 PDF | 🖥️ Zeno Model Output Overview VisualPuzzles is a multimodal benchmark specifically designed to evaluate reasoning abilitiesin large models while deliberately minimizing reliance on domain-specific knowledge. Key features: 1168 diverse puzzles 5 reasoning categories: Algorithmic, Analogical, Deductive, Inductive, Spatial… See the full description on the dataset page: https://huggingface.co/datasets/neulab/VisualPuzzles.imagevisual-question-answering1K<n<10K12 likes1.1k downloads6mo agoHugging Face07EpicPinkPenguin /visual_distracting_metaworld Visual Distracting Meta-World with Agent Masks Visual Distracting Meta-World with Agent Masks contains successful expert trajectories for all 50 Meta-World MT50 v3 manipulation tasks. Every visual observation includes agent masks: a ground-truth robot-arm mask obtained from the simulator, a SAM 2.1 mask predicted from the clean observation, and a separate SAM 2.1 mask predicted from the distracted observation. Each step pairs these masks with the same underlying simulator state… See the full description on the dataset page: https://huggingface.co/datasets/EpicPinkPenguin/visual_distracting_metaworld.imagereinforcement-learning10M<n<100M0 likes1.1k downloads22d agoHugging Face08TIGER-Lab /VisualWebInstruct-Seed Introduction This is the seed dataset we used to conduct Google Search. Links Github| Paper| Website Citation @article{visualwebinstruct, title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search}, author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu}, journal={arXiv preprint arXiv:2503.10582}, year={2025} } imagequestion-answering10K<n<100K19 likes1.1k downloads2y agoHugging Face09VisualSphinx /VisualSphinx-V1-Raw 🦁 VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL VisualSphinx is the largest fully-synthetic open-source dataset providing vision logic puzzles. It consists of over 660K automatically generated logical visual puzzles. Each logical puzzle is grounded with an interpretable rule and accompanied by both correct answers and plausible distractors. 🌐 Project Website - Learn more about VisualSphinx 📖 Technical Report - Discover the methodology and technical details… See the full description on the dataset page: https://huggingface.co/datasets/VisualSphinx/VisualSphinx-V1-Raw.imageimage-text-to-text100K<n<1M7 likes1.1k downloads1y agoHugging Face10multimodal-reasoning-lab /Visual-Searchimage10K<n<100K5 likes846 downloads1y agoHugging Face11tsakman23 /visual_masked_distracting_metaworld_sam Visual Masked Distracting Meta-World (ground-truth + SAM masks) Author: Georgios Tsakoumakis Thesis: Interaction-Masked Latent Action Models for Object-Aware Manipulation under Visual Distractors (MSc, Imperial College London) Expert Meta-World manipulation trajectories rendered with dynamic video-background distractors, carrying both ground-truth and SAM-predicted segmentation masks for the agent and the manipulated object, plus the manipulated object's pose. This is the SAM… See the full description on the dataset page: https://huggingface.co/datasets/tsakman23/visual_masked_distracting_metaworld_sam.imagerobotics1M<n<10M1 likes807 downloads18d agoHugging Face12visual-layer /oxford-iiit-pet-vl-enriched Visualize on Visual Layer Oxford-IIIT-Pets-VL-Enriched An enriched version of the Oxford IIIT Pets Dataset with image caption, bounding boxes, and label issues! With this additional information, the Oxford IIIT Pet dataset can be extended to various tasks such as image retrieval or visual question answering. The label issues help to curate a cleaner and leaner dataset. Description The dataset consists of 6 columns: image_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/oxford-iiit-pet-vl-enriched.imageimage-classification1K<n<10K9 likes707 downloads2y agoHugging Face13Arasaaf /visual_26k_reasoningimage10K<n<100K1 likes681 downloads6mo agoHugging Face14visualwebbench /VisualWebBench VisualWebBench Dataset for the paper: VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding? 🌐 Homepage | 🐍 GitHub | 📖 arXiv Introduction We introduce VisualWebBench, a multimodal benchmark designed to assess the understanding and grounding capabilities of MLLMs in web scenarios. VisualWebBench consists of seven tasks, and comprises 1.5K human-curated instances from 139 real websites, covering 87 sub-domains. We evaluate 14… See the full description on the dataset page: https://huggingface.co/datasets/visualwebbench/VisualWebBench.imageimage-to-text1K<n<10K18 likes643 downloads2y agoHugging Face15ScaleAI /VisualToolBench VisToolBench Dataset A benchmark dataset for evaluating vision-language models on tool-use tasks. Dataset Statistics Total samples: 1204 Single-turn: 603 Multi-turn: 601 Schema Column Type Description id string Unique task identifier turncase string Either "single-turn" or "multi-turn" num_turns int Number of conversation turns (1 for single-turn) prompt_category string Task category (e.g., "medical", "scientific", "general") eval_focus… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/VisualToolBench.imagevisual-question-answering1K<n<10K7 likes594 downloads9mo agoHugging Face16open-paws /visual-qa-llama-format Open Paws Visual Qa Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Multimodal Data Format: JSONL (JSON Lines) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/visual-qa-llama-format.imagetext-generation1M<n<10M1 likes487 downloads1y agoHugging Face17hamza-adnan /visual_distracting_metaworld_with_masksimage10M<n<100M0 likes481 downloads5mo agoHugging Face18hamza-adnan /visual_distracting_control_suite Visual Distracting Control Suite Benchmark This dataset contains expert trajectories generated by a Proximal Policy Optimization (PPO) reinforcement learning agent trained on 4 environments of the Distracting Control Suite. For each environment we collect data with different levels of distraction, which we define below, and masks for the agent. Levels of distraction: None: Vanilla DeepMind Control Suite without visual distractions. The environment uses the default static background… See the full description on the dataset page: https://huggingface.co/datasets/hamza-adnan/visual_distracting_control_suite.image100M<n<1B0 likes475 downloads6mo agoHugging Face19tanhuajie2001 /spatial-visual-reasoning-66kimage10K<n<100K2 likes441 downloads2y agoHugging Face20hf-internal-testing /document-visual-retrieval-test Model Card: Document Visual Retrieval Test (internal) Dataset Overview This dataset is designed to evaluate the performance of visual retrievers by testing their ability to match a query to a relevant image. Each of the three examples in this dataset contains a text query and an associated image, which is a scanned page from the foundational "Attention is All You Need" paper. The purpose of this dataset is to facilitate the evaluation of visual retrievers, where the… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/document-visual-retrieval-test.imagen<1K1 likes430 downloads2y agoHugging Face21knowledge-in-visual-synthesis /v1 Knowledge in Visual Synthesis This dataset contains prompt–image examples for evaluating and studying knowledge-intensive visual synthesis. Samples are organized by contributor as dataset subsets (configs), with each upload version exposed as a split. Dataset structure Subset Splits byx v1, v2 yuner v1, v2 zanyi v1, v2, v3 jiayu v1, v2, v3 sherry v1, v2 yujunz v1 The byx/v1 split contains 140 unique prompts and 300 generated images. For… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-in-visual-synthesis/v1.image1K<n<10K0 likes386 downloads2mo agoHugging Face22visual-layer /coco-2014-vl-enriched Visualize on Visual Layer COCO-2014-VL-Enriched An enriched version of the COCO 2014 dataset with label issues! The label issues help to curate a cleaner and leaner dataset. Description The dataset consists of 6 columns: image_id: The original image filename from the COCO dataset. image: Image data in the form of PIL Image. label_bbox: Bounding box annotations from the COCO dataset. Consists of bounding box coordinates, confidence scores, and labels… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/coco-2014-vl-enriched.imageobject-detection100K<n<1M2 likes380 downloads2y agoHugging Face23multimodal-reasoning-lab /Visual-Jigsawimage10K<n<100K4 likes334 downloads1y agoHugging Face24AI-4-Everyone /Visual-TableQA 🧠 Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images Welcome to Visual-TableQA, a project designed to generate high-quality synthetic question-answer datasets associated to images of tables. This resource is ideal for training and evaluating models on visually-grounded table understanding tasks such as document QA, table parsing, and multimodal reasoning. 🚀 Latest Update We have refreshed the dataset with newly generated QA pairs created by… See the full description on the dataset page: https://huggingface.co/datasets/AI-4-Everyone/Visual-TableQA.imagetable-question-answering1K<n<10K12 likes333 downloads1y agoHugging Face25ljnlonoljpiljm /visual_genomeimage100K<n<1M0 likes327 downloads5mo agoHugging Face26minglingfeng /Ocean_R1_collected_visual_dataimage100K<n<1M3 likes324 downloads2y agoHugging Face27weikaih /SOC-Training-Data-Visualization Paper Link SOS: Synthetic Object Segments Improve Detection, Segmentation, and Grounding Code repo Code for Generation Citation @misc{huang2025sossyntheticobjectsegments, title={SOS: Synthetic Object Segments Improve Detection, Segmentation, and Grounding}, author={Weikai Huang and Jieyu Zhang and Taoyang Jia and Chenhao Zheng and Ziqi Gao and Jae Sung Park and Ranjay Krishna}, year={2025}, eprint={2510.09110}, archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/SOC-Training-Data-Visualization.imagen<1K0 likes308 downloads1y agoHugging Face28VisualSphinx /VisualSphinx-V1-Raw-Panelsimage100K<n<1M0 likes306 downloads11mo agoHugging Face29gowitheflow /ARO-Visual-Relationimage10K<n<100K1 likes295 downloads2y agoHugging Face30visual-memory /PersonaChat-Qwen-Image-2512-enhancedimage10K<n<100K0 likes275 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.