CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01initiacms /XLRS-Bench_visual_grounding_en 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench_visual_grounding_en.image10K<n<100K0 likes37k downloads11mo agoHugging Face02ProgramComputer /avspeech-visual-audio AVSpeech Video + Audio This repository is a media-bearing reconstruction of the public AVSpeech annotations. Each row represents an already-trimmed segment and keeps the original source-video timing and target-face-center metadata. Dataset structure clip_id: identifier derived as {youtube_id}_{start_sec:.3f}_{end_sec:.3f}. avspeech_metadata: JSON containing youtube_id, start_sec, end_sec, x_center, and y_center from the AVSpeech annotation. video: video-only… See the full description on the dataset page: https://huggingface.co/datasets/ProgramComputer/avspeech-visual-audio.audio1M<n<10M5 likes18k downloads1mo agoHugging Face03Beakerman0101 /trex-visualizer T-Rex Dataset Visualizer A browseable subset of the T-Rex dataset — Tactile-Rich Bimanual Dexterous Manipulation — collected on a bimanual Dexmate Vega-1 robot equipped with two Sharpa Wave dexterous hands. This visualizer subset contains 3,838 short trajectory clips drawn from the full 100-hour T-Rex collection, organized by (verb, object, hand) so you can quickly inspect coverage across motion primitives and object categories. For the full dataset (multi-view RGB, robot… See the full description on the dataset page: https://huggingface.co/datasets/Beakerman0101/trex-visualizer.textrobotics1K<n<10K0 likes12k downloads3mo agoHugging Face04institutional /institutional-books-hl-visual-elementsgated 📚 Institutional Books: Harvard Library — Visual Elements 22 million visual elements extracted from the volumes that comprise the Institutional Books: Harvard Library dataset. 22,622,060 visual elements extracted from 983,004 volumes 766,992,447 o200k_base tokens in AI-generated captions 6 high-level classes of visual elements organized in splits 5 processing steps: Detection, Classification, Deduplication, Captioning, and Rotation The Institutional Data Initiative at Harvard… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-visual-elements.image10M<n<100M1 likes8.2k downloads1mo agoHugging Face05xiuhuywh /DRIM-VisualReasonHardThis repository contains the RL training datasets used in the paper Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images image10K<n<100K134 likes4.4k downloads9mo agoHugging Face06joujiboi /Galgame-VisualNovel-Reupload Galgame VisualNovel Reupload This repository is a reupload of the visual novel dataset OOPPEENN/56697375616C4E6F76656C5F44617461736574. The goal of this reupload is to restructure the data for easier and more efficient use with the datasets library, instead of having to manually extract each archive file and parse json files of the original dataset. Loading the entire dataset To load and stream all voice lines from all games combined, simply load the train split. The… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/Galgame-VisualNovel-Reupload.audioautomatic-speech-recognition1M<n<10M38 likes4.2k downloads1y agoHugging Face07visual-layer /imagenet-1k-vl-enriched Visualize on Visual Layer Imagenet-1K-VL-Enriched An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues! With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering. The label issues helps to curate a cleaner and leaner dataset. Description The dataset consists of 6 columns: image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.imageobject-detection1M<n<10M40 likes4.1k downloads2y agoHugging Face08facebook /cyberseceval3-visual-prompt-injection Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark Dataset Details Dataset Description This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains. Language(s): English License: MIT Dataset Sources Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.imagetext-generation1K<n<10K10 likes2.7k downloads2y agoHugging Face09TIGER-Lab /VisualWebInstruct VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search VisualWebInstruct is a large-scale, diverse multimodal instruction dataset designed to enhance vision-language models' reasoning capabilities. The dataset contains approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs associated with 163,743 unique images, while the remaining 60% are text-only QA pairs. Please also checkout our more recent verified version at Huggingface.… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct.imagequestion-answering1M<n<10M44 likes2.3k downloads8mo agoHugging Face10TIGER-Lab /VisualWebInstruct-Recall Introduction This is the dataset recalled from Google Search from the seed images. Links Github| Paper| Website Citation @article{visualwebinstruct, title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search}, author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu}, journal={arXiv preprint arXiv:2503.10582}, year={2025} } imagequestion-answering100K<n<1M4 likes2k downloads2y agoHugging Face11initiacms /XLRS-Bench_visual_grounding_zh 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench_visual_grounding_zh.image10K<n<100K0 likes1.7k downloads11mo agoHugging Face12AdamYao /3D_Visual_Illusion_Depth_Estimation 3D Visual Illusion Depth Estimation Dataset Dataset Summary The 3D Visual Illusion Depth Estimation Dataset is designed for research on stereo and monocular depth estimation in 3D visual illusion scenes.It contains left and right stereo images, depth maps estimated from DepthAnything V2, and illusion-region masks. Dataset Structure Each sample in the dataset includes: left: Left-view RGB image right: Right-view RGB image depth: Monocularly estimated depth… See the full description on the dataset page: https://huggingface.co/datasets/AdamYao/3D_Visual_Illusion_Depth_Estimation.image100K<n<1M2 likes1.4k downloads10mo agoHugging Face13VisualCloze /Graph200K VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning [Paper]   [Project Page]   [Github] [🤗 Online Demo] [🤗 Full Model Card (Diffusers)]   [🤗 LoRA Model Card (Diffusers)] Graph200k is a large-scale dataset containing a wide range of distinct tasks of image generation. If you find Graph200k is helpful, please consider to star ⭐ the Github Repo. Thanks! 📰 News [2025-5-15] 🤗🤗🤗 VisualCloze has been merged into the… See the full description on the dataset page: https://huggingface.co/datasets/VisualCloze/Graph200K.imageimage-to-image100K<n<1M19 likes1.3k downloads9mo agoHugging Face14AdaptLLM /food-visual-instructions Adapting Multimodal Large Language Models to Domains via Post-Training (EMNLP 2025) This repos contains the food visual instructions for post-training MLLMs in our paper: On Domain-Specific Post-Training for Multimodal Large Language Models. The main project page is: Adapt-MLLM-to-Domains Data Information Using our visual instruction synthesizer, we generate visual instruction tasks based on the image-caption pairs from extended Recipe1M+ dataset. These synthetic… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/food-visual-instructions.imagevisual-question-answering100K<n<1M3 likes1.2k downloads1y agoHugging Face15TIGER-Lab /VisualWebInstruct-verified 🧠 VisualWebInstruct-Verified: High-Confidence Multimodal QA for Reinforcement Learning VisualWebInstruct-Verified is a high-confidence subset of VisualWebInstruct, curated specifically for Reinforcement Learning (RL) and Reward Model training. It contains verified multimodal question–answer pairs where correctness, reasoning quality, and image–text alignment have been explicitly validated. This dataset is ideal for RLVR training pipelines. 📘 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct-verified.imagequestion-answering10K<n<100K7 likes1.2k downloads11mo agoHugging Face16neulab /VisualPuzzles VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge 🏠 Homepage | 📊 VisualPuzzles | 💻 Github | 📄 Arxiv | 📕 PDF | 🖥️ Zeno Model Output Overview VisualPuzzles is a multimodal benchmark specifically designed to evaluate reasoning abilitiesin large models while deliberately minimizing reliance on domain-specific knowledge. Key features: 1168 diverse puzzles 5 reasoning categories: Algorithmic, Analogical, Deductive, Inductive, Spatial… See the full description on the dataset page: https://huggingface.co/datasets/neulab/VisualPuzzles.imagevisual-question-answering1K<n<10K12 likes1.1k downloads6mo agoHugging Face17UWGZQ /Synthetic_Visual_Genome2 Synthetic Visual Genome 2 (SVG2) A large-scale panoptic video scene graph dataset containing object labels, attributes, relationships, and instance-level segmentation masks. Paper: Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos Website: Synthetic Visual Genome 2 Versions cleaned The cleaned version has two sources: PVD (~593K videos) and SA-V (~43K videos). SAM-3 outputs We also provide the instance masks and… See the full description on the dataset page: https://huggingface.co/datasets/UWGZQ/Synthetic_Visual_Genome2.textvideo-classification100K<n<1M8 likes1.1k downloads7d agoHugging Face18TIGER-Lab /VisualWebInstruct-Seed Introduction This is the seed dataset we used to conduct Google Search. Links Github| Paper| Website Citation @article{visualwebinstruct, title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search}, author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu}, journal={arXiv preprint arXiv:2503.10582}, year={2025} } imagequestion-answering10K<n<100K19 likes1.1k downloads2y agoHugging Face19VisualSphinx /VisualSphinx-V1-Raw 🦁 VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL VisualSphinx is the largest fully-synthetic open-source dataset providing vision logic puzzles. It consists of over 660K automatically generated logical visual puzzles. Each logical puzzle is grounded with an interpretable rule and accompanied by both correct answers and plausible distractors. 🌐 Project Website - Learn more about VisualSphinx 📖 Technical Report - Discover the methodology and technical details… See the full description on the dataset page: https://huggingface.co/datasets/VisualSphinx/VisualSphinx-V1-Raw.imageimage-text-to-text100K<n<1M7 likes934 downloads1y agoHugging Face20visualwebbench /VisualWebBench VisualWebBench Dataset for the paper: VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding? 🌐 Homepage | 🐍 GitHub | 📖 arXiv Introduction We introduce VisualWebBench, a multimodal benchmark designed to assess the understanding and grounding capabilities of MLLMs in web scenarios. VisualWebBench consists of seven tasks, and comprises 1.5K human-curated instances from 139 real websites, covering 87 sub-domains. We evaluate 14… See the full description on the dataset page: https://huggingface.co/datasets/visualwebbench/VisualWebBench.imageimage-to-text1K<n<10K18 likes839 downloads2y agoHugging Face21multimodal-reasoning-lab /Visual-Searchimage10K<n<100K5 likes756 downloads1y agoHugging Face22Mini-o3 /VisualProbe_Hardimagen<1K1 likes712 downloads1y agoHugging Face23visual-layer /oxford-iiit-pet-vl-enriched Visualize on Visual Layer Oxford-IIIT-Pets-VL-Enriched An enriched version of the Oxford IIIT Pets Dataset with image caption, bounding boxes, and label issues! With this additional information, the Oxford IIIT Pet dataset can be extended to various tasks such as image retrieval or visual question answering. The label issues help to curate a cleaner and leaner dataset. Description The dataset consists of 6 columns: image_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/oxford-iiit-pet-vl-enriched.imageimage-classification1K<n<10K9 likes685 downloads2y agoHugging Face24Arasaaf /visual_26k_reasoningimage10K<n<100K1 likes684 downloads5mo agoHugging Face25Mini-o3 /VisualProbe_Easyimagen<1K1 likes663 downloads1y agoHugging Face26Mini-o3 /VisualProbe_Mediumimagen<1K1 likes617 downloads1y agoHugging Face27ScaleAI /VisualToolBench VisToolBench Dataset A benchmark dataset for evaluating vision-language models on tool-use tasks. Dataset Statistics Total samples: 1204 Single-turn: 603 Multi-turn: 601 Schema Column Type Description id string Unique task identifier turncase string Either "single-turn" or "multi-turn" num_turns int Number of conversation turns (1 for single-turn) prompt_category string Task category (e.g., "medical", "scientific", "general") eval_focus… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/VisualToolBench.imagevisual-question-answering1K<n<10K7 likes614 downloads9mo agoHugging Face28Reza2kn /visualears-persian-asr-16k 🗂️ visualears-persian-asr-16k English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Main public Persian ASR audio/text dataset: 3.93M 16 kHz rows. مجموعه‌دادهٔ اصلی و عمومی شنوا برای آموزش بازشناسی گفتار فارسی؛ شامل صوت ۱۶ کیلوهرتز، متن و فرادادهٔ منشأ در مقیاس چندمیلیونی. 🧩 Role flagship training corpus پیکرهٔ اصلی آموزش 📦 Snapshot 176 files; approximately 512.17 GB 176… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-persian-asr-16k.audio1M<n<10M1 likes555 downloads2mo agoHugging Face29open-paws /visual-qa-llama-format Open Paws Visual Qa Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Multimodal Data Format: JSONL (JSON Lines) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/visual-qa-llama-format.imagetext-generation1M<n<10M1 likes489 downloads1y agoHugging Face30MapEval /MapEval-Visual MapEval-Visual This dataset was introduced in MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models Example Query I am presently visiting Mount Royal Park . Could you please inform me about the nearby historical landmark? Options Circle Stone Secret pool Maison William Caldwell Cottingham Poste de cavalerie du Service de police de la Ville de Montreal Correct Option Circle Stone Prerequisite Download… See the full description on the dataset page: https://huggingface.co/datasets/MapEval/MapEval-Visual.imagemultiple-choicen<1K2 likes479 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.