CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ProgramComputer /avspeech-visual-audio AVSpeech Video + Audio This repository is a media-bearing reconstruction of the public AVSpeech annotations. Each row represents an already-trimmed segment and keeps the original source-video timing and target-face-center metadata. Dataset structure clip_id: identifier derived as {youtube_id}_{start_sec:.3f}_{end_sec:.3f}. avspeech_metadata: JSON containing youtube_id, start_sec, end_sec, x_center, and y_center from the AVSpeech annotation. video: video-only… See the full description on the dataset page: https://huggingface.co/datasets/ProgramComputer/avspeech-visual-audio.audio1M<n<10M5 likes17k downloads1mo agoHugging Face02institutional /institutional-books-hl-visual-elementsgated 📚 Institutional Books: Harvard Library — Visual Elements 22 million visual elements extracted from the volumes that comprise the Institutional Books: Harvard Library dataset. 22,622,060 visual elements extracted from 983,004 volumes 766,992,447 o200k_base tokens in AI-generated captions 6 high-level classes of visual elements organized in splits 5 processing steps: Detection, Classification, Deduplication, Captioning, and Rotation The Institutional Data Initiative at Harvard… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-visual-elements.image10M<n<100M1 likes8.2k downloads1mo agoHugging Face03visual-layer /imagenet-1k-vl-enriched Visualize on Visual Layer Imagenet-1K-VL-Enriched An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues! With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering. The label issues helps to curate a cleaner and leaner dataset. Description The dataset consists of 6 columns: image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.imageobject-detection1M<n<10M40 likes4.5k downloads2y agoHugging Face04joujiboi /Galgame-VisualNovel-Reupload Galgame VisualNovel Reupload This repository is a reupload of the visual novel dataset OOPPEENN/56697375616C4E6F76656C5F44617461736574. The goal of this reupload is to restructure the data for easier and more efficient use with the datasets library, instead of having to manually extract each archive file and parse json files of the original dataset. Loading the entire dataset To load and stream all voice lines from all games combined, simply load the train split. The… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/Galgame-VisualNovel-Reupload.audioautomatic-speech-recognition1M<n<10M38 likes4k downloads1y agoHugging Face05TIGER-Lab /VisualWebInstruct VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search VisualWebInstruct is a large-scale, diverse multimodal instruction dataset designed to enhance vision-language models' reasoning capabilities. The dataset contains approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs associated with 163,743 unique images, while the remaining 60% are text-only QA pairs. Please also checkout our more recent verified version at Huggingface.… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct.imagequestion-answering1M<n<10M44 likes2.2k downloads8mo agoHugging Face06TIGER-Lab /VisualWebInstruct-Recall Introduction This is the dataset recalled from Google Search from the seed images. Links Github| Paper| Website Citation @article{visualwebinstruct, title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search}, author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu}, journal={arXiv preprint arXiv:2503.10582}, year={2025} } imagequestion-answering100K<n<1M4 likes1.8k downloads2y agoHugging Face07UWGZQ /Synthetic_Visual_Genome2 Synthetic Visual Genome 2 (SVG2) A large-scale panoptic video scene graph dataset containing object labels, attributes, relationships, and instance-level segmentation masks. Paper: Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos Website: Synthetic Visual Genome 2 Versions cleaned The cleaned version has two sources: PVD (~593K videos) and SA-V (~43K videos). SAM-3 outputs We also provide the instance masks and… See the full description on the dataset page: https://huggingface.co/datasets/UWGZQ/Synthetic_Visual_Genome2.textvideo-classification100K<n<1M8 likes1.2k downloads11d agoHugging Face08neulab /VisualPuzzles VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge 🏠 Homepage | 📊 VisualPuzzles | 💻 Github | 📄 Arxiv | 📕 PDF | 🖥️ Zeno Model Output Overview VisualPuzzles is a multimodal benchmark specifically designed to evaluate reasoning abilitiesin large models while deliberately minimizing reliance on domain-specific knowledge. Key features: 1168 diverse puzzles 5 reasoning categories: Algorithmic, Analogical, Deductive, Inductive, Spatial… See the full description on the dataset page: https://huggingface.co/datasets/neulab/VisualPuzzles.imagevisual-question-answering1K<n<10K12 likes1.1k downloads6mo agoHugging Face09TIGER-Lab /VisualWebInstruct-Seed Introduction This is the seed dataset we used to conduct Google Search. Links Github| Paper| Website Citation @article{visualwebinstruct, title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search}, author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu}, journal={arXiv preprint arXiv:2503.10582}, year={2025} } imagequestion-answering10K<n<100K19 likes1.1k downloads2y agoHugging Face10VisualSphinx /VisualSphinx-V1-Raw 🦁 VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL VisualSphinx is the largest fully-synthetic open-source dataset providing vision logic puzzles. It consists of over 660K automatically generated logical visual puzzles. Each logical puzzle is grounded with an interpretable rule and accompanied by both correct answers and plausible distractors. 🌐 Project Website - Learn more about VisualSphinx 📖 Technical Report - Discover the methodology and technical details… See the full description on the dataset page: https://huggingface.co/datasets/VisualSphinx/VisualSphinx-V1-Raw.imageimage-text-to-text100K<n<1M7 likes1.1k downloads1y agoHugging Face11visualwebbench /VisualWebBench VisualWebBench Dataset for the paper: VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding? 🌐 Homepage | 🐍 GitHub | 📖 arXiv Introduction We introduce VisualWebBench, a multimodal benchmark designed to assess the understanding and grounding capabilities of MLLMs in web scenarios. VisualWebBench consists of seven tasks, and comprises 1.5K human-curated instances from 139 real websites, covering 87 sub-domains. We evaluate 14… See the full description on the dataset page: https://huggingface.co/datasets/visualwebbench/VisualWebBench.imageimage-to-text1K<n<10K18 likes783 downloads2y agoHugging Face12multimodal-reasoning-lab /Visual-Searchimage10K<n<100K5 likes760 downloads1y agoHugging Face13visual-layer /oxford-iiit-pet-vl-enriched Visualize on Visual Layer Oxford-IIIT-Pets-VL-Enriched An enriched version of the Oxford IIIT Pets Dataset with image caption, bounding boxes, and label issues! With this additional information, the Oxford IIIT Pet dataset can be extended to various tasks such as image retrieval or visual question answering. The label issues help to curate a cleaner and leaner dataset. Description The dataset consists of 6 columns: image_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/oxford-iiit-pet-vl-enriched.imageimage-classification1K<n<10K9 likes716 downloads2y agoHugging Face14Arasaaf /visual_26k_reasoningimage10K<n<100K1 likes680 downloads6mo agoHugging Face15ScaleAI /VisualToolBench VisToolBench Dataset A benchmark dataset for evaluating vision-language models on tool-use tasks. Dataset Statistics Total samples: 1204 Single-turn: 603 Multi-turn: 601 Schema Column Type Description id string Unique task identifier turncase string Either "single-turn" or "multi-turn" num_turns int Number of conversation turns (1 for single-turn) prompt_category string Task category (e.g., "medical", "scientific", "general") eval_focus… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/VisualToolBench.imagevisual-question-answering1K<n<10K7 likes636 downloads9mo agoHugging Face16Reza2kn /visualears-persian-asr-16k 🗂️ visualears-persian-asr-16k English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Main public Persian ASR audio/text dataset: 3.93M 16 kHz rows. مجموعه‌دادهٔ اصلی و عمومی شنوا برای آموزش بازشناسی گفتار فارسی؛ شامل صوت ۱۶ کیلوهرتز، متن و فرادادهٔ منشأ در مقیاس چندمیلیونی. 🧩 Role flagship training corpus پیکرهٔ اصلی آموزش 📦 Snapshot 176 files; approximately 512.17 GB 176… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-persian-asr-16k.audio1M<n<10M1 likes550 downloads2mo agoHugging Face17open-paws /visual-qa-llama-format Open Paws Visual Qa Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Multimodal Data Format: JSONL (JSON Lines) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/visual-qa-llama-format.imagetext-generation1M<n<10M1 likes541 downloads1y agoHugging Face18DistantSky /placement_visual_conditionedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "hessian", "total_episodes": 197, "total_frames": 66397, "total_tasks": 1, "total_videos": 591, "total_chunks": 1, "chunks_size": 1000, "fps": 60, "splits": { "train": "0:197" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DistantSky/placement_visual_conditioned.tabularrobotics10K<n<100K0 likes517 downloads3mo agoHugging Face19tanhuajie2001 /spatial-visual-reasoning-66kimage10K<n<100K2 likes444 downloads2y agoHugging Face20hf-internal-testing /document-visual-retrieval-test Model Card: Document Visual Retrieval Test (internal) Dataset Overview This dataset is designed to evaluate the performance of visual retrievers by testing their ability to match a query to a relevant image. Each of the three examples in this dataset contains a text query and an associated image, which is a scanned page from the foundational "Attention is All You Need" paper. The purpose of this dataset is to facilitate the evaluation of visual retrievers, where the… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/document-visual-retrieval-test.imagen<1K1 likes400 downloads2y agoHugging Face21minglingfeng /Ocean_R1_collected_visual_dataimage100K<n<1M3 likes395 downloads2y agoHugging Face22visual-layer /coco-2014-vl-enriched Visualize on Visual Layer COCO-2014-VL-Enriched An enriched version of the COCO 2014 dataset with label issues! The label issues help to curate a cleaner and leaner dataset. Description The dataset consists of 6 columns: image_id: The original image filename from the COCO dataset. image: Image data in the form of PIL Image. label_bbox: Bounding box annotations from the COCO dataset. Consists of bounding box coordinates, confidence scores, and labels… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/coco-2014-vl-enriched.imageobject-detection100K<n<1M2 likes386 downloads2y agoHugging Face23knowledge-in-visual-synthesis /v1 Knowledge in Visual Synthesis This dataset contains prompt–image examples for evaluating and studying knowledge-intensive visual synthesis. Samples are organized by contributor as dataset subsets (configs), with each upload version exposed as a split. Dataset structure Subset Splits byx v1, v2 yuner v1, v2 zanyi v1, v2, v3 jiayu v1, v2, v3 sherry v1, v2 yujunz v1 The byx/v1 split contains 140 unique prompts and 300 generated images. For… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-in-visual-synthesis/v1.image1K<n<10K0 likes385 downloads2mo agoHugging Face24AI-4-Everyone /Visual-TableQA 🧠 Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images Welcome to Visual-TableQA, a project designed to generate high-quality synthetic question-answer datasets associated to images of tables. This resource is ideal for training and evaluating models on visually-grounded table understanding tasks such as document QA, table parsing, and multimodal reasoning. 🚀 Latest Update We have refreshed the dataset with newly generated QA pairs created by… See the full description on the dataset page: https://huggingface.co/datasets/AI-4-Everyone/Visual-TableQA.imagetable-question-answering1K<n<10K12 likes376 downloads1y agoHugging Face25taoye1992 /VisualWebInstruct VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search VisualWebInstruct is a large-scale, diverse multimodal instruction dataset designed to enhance vision-language models' reasoning capabilities. The dataset contains approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs associated with 163,743 unique images, while the remaining 60% are text-only QA pairs. Links GitHub Repository Research Paper Project Website… See the full description on the dataset page: https://huggingface.co/datasets/taoye1992/VisualWebInstruct.imagequestion-answering1M<n<10M0 likes335 downloads9mo agoHugging Face26ljnlonoljpiljm /visual_genomeimage100K<n<1M0 likes333 downloads5mo agoHugging Face27multimodal-reasoning-lab /Visual-Jigsawimage10K<n<100K4 likes315 downloads1y agoHugging Face28weikaih /SOC-Training-Data-Visualization Paper Link SOS: Synthetic Object Segments Improve Detection, Segmentation, and Grounding Code repo Code for Generation Citation @misc{huang2025sossyntheticobjectsegments, title={SOS: Synthetic Object Segments Improve Detection, Segmentation, and Grounding}, author={Weikai Huang and Jieyu Zhang and Taoyang Jia and Chenhao Zheng and Ziqi Gao and Jae Sung Park and Ranjay Krishna}, year={2025}, eprint={2510.09110}, archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/SOC-Training-Data-Visualization.imagen<1K0 likes297 downloads1y agoHugging Face29VisualSphinx /VisualSphinx-V1-Raw-Panelsimage100K<n<1M0 likes294 downloads11mo agoHugging Face30visual-memory /PersonaChat-Qwen-Image-2512-enhancedimage10K<n<100K0 likes282 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.