CoolFace
20 results

vision

stanford-vision-lab /gpicgated GPIC: A Giant Permissive Image Corpus for Visual Generation Keshigeyan&nbsp;Chandrasegaran*1,&nbsp; Kyle&nbsp;Sargent*1,&nbsp; Suchir&nbsp;Agarwal1,&nbsp; Michael&nbsp;Jang1,&nbsp; Michael&nbsp;Poli1,2,&nbsp; Juan&nbsp;Carlos&nbsp;Niebles1,4,&nbsp; Justin&nbsp;Johnson3,&nbsp; Jiajun&nbsp;Wu1,&nbsp; Li&nbsp;Fei-Fei1 1&nbsp;Stanford University&nbsp;&nbsp; 2&nbsp;Radical Numerics&nbsp;&nbsp; 3&nbsp;University of Michigan&nbsp;&nbsp; 4&nbsp;Salesforce… See the full description on the dataset page: https://huggingface.co/datasets/stanford-vision-lab/gpic.158 likes218k downloads2mo agoHugging Facehf-vision /course-assetsimagen<1K9 likes132k downloads2y agoHugging Faceut-vision /EgoBraingated [ICLR 2026] EgoBrain: Synergizing Minds and Eyes For Human Action Understanding 東京大学 The Univerisity of Tokyo X 微軟亞洲研究院 Microsoft Research Asia Nie Lin · Yansen Wang · Dongqi Han · Weibang Jiang · Jingyuan Li · Ryosuke Furuta · Yoichi Sato* · Dongsheng Li* · *(Co-corresponding authors)* This is the official dataset repository of our ICLR 2026 paper "EgoBrain: Synergizing Minds and Eyes For Human Action… See the full description on the dataset page: https://huggingface.co/datasets/ut-vision/EgoBrain.videon<1K11 likes31k downloads3mo agoHugging Facesensenova /SenseNova-Vision-Corpus-50M Vision as Unified Multimodal Generation English | 简体中文 This repository contains the dataset for the paper Vision as Unified Multimodal Generation. SenseNova Vision Corpus 50M Overview SenseNova Vision Corpus 50M (SN-VC-50M) is a large-scale multimodal vision corpus designed for unified training across diverse visual understanding and geometry-oriented tasks. The dataset is curated to address a common limitation of existing… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/SenseNova-Vision-Corpus-50M.imageany-to-anyn<1K58 likes27k downloads22d agoHugging Facenyu-visionx /VSI-590K VSI-590K Website | Paper | GitHub | Models Authors: Shusheng Yang*, Jihan Yang*, Pinzhi Huang†, Ellis Brown†, et al. VSI-590K is a large-scale spatially-focused instruction-tuning dataset focusing on spatial reasoning. The dataset is curated from diverse sources and carefully annotated. Quick Start import json # Load from JSONL file with open('vsi_590k.jsonl', 'r') as f: for line in f: sample = json.loads(line.strip()) print(sample) break… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/VSI-590K.visual-question-answering100K<n<1M26 likes27k downloads11mo agoHugging Facenyu-visionx /Cambrian-10M Cambrian-10M Dataset Please see paper & website for more information: https://cambrian-mllm.github.io/ https://arxiv.org/abs/2406.16860 Overview Cambrian-10M is a comprehensive dataset designed for instruction tuning, particularly in multimodal settings involving visual interaction data. The dataset is crafted to address the scarcity of high-quality multimodal instruction-tuning data and to maintain the language abilities of multimodal large language models (LLMs).… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-10M.visual-question-answering1M<n<10M131 likes17k downloads2y agoHugging Face