CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfu7 /Touch-Vision-Language-Dataset A Touch, Vision, and Language Dataset for Multimodal Alignment by Max (Letian) Fu, Gaurav Datta*, Huang Huang*, William Chung-Ho Panitch*, Jaimyn Drake*, Joseph Ortiz, Mustafa Mukadam, Mike Lambeta, Roberto Calandra, Ken Goldberg at UC Berkeley, Meta AI, TU Dresden and CeTI (*equal contribution). [Paper] | [Project Page] | [Checkpoints] | [Dataset] | [Citation] This repo contains the dataset for A Touch, Vision, and Language Dataset for Multimodal Alignment.… See the full description on the dataset page: https://huggingface.co/datasets/mlfu7/Touch-Vision-Language-Dataset.10 likes882 downloads1y agoHugging Face02kordelfrance /olfaction-vision-language-dataset Olfaction-Vision-Language Learning: A Multimodal Dataset Olfaction • Vision • Language An open-sourced dataset and dataset builder for prototyping and exploratory olfaction-vision-language tasks within the AI, robotics, and AR/VR domains. Whether this dataset is used for better vision-scent navigation with drones, triangulating the source of an odor in an image, extracting aromas from a scene, or augmenting a VR experience with scent, we hope its release will catalyze… See the full description on the dataset page: https://huggingface.co/datasets/kordelfrance/olfaction-vision-language-dataset.image-classification10K<n<100K3 likes408 downloads1y agoHugging Face03microsoft /VISION_LANGUAGEA key question for understanding multimodal vs. language capabilities of models is what is the relative strength of the spatial reasoning and understanding in each modality, as spatial understanding is expected to be a strength for multimodality? To test this we created a procedurally generatable, synthetic dataset to testing spatial reasoning, navigation, and counting. These datasets are challenging and by being procedurally generated new versions can easily be created to combat the effects… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/VISION_LANGUAGE.image10K<n<100K6 likes370 downloads2y agoHugging Face04NOVA-vision-language /calame-pt CALAME-PT Context-Aware LAnguage Modeling Evaluation for Portuguese CALAME-PT is a PT benchmark composed of small texts (contexts) and their respective last words. These contexts should, in theory, contain enough information so that a human or a model is capable of guessing its last word - without being too specific and/or too ambiguous. Composition CALAME-PT is composed of 2 "sets" of data - handwritten and generated. Handwritten Set: contains 406… See the full description on the dataset page: https://huggingface.co/datasets/NOVA-vision-language/calame-pt.text1K<n<10K2 likes342 downloads3y agoHugging Face05JimingYang /Touch-Vision-Language-Dataset A Touch, Vision, and Language Dataset for Multimodal Alignment by Max (Letian) Fu, Gaurav Datta*, Huang Huang*, William Chung-Ho Panitch*, Jaimyn Drake*, Joseph Ortiz, Mustafa Mukadam, Mike Lambeta, Roberto Calandra, Ken Goldberg at UC Berkeley, Meta AI, TU Dresden and CeTI (*equal contribution). [Paper] | [Project Page] | [Checkpoints] | [Dataset] | [Citation] This repo contains the dataset for A Touch, Vision, and Language Dataset for Multimodal Alignment.… See the full description on the dataset page: https://huggingface.co/datasets/JimingYang/Touch-Vision-Language-Dataset.0 likes188 downloads3mo agoHugging Face06ipranavks /visionlanguagemodelogimagen<1K4 likes123 downloads1y agoHugging Face07beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K3 likes113 downloads1mo agoHugging Face08Embodied-Vision-Language-Model /ShareRobot8 likes94 downloads2y agoHugging Face09NOVA-vision-language /MSCOCO_PT-BRtextn<1K2 likes81 downloads2y agoHugging Face10NOVA-vision-language /CC3M_PT-BR2 likes76 downloads2y agoHugging Face11Strangefiction /olfaction-vision-language-dataset Olfaction-Vision-Language Learning: A Multimodal Dataset Olfaction • Vision • Language An open-sourced dataset and dataset builder for prototyping and exploratory olfaction-vision-language tasks within the AI, robotics, and AR/VR domains. Whether this dataset is used for better vision-scent navigation with drones, triangulating the source of an odor in an image, extracting aromas from a scene, or augmenting a VR experience with scent, we hope its release will catalyze… See the full description on the dataset page: https://huggingface.co/datasets/Strangefiction/olfaction-vision-language-dataset.image-classification10K<n<100K0 likes75 downloads5mo agoHugging Face12hassan-wajid /Spatial-Blind-Spots-in-Vision-Language-Modelslicense: mit model_evaluated: name: Qwen3-VL-2B-Instruct url: https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct evaluation_notebook: https://www.kaggle.com/code/wajidhassanmoosa/blind-spot-qwen3-2b evaluation_setup: | The model evaluated in this study is Qwen3-VL-2B-Instruct. Evaluation was conducted using the Hugging Face Transformers library with automatic device mapping (device_map="auto") and "bfloat16" dtype selection. For each example: The image was provided as part of a… See the full description on the dataset page: https://huggingface.co/datasets/hassan-wajid/Spatial-Blind-Spots-in-Vision-Language-Models.imagen<1K2 likes70 downloads7mo agoHugging Face13ipranavks /visionlanguagemodel2 likes36 downloads1y agoHugging Face14wrom /Language-Vision-Hallucinations Dataset for Techen Project 095280 A comprehensive dataset for the Techen Project, focused on examining hallucinations in multi-modal AI-generated text by investigating model uncertainty, text generation patterns, and linguistic factors. Columns Overview image_link: URL to the image associated with each data row. temperature: Temperature setting for text generation, controlling output randomness. description: Text generated by the model for each image, using the… See the full description on the dataset page: https://huggingface.co/datasets/wrom/Language-Vision-Hallucinations.imagen<1K2 likes17 downloads2y agoHugging Face15fineset-io /vision-language-action-papers Vision-Language-Action (VLA) & Robot Learning Papers — FineSet A research-paper dataset on Vision-Language-Action (VLA) & Robot Learning Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Vision-Language-Action (VLA) & Robot Learning Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/vision-language-action-papers.tabulartext-classificationn<1K0 likes14 downloads3mo agoHugging Face16NOVA-vision-language /VisualDialogue_PT-BRtextn<1K0 likes10 downloads2y agoHugging Face17Omarrran /vision_language_pairs_data Vision-Language Pairs Dataset This dataset contains metadata about image-text pairs from various popular vision-language datasets. Contents vision_language_data/all_vision_language_images.csv: Combined metadata for all images (75629 records) vision_language_data/all_vision_language_captions.csv: Combined captions for all images (86676 records) dataset_statistics.csv: Summary statistics for each dataset category_distribution.csv: Distribution of image categories across… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/vision_language_pairs_data.image0 likes4 downloads1y agoHugging Face18kraimon /humanoid-vision-language-commands0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.