CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01microsoft /VISION_LANGUAGEA key question for understanding multimodal vs. language capabilities of models is what is the relative strength of the spatial reasoning and understanding in each modality, as spatial understanding is expected to be a strength for multimodality? To test this we created a procedurally generatable, synthetic dataset to testing spatial reasoning, navigation, and counting. These datasets are challenging and by being procedurally generated new versions can easily be created to combat the effects… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/VISION_LANGUAGE.image10K<n<100K6 likes367 downloads2y agoHugging Face02ipranavks /visionlanguagemodelogimagen<1K4 likes128 downloads1y agoHugging Face03hassan-wajid /Spatial-Blind-Spots-in-Vision-Language-Modelslicense: mit model_evaluated: name: Qwen3-VL-2B-Instruct url: https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct evaluation_notebook: https://www.kaggle.com/code/wajidhassanmoosa/blind-spot-qwen3-2b evaluation_setup: | The model evaluated in this study is Qwen3-VL-2B-Instruct. Evaluation was conducted using the Hugging Face Transformers library with automatic device mapping (device_map="auto") and "bfloat16" dtype selection. For each example: The image was provided as part of a… See the full description on the dataset page: https://huggingface.co/datasets/hassan-wajid/Spatial-Blind-Spots-in-Vision-Language-Models.imagen<1K2 likes76 downloads7mo agoHugging Face04wrom /Language-Vision-Hallucinations Dataset for Techen Project 095280 A comprehensive dataset for the Techen Project, focused on examining hallucinations in multi-modal AI-generated text by investigating model uncertainty, text generation patterns, and linguistic factors. Columns Overview image_link: URL to the image associated with each data row. temperature: Temperature setting for text generation, controlling output randomness. description: Text generated by the model for each image, using the… See the full description on the dataset page: https://huggingface.co/datasets/wrom/Language-Vision-Hallucinations.imagen<1K2 likes16 downloads2y agoHugging Face05Omarrran /vision_language_pairs_data Vision-Language Pairs Dataset This dataset contains metadata about image-text pairs from various popular vision-language datasets. Contents vision_language_data/all_vision_language_images.csv: Combined metadata for all images (75629 records) vision_language_data/all_vision_language_captions.csv: Combined captions for all images (86676 records) dataset_statistics.csv: Summary statistics for each dataset category_distribution.csv: Distribution of image categories across… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/vision_language_pairs_data.image0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.