vision-language
Touch-Vision-Language-Dataset
A Touch, Vision, and Language Dataset for Multimodal Alignment
by Max (Letian) Fu, Gaurav Datta*, Huang Huang*, William Chung-Ho Panitch*, Jaimyn Drake*, Joseph Ortiz, Mustafa Mukadam, Mike Lambeta, Roberto Calandra, Ken Goldberg at UC Berkeley, Meta AI, TU Dresden and CeTI (*equal contribution).
[Paper] | [Project Page] | [Checkpoints] | [Dataset] | [Citation]
This repo contains the dataset for A Touch, Vision, and Language Dataset for Multimodal Alignment.… See the full description on the dataset page: https://huggingface.co/datasets/mlfu7/Touch-Vision-Language-Dataset.olfaction-vision-language-dataset
Olfaction-Vision-Language Learning: A Multimodal Dataset
Olfaction • Vision • Language
An open-sourced dataset and dataset builder for prototyping and exploratory olfaction-vision-language tasks within the AI, robotics, and AR/VR domains.
Whether this dataset is used for better vision-scent navigation with drones, triangulating the source of an odor in an image, extracting aromas from a scene, or augmenting a VR experience with scent, we hope its release will catalyze… See the full description on the dataset page: https://huggingface.co/datasets/kordelfrance/olfaction-vision-language-dataset.VISION_LANGUAGEA key question for understanding multimodal vs. language capabilities of models is what is
the relative strength of the spatial reasoning and understanding in each modality, as spatial understanding is
expected to be a strength for multimodality? To test this we created a procedurally generatable, synthetic dataset
to testing spatial reasoning, navigation, and counting. These datasets are challenging and by
being procedurally generated new versions can easily be created to combat the effects… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/VISION_LANGUAGE.calame-pt
CALAME-PT
Context-Aware LAnguage Modeling Evaluation for Portuguese
CALAME-PT is a PT benchmark composed of small texts (contexts) and their respective last words.
These contexts should, in theory, contain enough information so that a human or a model is capable of guessing its last word - without being too specific and/or too ambiguous.
Composition
CALAME-PT is composed of 2 "sets" of data - handwritten and generated.
Handwritten Set: contains 406… See the full description on the dataset page: https://huggingface.co/datasets/NOVA-vision-language/calame-pt.Touch-Vision-Language-Dataset
A Touch, Vision, and Language Dataset for Multimodal Alignment
by Max (Letian) Fu, Gaurav Datta*, Huang Huang*, William Chung-Ho Panitch*, Jaimyn Drake*, Joseph Ortiz, Mustafa Mukadam, Mike Lambeta, Roberto Calandra, Ken Goldberg at UC Berkeley, Meta AI, TU Dresden and CeTI (*equal contribution).
[Paper] | [Project Page] | [Checkpoints] | [Dataset] | [Citation]
This repo contains the dataset for A Touch, Vision, and Language Dataset for Multimodal Alignment.… See the full description on the dataset page: https://huggingface.co/datasets/JimingYang/Touch-Vision-Language-Dataset.visionlanguagemodelog
