datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
video-lvlm-datalvlm-ood-fake-dataUncovering_LVLM_BiasLVLM_NLFNOTE: LVLM_NLF and VLSafe are constructed based on COCO and LLaVA. So the image can be directly retrieved from the COCO train-2017 version using the image id.
LVLM_NLF (Large Vision Language Model with Natural Language Feedback) Dataset Card
Dataset details
Dataset type: LVLM_NLF is a GPT-4-Annotated natural language feedback dataset that aims to improve the 3H alignment and interaction ability of large vision-language models (LVLMs).
Dataset date: LVLM_NLF was collected between September… See the full description on the dataset page: https://huggingface.co/datasets/YangyiYY/LVLM_NLF.SOCO-LVLM
SOCO-LVLM
SOCO-LVLM provides multiple-choice semantic object correspondence evaluation data for
LVLMs. This is the SOCO-LVLM v1 release, derived from SOCOv1. The original SOCO
correspondence benchmark is available in
the GenIntelLab/SOCO dataset repository.
Repository Layout
GenIntelLab/SOCO-LVLM
SOCO_LVLM/
soco_lvlm_img.tsv
soco_lvlm_imgtxt.tsv
soco_lvlm_txt.tsv
README.md
Variants
soco_lvlm_img.tsv: image-input evaluation variant… See the full description on the dataset page: https://huggingface.co/datasets/GenIntelLab/SOCO-LVLM.Extending-LVLMs-NonEnglish-DataResource associated to the paper "Extending Large Language Models to Multimodality for non-English Languages".
We provide the following resources:
align/: projector alignment dataset of LLaVA translated to Italian and Spanish using madlad
test/: contains the MultiInstruct test set formatted in English, as well as the formatted and translated version in Italian and Spanish using MadLad and human written formattings
test_conversation/: contains the ImageDialog type 1 subset of the OmniDialog… See the full description on the dataset page: https://huggingface.co/datasets/swap-uniba/Extending-LVLMs-NonEnglish-Data.stic-coco-preference-6kstic-llava-instruct-desc-5kLVLM_InterpretationThis repository contains the IDs of a subset of question used in the project: Where do Large Vision-Language Models Look at when Answering Questions? [paper] [code]
It is a heatmap visualization method for interpreting Large Vision-Language Models (LVLMs) when generating open-ended answers.
The original datasets can be obtained at CV-Bench, MMVP, MMStar. We sincerely appreciate the authors of these datasets for their contributions. This selected subseted is based on the relevance of the… See the full description on the dataset page: https://huggingface.co/datasets/xiaoying0505/LVLM_Interpretation.datasets
SEM anomaly-detection datasets
Datasets for reproducing the SEM anomaly-detection results (detection + PALM domain
adaptation). Delivered as *.zip files; download and unzip into data/datasets/.
Contents
file
size
what it is
used for
miic_test.zip
14G
MIIC SEM test set — 5,000 normal + 116 abnormal
MIIC detection eval
oasis.zip
3.9G
OASIS instruction-tuning data — synthetic particle/cut/bridge defects on normal SEM (per-class *_llava.json + images)… See the full description on the dataset page: https://huggingface.co/datasets/lvlm-anomaly-detection/datasets.
