laughatwill/TIGER-Lab_MMEB-train
MMEB Training Dataset (Lance Format) This is a Lance-format version of the TIGER-Lab/MMEB-train dataset, optimized for efficient storage and fast random access. The original dataset is used for training VLM2Vec models in the paper VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks (ICLR 2025). Directory Structure TIGER-Lab_MMEB-train/ └── data/ ├── A-OKVQA/ │ ├── train.lance │ ├── original.lance │ └──… See the full description on the dataset page: https://huggingface.co/datasets/laughatwill/TIGER-Lab_MMEB-train.
This repository belongs to laughatwill on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
