VDBBench/multimodal-embedding-1M
Multimodal Embedding 1M Benchmark Dataset A vector search benchmark dataset containing 1M base vectors and 10K query vectors with pre-computed ground truth, generated from multimodal (image + text) inputs. Dataset Description Each embedding is produced by encoding an image-text pair into a single 4096-dimensional vector using Qwen3-VL-Embedding-8B, a state-of-the-art multimodal embedding model. Source data: pixparse/cc3m-wds (Conceptual Captions 3M in WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/VDBBench/multimodal-embedding-1M.
0108
Upload README.md with huggingface_hub
Upload folder using huggingface_hub
initial commit
