multi-vector
multi-vector-search-datasets
Multi-Vector Search Datasets
The datasets listed below are used in the Multi-Vector HNSW project for testing and benchmarking multi-vector approximate nearest neighbor search algorithms and their implementations.
Stack Exchange Datasets
Source: habedi/stack-exchange-dataset
Each row contains:
id: unique post ID
title: the post title
body: the main body content (with HTML tags removed)
tags: associated tags
embedding: a list of three 768-dimensional vectors for [title… See the full description on the dataset page: https://huggingface.co/datasets/habedi/multi-vector-search-datasets.multi-vector-hnsw-datasets
Multi-Vector HNSW Benchmark Datasets
This repository contains benchmark datasets used by the Multi-Vector HNSW project.
The datasets are from habedi/multi-vector-search-datasets.
Each record includes a question ID and three distinct 768-dimensional vectors representing the title, body, and tags of the question.
The text embeddings were generated using the all-mpnet-base-v2 text embedding model.
There are three datasets; each includes questions from a separate Q&A community hosted on… See the full description on the dataset page: https://huggingface.co/datasets/habedi/multi-vector-hnsw-datasets.arXiv-AI-papers-multi-vector
Overview
This is a dataset containing individual pages from the top-40 most cited AI papers on arXiv](https://arxiv.org/abs/2412.12121) from the period 2023-01-01 to 2024-09-30.
Only the first 10 pages from each paper is included.
The dataset includes an image of each page as well as a multi-vector embedding using vidore/colqwen2-v1.0.
emotions-multi-vectorsmulti-vector-sketchingHydrus-AM-Thinking-Multi-Turn
