qdrant
Datasets
All datasets matching “qdrant”FineWeb-10B
Qdrant-FineWeb-10B
Overview
Qdrant-FineWeb-10B (Q-FineWeb-10B) is a 10-billion-vector retrieval benchmark derived from FineWeb. Each document is represented with dense and sparse embeddings from Alibaba-NLP/gte-multilingual-base, alongside its original FineWeb payload and metadata. The benchmark also includes exact brute-force ground truth for ~120,000 MS MARCO queries.
The dataset includes:
10 billion dense embeddings
10 billion sparse embeddings
FineWeb… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/FineWeb-10B.gaianet-qdrant-snapshotPubMed-MV
PubMed-MultiVector (PubMed-VE)
Overview
PubMed-VE is a retrieval benchmark built from PubMed abstracts using BGE-M3.
Each document is represented in all three formats produced by BGE-M3:
~24 million dense embeddings
~24 million sparse embeddings
~8.77B multi-vector (token-level) embeddings
Statistic
Value
Documents
23,898,701
Queries
1,000
Ground truth
exact top-1000, one list per modality (dense / sparse / multi-vector)
Embeddings
BAAI/bge-m3… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/PubMed-MV.wikipedia_20240401_10-languages_bge-m3_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked Wikipedia HTML dumps from 10 languages. The embedding model used is BAAI/bge-m3
This index is compatible with WikiChat v2.0.
Refer to the following for more information:
GitHub repository: https://github.com/stanford-oval/WikiChat
Papers:
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia_20240401_10-languages_bge-m3_qdrant_index.Coyo-VE
Coyo-Vector-Embeddings (Coyo-VE)
Overview
Coyo-VE is a large-scale visual-text-embedding slice of the COYO subset from the LLaVA-OneVision-1.5-Mid-Training-85M dataset, with the original images being sourced from coyo-700m. Each image-caption pair is jointly embedded using Qwen3-VL-Embedding-2B, which accepts mixed image-and-text inputs and encodes their combined visual and textual content into a single 2,048-dimensional dense vector. This produces a unified… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/Coyo-VE.arxiv-titles-instructorxl-embeddings
arxiv-titles-instructorxl-embeddings
This dataset contains 768-dimensional embeddings generated from the arxiv
paper titles using InstructorXL model. Each
vector has an abstract used to create it, along with the DOI (Digital Object Identifier). The
dataset was created using precomputed embeddings exposed by the Alexandria Index.
Generation process
The embeddings have been generated using the following instruction:
Represent the Research Paper title for retrieval;… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings.
