vector-data
cwicr-vector-db-bgem3-v3
CWICR Vector Database — BGE-M3 V3 Snapshots
Production Qdrant snapshots for CWICR (Construction Works Items, Costs & Resources) — a multilingual catalogue of construction rate databases covering 30 countries / language locales. Each snapshot encodes one country's rate book using the BAAI/bge-m3 embedder and is ready to restore directly into a Qdrant server for hybrid semantic search.
These snapshots are the V3 production artifacts produced by the OpenConstructionEstimate / CWICR… See the full description on the dataset page: https://huggingface.co/datasets/DataDrivenConstruction/cwicr-vector-db-bgem3-v3.hackernews-vector-search-datasetThe Hacker News dataset contains 28.74 million postings and their vector embeddings. The embeddings were generated using SentenceTransformers model all-MiniLM-L6-v2. The dimension of each embedding vector is 384.
Created by clickhouse more info: https://clickhouse.com/docs/getting-started/example-datasets/hackernews-vector-search-dataset
vector_databasesvector_dataset_roberta-fine-tunedunbias-plus-dataset
Unbias Dataset
This dataset contains configurations used for the Unbias project at the Vector Institute:
train_4 (config, default): Our newest and highest quality training split.
other_splits (config): Contains the earlier splits below.
train_1: Training split sourced from VLDBench (regenerated version).
train_2: Another training split.
train_3: Another training split.
test_set: Test split sourced from BABE Golden 500.
⭐ train_4 is our newest, highest quality… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/unbias-plus-dataset.multi-vector-search-datasets
Multi-Vector Search Datasets
The datasets listed below are used in the Multi-Vector HNSW project for testing and benchmarking multi-vector approximate nearest neighbor search algorithms and their implementations.
Stack Exchange Datasets
Source: habedi/stack-exchange-dataset
Each row contains:
id: unique post ID
title: the post title
body: the main body content (with HTML tags removed)
tags: associated tags
embedding: a list of three 768-dimensional vectors for [title… See the full description on the dataset page: https://huggingface.co/datasets/habedi/multi-vector-search-datasets.
