hnsw
Datasets
All datasets matching “hnsw”wiki-18-e5-index-HNSW64omnimcp_pgvector_hnsw_index_tuner_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_pgvector_hnsw_index_tuner_teaser.emgena_postgres_pgvector_hnsw_tuner_mcp_teaser
🚀 Database - PostgreSQL & PgVector HNSW Index & Buffer Pool Tuner (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout!
🌟 Domain Overview & Features
EXPLAIN (BUFFERS) deep profiling, HNSW M/ef_construction index tuning, and shared buffers cache thrashing… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_postgres_pgvector_hnsw_tuner_mcp_teaser.hnsw_indexThis is the vector index file of our paper SearchAgent-X. It contains a HNSW-based index file of approxiamte nearest neighbor search built for Wikipedia English corpus.
Details of this dataset:
Chunked Text: wiki-18-corpus
Embedding Model: all-MiniLM-L6-v2
HNSW Implementation Base: HNSWLib
Indexing Params:
32 neighbors per node (i.e., M)
candidate list of 500 (i.e., efConstruction)
DIY: How to Encode and Index Your Own Corpus?😄 Follow our github instructions and code here
multi-vector-hnsw-datasets
Multi-Vector HNSW Benchmark Datasets
This repository contains benchmark datasets used by the Multi-Vector HNSW project.
The datasets are from habedi/multi-vector-search-datasets.
Each record includes a question ID and three distinct 768-dimensional vectors representing the title, body, and tags of the question.
The text embeddings were generated using the all-mpnet-base-v2 text embedding model.
There are three datasets; each includes questions from a separate Q&A community hosted on… See the full description on the dataset page: https://huggingface.co/datasets/habedi/multi-vector-hnsw-datasets.tool_HNSW
