sift
Datasets
All datasets matching “sift”SIFT1B-DiskANN
SIFT-1B Dataset & Disk Index
The SIFT-1B (BigANN) dataset and pre-built disk-based ANN index.
Built February 2026 on Intel Xeon 8462Y+ (Sapphire Rapids) with 800GB RAM.
Build Parameters
Parameter
Value
Dataset
SIFT-1B (1,000,000,000 vectors, 128-dim, uint8)
Graph R
128 (max degree)
Build L
200 (search list size during construction)
PQ chunks
32 (4 dimensions per sub-quantizer)
Build time
~2 days
Files
Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/Nanvivi/SIFT1B-DiskANN.sift1b
SIFT1B - Sharded DiskANN Indices
Pre-built DiskANN indices for the SIFT1B (BigANN) dataset, sharded for distributed vector search.
Dataset Info
Source: BigANN Benchmarks
Vectors: 1,000,000,000 (1 billion)
Dimensions: 128
Data type: uint8
Queries: 10,000
Distance: L2
DiskANN Parameters
R (graph degree): 64
L (build beam width): 100
PQ bytes: 32
Shard Configurations
shard_2: 2 shards x 500,000,000 vectors
shard_3: 3 shards x ~333,333,333 vectors… See the full description on the dataset page: https://huggingface.co/datasets/makneeeee/sift1b.SIFT-50M
Dataset Card for SIFT-50M
SIFT-50M (Speech Instruction Fine-Tuning) is a 50-million-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). It is built from publicly available speech corpora containing a total of 14K hours of speech and leverages LLMs and off-the-shelf expert models. The dataset spans five languages, covering diverse aspects of speech understanding and controllable speech generation instructions. SIFT-50M… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/SIFT-50M.sift1m
sift1m
sift1m data, copied from http://corpus-texmex.irisa.fr/, published:
Jégou H, Douze M, Schmid C. Improving bag-of-features for large scale image search[J]. International journal of computer vision, 2010, 87(3): 316-336.
siftsmallsift100m
