sif
Datasets
All datasets matching “sif”omnibioai-sif-images
OmniBioAI SIF Images 🧬
500+ native ARM64 Singularity (SIF) container images for bioinformatics,
built on NVIDIA DGX (aarch64).
Tool Categories
Category
Tools
Genomics & Alignment
BWA, STAR, HISAT2, Minimap2, Bowtie2
Variant Calling
GATK, DeepVariant, Clair3, Mutect2
RNA-seq
Salmon, Kallisto, DESeq2, edgeR
Single Cell
Seurat, Scanpy, Cell Ranger, Harmony
Epigenomics
MACS2, deepTools, Bismark
Metagenomics
Kraken2, MetaPhlAn, QIIME2
Proteomics… See the full description on the dataset page: https://huggingface.co/datasets/omnibioai/omnibioai-sif-images.sif-files-for-swesi_for_sdSIFT1B-DiskANN
SIFT-1B Dataset & Disk Index
The SIFT-1B (BigANN) dataset and pre-built disk-based ANN index.
Built February 2026 on Intel Xeon 8462Y+ (Sapphire Rapids) with 800GB RAM.
Build Parameters
Parameter
Value
Dataset
SIFT-1B (1,000,000,000 vectors, 128-dim, uint8)
Graph R
128 (max degree)
Build L
200 (search list size during construction)
PQ chunks
32 (4 dimensions per sub-quantizer)
Build time
~2 days
Files
Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/Nanvivi/SIFT1B-DiskANN.sift1b
SIFT1B - Sharded DiskANN Indices
Pre-built DiskANN indices for the SIFT1B (BigANN) dataset, sharded for distributed vector search.
Dataset Info
Source: BigANN Benchmarks
Vectors: 1,000,000,000 (1 billion)
Dimensions: 128
Data type: uint8
Queries: 10,000
Distance: L2
DiskANN Parameters
R (graph degree): 64
L (build beam width): 100
PQ bytes: 32
Shard Configurations
shard_2: 2 shards x 500,000,000 vectors
shard_3: 3 shards x ~333,333,333 vectors… See the full description on the dataset page: https://huggingface.co/datasets/makneeeee/sift1b.SIFT-50M
Dataset Card for SIFT-50M
SIFT-50M (Speech Instruction Fine-Tuning) is a 50-million-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). It is built from publicly available speech corpora containing a total of 14K hours of speech and leverages LLMs and off-the-shelf expert models. The dataset spans five languages, covering diverse aspects of speech understanding and controllable speech generation instructions. SIFT-50M… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/SIFT-50M.
