datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
faiss-integration-testbio-faiss-longevity-v1
bio-faiss-longevity-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
index.info.json: (optional) dimensions, index type, faiss version.
Build provenance
Chunking: hierarchical (section→paragraph→~480-token chunks, ~15% overlap)
Embedder:… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-longevity-v1.neophyte-faiss-index-v1
neophyte-faiss-index-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
index.info.json: (optional) dimensions, index type, faiss version.
Build provenance
Chunking: hierarchical (section→paragraph→~480-token chunks, ~15% overlap)
Embedder:… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/neophyte-faiss-index-v1.bio-faiss-d1ckgpt-v1
bio-faiss-d1ckgpt-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
Build provenance
Chunking: hierarchical (section→paragraph→~380-token chunks, ~15% overlap)
Embedder: bio-protocol/scientific-retriever (mean-pooled, L2-normalized)
Similarity: cosine via inner… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-d1ckgpt-v1.Wiki_Faiss_Indexes
dataset_info:
features:
- name: text
dtype: string
- name: embeddings
dtype: float32
shape: [384]
configs:
- config_name: default
data_files: "*.parquet"
Wikipedia IVF-OPQ-PQ Vector Database (GPU-Optimized)
A high-performance, GPU-accelerated FAISS vector database built from Wikipedia articles with pre-computed embeddings. This dataset contains approximately 35 million Wikipedia articles with 384-dimensional embeddings using the all-MiniLM-L6-v2… See the full description on the dataset page: https://huggingface.co/datasets/Ram-G/Wiki_Faiss_Indexes.bio-faiss-microbiome-v1
bio-faiss-microbiome-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
Build provenance
Chunking: hierarchical (section→paragraph→~380-token chunks, ~15% overlap)
Embedder: bio-protocol/scientific-retriever (mean-pooled, L2-normalized)
Similarity: cosine via inner… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-microbiome-v1.Caselaw_Access_Project_FAISS_index
The Caselaw Access Project
In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/
Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/free-law/Caselaw_Access_Project_FAISS_index.daemon-wiki-faiss
Daemon Wiki FAISS Index
Pre-built FAISS IVFPQ index and metadata for the Daemon conversational RAG system.
Contents
File
Size
Description
vector_index_ivf.faiss
~2.2 GB
FAISS IVFPQ index (48 subquantizers x 8 bits, ~32x compression)
metadata.parquet
~12 GB
Row-group metadata (titles, text, timestamps) for zero-copy lookup
Coverage: ~41 million vectors from 6.5M+ English Wikipedia articles, embedded with sentence-transformers/all-MiniLM-L6-v2 (384-dim).… See the full description on the dataset page: https://huggingface.co/datasets/PaczkiLives/daemon-wiki-faiss.ragtime-qwen3-8b-faiss-PQ2048x4fswiki-movie-plots-with-summaries-faiss-embeddings
Dataset Card for "wiki-movie-plots-with-summaries-faiss-embeddings"
More Information needed
twin-uniref50-faiss
Twin-Model UniRef50 FAISS Index
FAISS index over Twin model mean-pooled embeddings of all UniRef50
representative proteins (~49.8M). The Twin model is a two-tower contrastive
encoder fine-tuned on Resnik GO similarity:
Custom tower: AA-vocab Transformer → padding-masked mean pool → MLP → 512-dim
ESM tower: facebook/esm2_t33_650M_UR50D (frozen) → masked mean pool → MLP → 512-dim
Output: concat(custom, esm) → 1024-dim
Checkpoint:… See the full description on the dataset page: https://huggingface.co/datasets/genomenet/twin-uniref50-faiss.Salaries_ds_prepared_for_FAISSfaiss-indexkilt-qwen-faiss-indexesm2-uniref50-faiss
ESM2 UniRef50 FAISS Index
FAISS index over ESM2 (esm2_t33_650M_UR50D) mean-pooled embeddings of
GO-annotated UniRef50 proteins. Used by the
genomenet/functional-distance
Space for nearest-neighbor search.
Files
File
Description
esm2_uniref50.index
FAISS index (OPQ + IVF + PQ, cosine / inner product on L2-normalized vectors)
ids.npy
UniRef50 cluster IDs aligned with FAISS positions (dtype='S24')
metadata.json
Build parameters (dim, factory, nprobe, n_vectors… See the full description on the dataset page: https://huggingface.co/datasets/genomenet/esm2-uniref50-faiss.aic2026-faiss-meili-dbmorrowind-faiss-indexmorrowind-faiss-index-v2morrowind-faiss-index-v3
