datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bio-faiss-longevity-v1
bio-faiss-longevity-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
index.info.json: (optional) dimensions, index type, faiss version.
Build provenance
Chunking: hierarchical (section→paragraph→~480-token chunks, ~15% overlap)
Embedder:… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-longevity-v1.neophyte-faiss-index-v1
neophyte-faiss-index-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
index.info.json: (optional) dimensions, index type, faiss version.
Build provenance
Chunking: hierarchical (section→paragraph→~480-token chunks, ~15% overlap)
Embedder:… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/neophyte-faiss-index-v1.eduintel-faiss
EduIntel FAISS Index v1
A FAISS vector index of Sepedi (Sesotho sa Leboa) text chunks for the
EduIntel RAG pipeline, built by Sediba AI.
This is not a model — it is a retrieval index. It stores chunked Sepedi text and their
embeddings so that a retriever can find relevant Sepedi content for a query, which is then
fed to a generative model (e.g. Sedibaai/SedibaLM)
for answer generation.
Contents
File
Purpose
eduintel.index
The FAISS index (flat or IVF… See the full description on the dataset page: https://huggingface.co/datasets/Sediba-AI/eduintel-faiss.bio-faiss-d1ckgpt-v1
bio-faiss-d1ckgpt-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
Build provenance
Chunking: hierarchical (section→paragraph→~380-token chunks, ~15% overlap)
Embedder: bio-protocol/scientific-retriever (mean-pooled, L2-normalized)
Similarity: cosine via inner… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-d1ckgpt-v1.Wiki_Faiss_Indexes
dataset_info:
features:
- name: text
dtype: string
- name: embeddings
dtype: float32
shape: [384]
configs:
- config_name: default
data_files: "*.parquet"
Wikipedia IVF-OPQ-PQ Vector Database (GPU-Optimized)
A high-performance, GPU-accelerated FAISS vector database built from Wikipedia articles with pre-computed embeddings. This dataset contains approximately 35 million Wikipedia articles with 384-dimensional embeddings using the all-MiniLM-L6-v2… See the full description on the dataset page: https://huggingface.co/datasets/Ram-G/Wiki_Faiss_Indexes.bio-faiss-microbiome-v1
bio-faiss-microbiome-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
Build provenance
Chunking: hierarchical (section→paragraph→~380-token chunks, ~15% overlap)
Embedder: bio-protocol/scientific-retriever (mean-pooled, L2-normalized)
Similarity: cosine via inner… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-microbiome-v1.reddit_faiss_dbfaiss_search
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/sloppysid/faiss_search.
