CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gunnybd01 /faiss-integration-testtabularn<1K0 likes670 downloads4mo agoHugging Face02bio-protocol /bio-faiss-longevity-v1 bio-faiss-longevity-v1 A FAISS index + metadata for scientific retrieval Contents index.faiss: FAISS index (cosine w/ inner product). meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost. index.info.json: (optional) dimensions, index type, faiss version. Build provenance Chunking: hierarchical (section→paragraph→~480-token chunks, ~15% overlap) Embedder:… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-longevity-v1.tabulartext-retrieval100K<n<1M0 likes118 downloads1y agoHugging Face03bio-protocol /neophyte-faiss-index-v1 neophyte-faiss-index-v1 A FAISS index + metadata for scientific retrieval Contents index.faiss: FAISS index (cosine w/ inner product). meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost. index.info.json: (optional) dimensions, index type, faiss version. Build provenance Chunking: hierarchical (section→paragraph→~480-token chunks, ~15% overlap) Embedder:… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/neophyte-faiss-index-v1.tabulartext-retrieval100K<n<1M0 likes92 downloads11mo agoHugging Face04bio-protocol /bio-faiss-d1ckgpt-v1 bio-faiss-d1ckgpt-v1 A FAISS index + metadata for scientific retrieval Contents index.faiss: FAISS index (cosine w/ inner product). meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost. Build provenance Chunking: hierarchical (section→paragraph→~380-token chunks, ~15% overlap) Embedder: bio-protocol/scientific-retriever (mean-pooled, L2-normalized) Similarity: cosine via inner… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-d1ckgpt-v1.tabulartext-retrieval1K<n<10K0 likes65 downloads1y agoHugging Face05Ram-G /Wiki_Faiss_Indexes dataset_info: features: - name: text dtype: string - name: embeddings dtype: float32 shape: [384] configs: - config_name: default data_files: "*.parquet" Wikipedia IVF-OPQ-PQ Vector Database (GPU-Optimized) A high-performance, GPU-accelerated FAISS vector database built from Wikipedia articles with pre-computed embeddings. This dataset contains approximately 35 million Wikipedia articles with 384-dimensional embeddings using the all-MiniLM-L6-v2… See the full description on the dataset page: https://huggingface.co/datasets/Ram-G/Wiki_Faiss_Indexes.tabularfeature-extractionn<1K1 likes52 downloads1y agoHugging Face06bio-protocol /bio-faiss-microbiome-v1 bio-faiss-microbiome-v1 A FAISS index + metadata for scientific retrieval Contents index.faiss: FAISS index (cosine w/ inner product). meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost. Build provenance Chunking: hierarchical (section→paragraph→~380-token chunks, ~15% overlap) Embedder: bio-protocol/scientific-retriever (mean-pooled, L2-normalized) Similarity: cosine via inner… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-microbiome-v1.tabulartext-retrieval10K<n<100K0 likes46 downloads1y agoHugging Face07free-law /Caselaw_Access_Project_FAISS_index The Caselaw Access Project In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/ Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/free-law/Caselaw_Access_Project_FAISS_index.tabulartext-generationn<1K10 likes42 downloads3y agoHugging Face08PaczkiLives /daemon-wiki-faiss Daemon Wiki FAISS Index Pre-built FAISS IVFPQ index and metadata for the Daemon conversational RAG system. Contents File Size Description vector_index_ivf.faiss ~2.2 GB FAISS IVFPQ index (48 subquantizers x 8 bits, ~32x compression) metadata.parquet ~12 GB Row-group metadata (titles, text, timestamps) for zero-copy lookup Coverage: ~41 million vectors from 6.5M+ English Wikipedia articles, embedded with sentence-transformers/all-MiniLM-L6-v2 (384-dim).… See the full description on the dataset page: https://huggingface.co/datasets/PaczkiLives/daemon-wiki-faiss.tabularfeature-extraction10M<n<100M0 likes31 downloads6mo agoHugging Face09pashaa /ragtime-qwen3-8b-faiss-PQ2048x4fstabularn<1K0 likes30 downloads5mo agoHugging Face10vishnupriyavr /wiki-movie-plots-with-summaries-faiss-embeddings Dataset Card for "wiki-movie-plots-with-summaries-faiss-embeddings" More Information needed tabular10K<n<100K3 likes28 downloads3y agoHugging Face11genomenet /twin-uniref50-faiss Twin-Model UniRef50 FAISS Index FAISS index over Twin model mean-pooled embeddings of all UniRef50 representative proteins (~49.8M). The Twin model is a two-tower contrastive encoder fine-tuned on Resnik GO similarity: Custom tower: AA-vocab Transformer → padding-masked mean pool → MLP → 512-dim ESM tower: facebook/esm2_t33_650M_UR50D (frozen) → masked mean pool → MLP → 512-dim Output: concat(custom, esm) → 1024-dim Checkpoint:… See the full description on the dataset page: https://huggingface.co/datasets/genomenet/twin-uniref50-faiss.tabularn<1K0 likes25 downloads5mo agoHugging Face12NuriAk /Salaries_ds_prepared_for_FAISStabular100K<n<1M0 likes24 downloads3y agoHugging Face13Kitk568 /faiss-indextabular100K<n<1M0 likes17 downloads2y agoHugging Face14azusa-nami /kilt-qwen-faiss-indextabularn<1K0 likes8 downloads5mo agoHugging Face15genomenet /esm2-uniref50-faiss ESM2 UniRef50 FAISS Index FAISS index over ESM2 (esm2_t33_650M_UR50D) mean-pooled embeddings of GO-annotated UniRef50 proteins. Used by the genomenet/functional-distance Space for nearest-neighbor search. Files File Description esm2_uniref50.index FAISS index (OPQ + IVF + PQ, cosine / inner product on L2-normalized vectors) ids.npy UniRef50 cluster IDs aligned with FAISS positions (dtype='S24') metadata.json Build parameters (dim, factory, nprobe, n_vectors… See the full description on the dataset page: https://huggingface.co/datasets/genomenet/esm2-uniref50-faiss.tabularn<1K0 likes7 downloads5mo agoHugging Face16manhdungcr7 /aic2026-faiss-meili-dbtabular100K<n<1M0 likes7 downloads2mo agoHugging Face17Boomgaard /morrowind-faiss-indextabularn<1K0 likes5 downloads6mo agoHugging Face18Boomgaard /morrowind-faiss-index-v2tabularn<1K0 likes5 downloads6mo agoHugging Face19Boomgaard /morrowind-faiss-index-v3tabularn<1K0 likes5 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.