CoolFace
20 results

mpnet

deadbits /vigil-jailbreak-all-mpnet-base-v2 Vigil: LLM Jailbreak all-mpnet-base-v2 Repo: github.com/deadbits/vigil-llm Vigil is a Python framework and REST API for assessing Large Language Model (LLM) prompts against a set of scanners to detect prompt injections, jailbreaks, and other potentially risky inputs. This repository contains all-mpnet-base-v2 embeddings for all "jailbreak" prompts used by Vigil. You can use the parquet2vdb.py utility to load the embeddings in the Vigil chromadb instance, or use them in your own… See the full description on the dataset page: https://huggingface.co/datasets/deadbits/vigil-jailbreak-all-mpnet-base-v2.textn<1K1 likes16k downloads3y agoHugging Facesentence-transformers /msmarco-mpnet-margin-mse-mean-v1 MS MARCO with hard negatives from mpnet-margin-mse-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-mpnet-margin-mse-mean-v1.tabularfeature-extraction10M<n<100M1 likes1.6k downloads2y agoHugging Facerpaut03l /trishieldrag-nq-mpnet-embeddings TriShieldRAG — BeIR NQ embeddings and FAISS index Artifacts for TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation. File Size Description emb_*.npy (27) 7.7 GB Embeddings, 100k passages per chunk, corpus order nq_ivf_nlist6550.index 7.8 GB FAISS IVF index, nlist=6550, inner product Corpus: BeIR/nq corpus split, 2,681,468 passages Model: sentence-transformers/all-mpnet-base-v2, 768-d… See the full description on the dataset page: https://huggingface.co/datasets/rpaut03l/trishieldrag-nq-mpnet-embeddings.feature-extraction0 likes1.2k downloads28d agoHugging Faceolmer /wiki_mpnet_embeddingsEmbeddings of the english Wikipedia paragraphs using all-mpnet-base-v2 sentence transformers encoder.The dataset contains 43 911 155 paragraphs from 6 458 670 Wikipedia articles.The size of each paragraph varies from 20 to 2000 characters.For each paragraph there is an embedding of size 768.Embeddings are stored in numpy files, 1 000 000 embeddings per file.For each embedding file, there is an ids file that contains the list of ids of the corresponding paragraphs.Be careful, dataset size is… See the full description on the dataset page: https://huggingface.co/datasets/olmer/wiki_mpnet_embeddings.texttext-retrieval10M<n<100M1 likes99 downloads3y agoHugging FaceGBaker /MedQA-USMLE-4-options-hf-MPNet-IR Dataset Card for "MedQA-USMLE-4-options-hf-MPNet-IR" More Information needed text10K<n<100K5 likes55 downloads4y agoHugging Facecat-claws /hotpotqa_clustered_dbscan_all-mpnet-base-v2_autotext10K<n<100K0 likes41 downloads1y agoHugging Face