CoolFace
20 results

bge-m3

stanford-oval /wikipedia_20240401_10-languages_bge-m3_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked Wikipedia HTML dumps from 10 languages. The embedding model used is BAAI/bge-m3 This index is compatible with WikiChat v2.0. Refer to the following for more information: GitHub repository: https://github.com/stanford-oval/WikiChat Papers: WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia_20240401_10-languages_bge-m3_qdrant_index.text-retrieval100M<n<1B0 likes3.7k downloads2y agoHugging FaceUpstash /wikipedia-2024-06-bge-m3 Wikipedia Embeddings with BGE-M3 This dataset contains embeddings from the June 2024 Wikipedia dump for the 11 most popular languages. The embeddings are generated with the multilingual BGE-M3 model. The dataset consists of Wikipedia articles split into paragraphs, and embedded with the aforementioned model. To enhance search quality, the paragraphs are prefixed with their respective article titles before embedding. Additionally, paragraphs containing fewer than 100 characters… See the full description on the dataset page: https://huggingface.co/datasets/Upstash/wikipedia-2024-06-bge-m3.text100M<n<1B41 likes2.9k downloads2y agoHugging FaceShitao /bge-m3-data Dataset Summary This depository contains all the fine-tuning data for the bge-m3 model, including: Dataset Language MS MARCO English NQ English HotpotQA English TriviaQA English SQuAD English COLIEE English PubMedQA English NLI from SimCSE English DuReader Chinese mMARCO-zh Chinese T2Ranking Chinese Law-GPT Chinese cMedQAv2 Chinese NLI-zh Chinese LeCaRDv2 Chinese Mr.TyDi 11 languages MIRACL 16 languages MLDR 13 languages Note: The… See the full description on the dataset page: https://huggingface.co/datasets/Shitao/bge-m3-data.text100K<n<1M55 likes2k downloads2y agoHugging Facestanford-oval /wikipedia_20240801_10-languages_bge-m3_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked Wikipedia HTML dumps from 10 languages. The embedding model used is BAAI/bge-m3 This index is compatible with WikiChat v2.0. Refer to the following for more information: GitHub repository: https://github.com/stanford-oval/WikiChat Papers: WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia_20240801_10-languages_bge-m3_qdrant_index.text-retrieval100M<n<1B3 likes1.1k downloads2y agoHugging FaceDataDrivenConstruction /cwicr-vector-db-bgem3-v3 CWICR Vector Database — BGE-M3 V3 Snapshots Production Qdrant snapshots for CWICR (Construction Works Items, Costs & Resources) — a multilingual catalogue of construction rate databases covering 30 countries / language locales. Each snapshot encodes one country's rate book using the BAAI/bge-m3 embedder and is ready to restore directly into a Qdrant server for hybrid semantic search. These snapshots are the V3 production artifacts produced by the OpenConstructionEstimate / CWICR… See the full description on the dataset page: https://huggingface.co/datasets/DataDrivenConstruction/cwicr-vector-db-bgem3-v3.imagen<1K0 likes531 downloads4mo agoHugging Faceminishlab /tokenlearn-c4-multilingual-bge-m3 minishlab/tokenlearn-c4-multilingual-bge-m3 Dataset Card This dataset was created with Tokenlearn for training Model2Vec models. It contains mean token embeddings produced by a sentence transformer, used as training targets for static embedding distillation. Dataset Details Field Value Source dataset allenai/c4 Source split train Embedding model BAAI/bge-m3 Embedding dimension 1024 Rows 9999954 Dataset Structure Column Type… See the full description on the dataset page: https://huggingface.co/datasets/minishlab/tokenlearn-c4-multilingual-bge-m3.text10M<n<100M2 likes273 downloads6mo agoHugging Face